leminkozey

Themistic / My own NLP model

Skira 6.

Find sensitive information. Keep the context.

I develop my own multilingual NLP models. Skira 6 detects names, contact details and descriptions that can identify someone without naming them.

Try it yourself
How Valravn became SkiraThe full story. Generations, experiments and results.

June 2026

It started with
eight billion.

I wanted a model that could find sensitive information in German text. Not just names and email addresses, but descriptions that identify someone indirectly. That became Valravn: an 8B model fine-tuned for anonymisation, originally called Corvus.

The base was Ministral 8B Instruct, quantised to 4 bits and fine-tuned with LoRA through MLX on my Mac. I was adapting an existing language model. Its decoder read the text and generated a JSON list of spans, classes and replacement values.

Valravn8 B

A decoder generates its answer
token by token.

Skira 5141 M

An encoder classifies spans
directly in the input text.

Model size, not a speed comparison. Skira 5 has exactly 140,651,161 parameters, counted from the saved tensor shapes.

Progress wasn't a straight line.

On my fixed test set of 14 documents, historical span F1 rose from 87.0 to 93.2%. Version 5 introduced the name Valravn. The focus then shifted towards longer texts. Experiments through v12 showed how little a good score on short examples says about dense, long documents.

Valravn's documented results Historical span F1 · 14 documents · 0–100% scale

Versions 6–11 were long-text experiments. I don't have a reliable individual score here for those versions or v13, so the gaps remain visible.

21 % → 42 %

Indirect identifiers got better.

Between v12 and v14, QID F1 doubled on the historical test. Overall F1 fell slightly, from 88.6 to 88.2%. More than 3,300 new valid synthetic documents helped with a specific weakness, but didn't automatically improve every class.

Valravn remained relatively heavy. Memory and training length limited experiments on my Mac at the time. Above all, I wanted precise text boundaries without generating a free-form answer for every detection. I archived the decoder branch on 21 July.

July to August 2026

Generate less.
Look closer.

Encoders changed the task: assign a class to each token. Where does a name begin? Where does a description end? Early encoder experiments still used Valravn names before becoming the Veyra series. This was an architecture change, not just another version number.

One of the biggest lessons came from the data. When the same kind of description is annotated with its leading article in one example and without it in another, the model learns conflicting boundaries. In early GBERT experiments, consistent boundaries raised exact QID F1 from 13 to 49%. Clear conventions solved a problem that more training runs alone couldn't fix.

Veyra 1

Combined weights from two checkpoints. WiSE-FT helped keep improvements without losing everything that came before.

79.58% Macro F1 · norm2

Veyra 1.1

Another weight blend. Small gains, with improvements in ORG and QID weighed against losses in other classes.

79.83% Macro F1 · norm2

Veyra 2

Kept the 1.1 weights. Word boundaries, organisation names, identifiers and a consistent BIO decoder improved the evaluation.

84.49% Macro F1 · norm3

Veyra 2 was still called Veyra 1.2 in the original result log. The reference was corrected too: 84.49% on norm3, then 84.54% on norm4 without changing the model. This is not an isolated pipeline improvement.

The biggest model wasn't the one I chose.

In August, I compared several encoders on the same 200 long texts. EuroBERT 610M had the highest macro F1 in that selection at 91.87%. mmBERT-small came close at 91.41% and processed 11.1 rather than 1.6 documents per second in that setup. That combination became the foundation for Skira.

EncoderMacro F1Documents / s
mmBERT-small91.41%11.1
EuroBERT 610M91.87%1.6
GBERT-large91.74%2.6
mmBERT-base90.92%5.7

Historical comparison from 16 August: 200 long documents, Apple M4 / MPS, batch size 1, 2,048-token context. This is not a current speed benchmark for Skira 6.

August to September 2026

The generations behind
the name Skira.

  1. Skira 1

    pool_mix17

    The small mmBERT encoder became my new starting point. From here, the goal was to systematically find the remaining errors: missed information, cut-off descriptions and inconsistent classes.

  2. Skira 2

    pool_mix18

    More work on relative dates, organisation names and repeated mentions. In the shared norm22 comparison, completely missed gold spans fell from 531 to 270.

  3. Skira 3

    pool_mix22

    QID boundaries and other annotations became more consistent. On the same norm22 reference, 233 spans were still completely missed. Indirect identifiers in particular needed full phrases rather than isolated words.

  4. Skira 4

    4,000 documents

    Skira 4 came from the 4,000-document dataset. The work leading up to it involved several intermediate checkpoints: removing generation artefacts, consistent age annotations and more complete QID boundaries. The early pool23 result below records one of those steps, not the later 4k result.

Fewer spans missed entirely bench1878 norm22 · 1,878 documents

Completely missed gold spans out of 62,023. Partially cut-off spans aren't counted here. Lower is better.

Boundaries matter too Skira 3 / early Skira 4 checkpoint · norm24

Unprotected gold characters: 6,909 → 6,306, down 8.7%. Completely missed spans: 233 → 222. Same reference within each comparison; O-bias 3 and 2,048-token context.

Then one line became two branches.

The 4,000 documents belong to the Skira 4 development. Two training objectives on that dataset led to two variants: Skira 4.1 targeted exact classification. Skira 5 was trained directly to mask sensitive information as completely as possible. Both branches come from the Skira 4 work. Skira 5 does not continue the 4.1 checkpoint.

Both runs selected the checkpoint from epoch 1 of 2, using 2,799 validation documents before evaluating on Test300. The selection metric was micro F1 for 4.1 and full-document coverage for 5. Skira 5 was released on 6 September. On 7 September, the F1 branch was officially named Skira 4.1 without retraining.

Newer doesn't simply mean better. Skira 5 leaves far less sensitive information exposed, but its classes and span boundaries are less precise. That's the difference I want to show. Skira 6 brings both goals together: an additional trained label head for more precise classes, while preserving Skira 5's masking. The finished result is at the end of this page.

Try a sentence. See the difference.

Compare the original and protected text. “Without names” shows how context can identify someone.

Skira 6Local demo

Original

Protected

Your own text is available when the model service is ready.

Four precomputed model outputs.

  • Skira 6 · 7 Sep 2026
  • 12 classes for sensitive information
  • Four real model outputs

Same texts. Different strengths.

Skira 4.1 / Skira 6

300 documents. 9,551 sensitive spans.
How much stays protected, and how accurate are the classes?

The cost of
a wider net.

Skira 6 also masks 58,355 characters outside the gold annotations, compared with 2,998 for Skira 4.1. These extra masks aren't automatically false positives, since references can have gaps too. They do show how much additional text is affected.

A span can be fully protected and still have the wrong class or imprecise boundaries. Coverage and label F1 answer different questions.

All twelve classes in detail

Exact class and exact text boundaries. Each class is evaluated separately.

Skira 4.1Skira 60–100 %
How I compare the results

Skira 4.1 uses its official evaluation from 6 September 2026; Skira 6 uses the final package evaluation from 7 September 2026. Test300 is a known regression set, not a new independent holdout. The reference contains 300 documents, 9,551 gold spans and 130,996 gold characters. Every document contains gold annotations.

  • Document coverage: Every annotated sensitive character in a document is protected, regardless of its class label.
  • Exact label F1: A match needs the same class and exactly the same start and end positions. Macro F1 gives equal weight to all twelve classes.
  • Skira 4.1: Viterbi, base chain and QID boundary at O-bias 0, evaluation stage 4 for labels and masking.
  • Skira 6: Labels from the trained, separate label head at bias 0. Masks come from the frozen Skira 5 masking path, including its binary PII decision at a 0.5 threshold. The masks are not rebuilt from the new labels.

The pipeline is part of the result too: the same 4.1 checkpoint reaches 279/300 fully covered documents and 91.99% macro F1 with the more aggressive stage 10 at bias 4. The figures above use the official profiles, not an isolated measurement of the model weights.

Historical test sets stay separate. The small Valravn test, changing Veyra references, norm22, norm24 and Test300 don't form one continuous progress curve.

Download all aggregated metrics as JSON ↓Sources and calculations ↗

7 September 2026 · Released

The safety net stays.
The labels get sharper.

Skira 6

My goal was clear: keep Skira 5's coverage and get closer to Skira 4.1's label quality. The finished model reaches 94.68% label macro F1. A separate label head was trained on the existing 4,000 documents. The encoder and original masking head stay frozen.

Validation · 2,799 documents83.76 92.48%

From the decoder preview to the trained label head. Selection used only the validation set.

Test300 · Skira 5 → Skira 675.58 94.68%

Label macro F1. Micro F1: 77.37 → 95.87%.

300 / 300

Exactly the same masks.

Still 298 fully protected documents, two unprotected gold spans and 42 unprotected gold characters. The 58,355 extra masked non-gold characters also remain. Masks match across all 2,799 validation documents too.

Skira 6 is now only 1.50 percentage points below Skira 4.1's 96.17% label macro F1, while retaining Skira 5's higher document coverage. Classification quality and protection remain two separate measures.

All classes: Skira 5 and Skira 6

Exact label F1, precision and recall on Test300. PII coverage measures fully protected gold spans regardless of their label. All values are percentages.

From the preview to the finished label head

The decoder preview first brought Skira 6 to 87.64% macro F1. A bias experiment then reached 89.30%. The decisive step was a separate label head: initialised from Skira 4.1's head weights, then trained on 4,000 unchanged documents to adapt to the frozen text representations.

  • Three epochs, 157,465 trained parameters. Epoch 3 was selected by the highest label macro F1 across all 2,799 validation documents. Test300 was not used for selection.
  • One encoder, two heads. The same text representations produce better labels and unchanged masking. 390 windows, exactly 390 encoder calls in the full Test300 package run.
  • Use the masks separately. The demo uses the returned masks for protection and the new entity list for labels. It does not rebuild the masking from the labels.

Masks were compared using the finished package loader across 2,799 validation and 300 test documents: no differences. Test300 remains a known regression set, not a new independent holdout. Two existing protection gaps remain. The effect of head initialisation alone was not measured in isolation.

The interactive demo uses the finished Skira 6 label head. The full comparison results come from package checks on MPS; the Pi service is checked separately with demo inputs. This is not a new full CPU evaluation of all test documents.

Download the metrics ↓
My models on Hugging Face