June 2026
It started with
eight billion.
I wanted a model that could find sensitive information in German text. Not just names and email addresses, but descriptions that identify someone indirectly. That became Valravn: an 8B model fine-tuned for anonymisation, originally called Corvus.
The base was Ministral 8B Instruct, quantised to 4 bits and fine-tuned with LoRA through MLX on my Mac. I was adapting an existing language model. Its decoder read the text and generated a JSON list of spans, classes and replacement values.
A decoder generates its answer
token by token.
An encoder classifies spans
directly in the input text.
Model size, not a speed comparison. Skira 5 has exactly 140,651,161 parameters, counted from the saved tensor shapes.
Progress wasn't a straight line.
On my fixed test set of 14 documents, historical span F1 rose from 87.0 to 93.2%. Version 5 introduced the name Valravn. The focus then shifted towards longer texts. Experiments through v12 showed how little a good score on short examples says about dense, long documents.
Versions 6–11 were long-text experiments. I don't have a reliable individual score here for those versions or v13, so the gaps remain visible.
Indirect identifiers got better.
Between v12 and v14, QID F1 doubled on the historical test. Overall F1 fell slightly, from 88.6 to 88.2%. More than 3,300 new valid synthetic documents helped with a specific weakness, but didn't automatically improve every class.
Valravn remained relatively heavy. Memory and training length limited experiments on my Mac at the time. Above all, I wanted precise text boundaries without generating a free-form answer for every detection. I archived the decoder branch on 21 July.
July to August 2026
Generate less.
Look closer.
Encoders changed the task: assign a class to each token. Where does a name begin? Where does a description end? Early encoder experiments still used Valravn names before becoming the Veyra series. This was an architecture change, not just another version number.
One of the biggest lessons came from the data. When the same kind of description is annotated with its leading article in one example and without it in another, the model learns conflicting boundaries. In early GBERT experiments, consistent boundaries raised exact QID F1 from 13 to 49%. Clear conventions solved a problem that more training runs alone couldn't fix.
Veyra 1
Combined weights from two checkpoints. WiSE-FT helped keep improvements without losing everything that came before.
79.58% Macro F1 · norm2Veyra 1.1
Another weight blend. Small gains, with improvements in ORG and QID weighed against losses in other classes.
79.83% Macro F1 · norm2Veyra 2
Kept the 1.1 weights. Word boundaries, organisation names, identifiers and a consistent BIO decoder improved the evaluation.
84.49% Macro F1 · norm3Veyra 2 was still called Veyra 1.2 in the original result log. The reference was corrected too: 84.49% on norm3, then 84.54% on norm4 without changing the model. This is not an isolated pipeline improvement.
The biggest model wasn't the one I chose.
In August, I compared several encoders on the same 200 long texts. EuroBERT 610M had the highest macro F1 in that selection at 91.87%. mmBERT-small came close at 91.41% and processed 11.1 rather than 1.6 documents per second in that setup. That combination became the foundation for Skira.
| Encoder | Macro F1 | Documents / s |
|---|---|---|
| mmBERT-small | 91.41% | 11.1 |
| EuroBERT 610M | 91.87% | 1.6 |
| GBERT-large | 91.74% | 2.6 |
| mmBERT-base | 90.92% | 5.7 |
Historical comparison from 16 August: 200 long documents, Apple M4 / MPS, batch size 1, 2,048-token context. This is not a current speed benchmark for Skira 6.
August to September 2026
The generations behind
the name Skira.
Skira 1
pool_mix17The small mmBERT encoder became my new starting point. From here, the goal was to systematically find the remaining errors: missed information, cut-off descriptions and inconsistent classes.
Skira 2
pool_mix18More work on relative dates, organisation names and repeated mentions. In the shared norm22 comparison, completely missed gold spans fell from 531 to 270.
Skira 3
pool_mix22QID boundaries and other annotations became more consistent. On the same norm22 reference, 233 spans were still completely missed. Indirect identifiers in particular needed full phrases rather than isolated words.
Skira 4
4,000 documentsSkira 4 came from the 4,000-document dataset. The work leading up to it involved several intermediate checkpoints: removing generation artefacts, consistent age annotations and more complete QID boundaries. The early pool23 result below records one of those steps, not the later 4k result.
Completely missed gold spans out of 62,023. Partially cut-off spans aren't counted here. Lower is better.
Unprotected gold characters: 6,909 → 6,306, down 8.7%. Completely missed spans: 233 → 222. Same reference within each comparison; O-bias 3 and 2,048-token context.
Then one line became two branches.
The 4,000 documents belong to the Skira 4 development. Two training objectives on that dataset led to two variants: Skira 4.1 targeted exact classification. Skira 5 was trained directly to mask sensitive information as completely as possible. Both branches come from the Skira 4 work. Skira 5 does not continue the 4.1 checkpoint.
The right category.
The right text boundaries.
Catch sensitive spans.
Even tricky descriptions.
Both runs selected the checkpoint from epoch 1 of 2, using 2,799 validation documents before evaluating on Test300. The selection metric was micro F1 for 4.1 and full-document coverage for 5. Skira 5 was released on 6 September. On 7 September, the F1 branch was officially named Skira 4.1 without retraining.
Newer doesn't simply mean better. Skira 5 leaves far less sensitive information exposed, but its classes and span boundaries are less precise. That's the difference I want to show. Skira 6 brings both goals together: an additional trained label head for more precise classes, while preserving Skira 5's masking. The finished result is at the end of this page.