Two different things are measured here. Detection quality decides whether a release may ship; timing is reviewed rather than gated.
uv run --extra ml --with datasets python benchmarks/evaluate_quality.py \
--samples 1000 --ml --dataset-revision a785eb528e28be2693c3718a27e066970de5dadb --output result.json--allow-unverified-checksums is available only for synthetic-corpus diagnostics where invalid
checksum-shaped values are intentional. It is not enabled by default, does not alter normal library
processing, and results produced with it must not be compared to the strict published baseline.
Scored against the ai4privacy/pii-masking-openpii-1.5m validation split (revision a785eb528e28be2693c3718a27e066970de5dadb), English rows,
shuffled with a fixed seed.
The revision is required and the JSON result records its inputs, corpus or model SHA-256 values,
environment, aggregate counts, per-entity counts, and metrics.
How a number is produced matters as much as the number:
- Each detection is paired with at most one annotation and each annotation with at most one detection, taking the largest overlap first. Precision and recall therefore count the same unit on both sides.
- A pair must agree on entity type. Finding an email where the corpus annotated a surname
is not a hit.
--span-onlyscores any overlap regardless of type, which is looser and is reported separately when quoted. - Annotations carrying a label outside
SUPPORTED_LABELSare not scored. A detection landing on one is counted as unscored rather than as a false positive, because the package is not claiming to detect that label here and should not be penalised for being right about it. - Develop against the
trainsplit withbenchmarks/train_eval.py, which runs the same scoring code and prints every miss and false positive. The validation split is for measurement only.
Figures published before this scoring was corrected counted true positives per detection against false negatives per annotation, and ignored entity types entirely. They are not comparable to figures produced afterwards, and a release comparing against them is comparing against a different measurement rather than a different implementation.
The released code, unchanged, measured under the old and the corrected scoring. 300 validation rows with the ONNX backend enabled.
| Version | Scoring | Precision | Recall | F1 |
|---|---|---|---|---|
0.17.0 |
Corrected, --span-only |
0.9316 | 0.7529 | 0.8328 |
0.17.0 |
Corrected, entity types compared | 0.8094 | 0.6536 | 0.7232 |
1.0.0 |
One-to-one overlap and entity-type match (1000 rows) | 0.9141 | 0.5465 | 0.6840 |
1.10.0 |
One-to-one overlap and entity-type match (1000 rows) | 0.8425 | 0.6293 | 0.7205 |
1.20.0 |
One-to-one overlap and entity-type match (1000 rows) | 0.8587 | 0.8016 | 0.8292 |
1.26.0 |
One-to-one overlap and entity-type match (1000 rows) | 0.8611 | 0.8026 | 0.8308 |
1.34.0 |
One-to-one overlap and entity-type match (1000 rows) | 0.8608 | 0.8026 | 0.8307 |
0.19.0 |
Corrected, --span-only |
0.9317 | 0.7542 | 0.8336 |
0.19.0 |
Corrected, entity types compared | 0.8097 | 0.6549 | 0.7241 |
The gap is the measurement, not a change in behaviour. The 90% threshold that releases
0.14.0 onwards were gated against was never cleared under scoring that counts one unit
on both sides of the ratio and checks that a detection agrees with the annotation it is
credited for.
The two rows per version also show what a release is worth measuring against. Under
corrected scoring 0.18.0 and 0.19.0 together moved F1 by 0.0009, two additional true
positives out of 1611 annotations, which is within run-to-run noise. The previously
reported figures put the same interval at 0.9184 to 0.9224 and attributed it to a
release. A gate that cannot distinguish a change from noise cannot gate anything, which
is why per-entity counts belong in the published table alongside the aggregate.
You can evaluate quality on your own local datasets using the built-in benchmarking harness:
python -m pseudonymize.bench path/to/your/dataset.jsonl --mlThe dataset file must be in JSONL format, where each line represents a row with the following schema:
source_text: The raw text to process.privacy_mask: A list of ground-truth annotations. Each annotation is a dictionary withstart,end, andlabel(e.g.,EMAIL,PERSON,LOCATION).
By default, the harness will calculate both global performance (Precision, Recall, F1) and print a detailed Per-Entity Metrics table. This is extremely useful for verifying exactly which types of PII are passing successfully versus failing on your specific data distribution.
Run uv run pytest benchmarks --benchmark-only. Record CPU, operating system, Python version,
input construction, rounds, package revision, wheel size, import time, latency, and memory before
publishing results. Timing regressions are reviewed rather than gated on noisy shared runners.
Measured on the 0.1.0b1 release candidate with Python 3.14.3, Windows 11 Home 10.0.26200,
an Intel Core Ultra 7 155H (16 cores, 22 logical processors), and 31.5 GiB RAM. The engine input
repeats the committed synthetic message corpus to the named size and uses deterministic mode.
Pytest-benchmark ran at least 20 rounds with garbage collection disabled.
| Measurement | Result |
|---|---|
| 4 KiB processing, median | 8.05 ms |
| 4 KiB processing, mean | 7.73 ms |
| 64 KiB processing, median | 282.49 ms |
| 64 KiB processing, mean | 269.77 ms |
Import, median of 10 -X importtime processes |
78.93 ms |
| Import peak traced allocation | 4.49 MiB |
| 4 KiB processing peak traced allocation | 217.18 KiB |
| 64 KiB processing peak traced allocation | 2.48 MiB |
| Wheel size | 40,222 bytes |
Peak allocations use tracemalloc in a fresh process with the built wheel installed. The current
aspirational latency and import budgets are not met on this Windows reference machine; they remain
optimization targets rather than release gates. The wheel remains below its enforced 250 KiB
limit.
The release candidate was remeasured on the same Windows machine and Python 3.14.3 environment.
There are no core implementation changes from 0.1.0b1.
| Measurement | Result |
|---|---|
| 4 KiB processing, median | 7.72 ms |
| 4 KiB processing, mean | 7.57 ms |
| 64 KiB processing, median | 336.78 ms |
| 64 KiB processing, mean | 332.13 ms |
| Isolated import with allocation tracing, two runs | 234.90–446.66 ms |
| Import peak traced allocation | 1.92 MiB |
| Wheel size | 40,235 bytes |
The instrumented import range is not comparable to the beta's uninstrumented import-time sample. Its variance and the 64 KiB variance reinforce the decision to review timing without making it a release gate.