This repository is a Python 3.11+ library and CLI for local-first PII pseudonymization for text, structured data, and LLM payloads. uv owns dependencies, packaging, scripts, and the lock file. The public package is pseudonymize.
Read the relevant implementation, tests, and documentation before editing.
Keep changes focused and preserve existing public interfaces, defaults, JSON
keys, environment variables, console scripts, and tool signatures through
the 0.1.x series as required by docs/compatibility.md.
- Strict PII Scope: The sole purpose of this package is PII pseudonymization. Any ML capabilities (like
LocalONNXPIIBackend) must be explicitly framed, designed, and tested exclusively for detecting and pseudonymizing PII. Never treat, use, or document ML detection as a general-purpose NLP or entity extraction tool.
- Use the existing style and the simplest working implementation.
- Add a regression test for every bug fix and offline tests for new behavior.
- Update user-facing documentation and
CHANGELOG.mdfor behavior changes. - Do not edit
uv.lockmanually; useuvcommands. - Never expose or commit
.envcontents, passwords, credentials, tokens, captured request headers, or local service-account files. - Preserve user changes in a dirty worktree. Do not reset, restore, or delete unrelated work.
- File Deletion Pre-Approved: You are explicitly authorized to remove, overwrite, or delete files and directories inside this workspace (e.g., build artifacts, debug scripts, caches) without asking for my permission. You do not need to pause or seek confirmation before executing file deletion commands.
- Never use dummy models, fake artifacts, or excessive mocking for integration tests. If an external model or binary is needed to test an inference pipeline, write a setup script or test fixture to dynamically download a real, lightweight version (e.g., a quantized BERT model) to a local cache directory excluded from version control. Tests must exercise actual logic against real weights.
- Strict Benchmark Integrity & Zero-Tolerance for Cheating: You are absolutely forbidden from "cheating" the quality benchmarks. You must NEVER disable, bypass, or comment out validation logic (such as Luhn checks, checksums, or structural parsers) just to artificially inflate recall. You must NEVER hardcode strings, names, or regexes designed specifically to patch failures found only in the benchmark dataset. You must NEVER dynamically inflate the engine at runtime using the evaluation or train splits (e.g. injecting benchmark data into a dictionary/DAWG during the test run). You must NEVER fine-tune a model on the benchmark evaluation slice.
- BLIND EVALUATION: You must never inspect, print, or analyze the evaluation/holdout dataset (
validationsplit ofai4privacy/pii-masking-openpii-1.5m) to find missed edge cases. If you need to debug False Positives or False Negatives to build heuristics, you MUST use thetrainsplit exclusively. All improvements to precision, recall, and F1 must stem from generalized heuristics, better token alignment, and robust ML calibration that apply to unseen, real-world text. - Adversarial Testing: New tests must include extremely hard, adversarial edge cases (e.g., bidirectional overrides, deep nesting, escaped structures) to challenge the parsers and detectors.
- LRU Cache Test Isolation: When mocking or monkeypatching configurations on a backend using fast-path LRU caching (e.g.,
LocalONNXPIIBackend._infer_text_cached), always call.cache_clear()on the cache method right before executing tests on the same text to prevent stale cache entries from contaminating test assertions.
Install and validate with:
uv sync --all-extras --all-groups --frozen
uv run ruff format --check .
uv run ruff check .
uv run mypy
uv run pytestFollow docs/releasing.md exactly. Use /release:prepare X.Y.Z to prepare a
candidate and /release:verify to validate it.
- Keep version files unchanged during ordinary development.
- For a release candidate, synchronize
pyproject.toml, prependCHANGELOG.md, and update README examples using only verified changes. - Treat the Git tag, GitHub release, and built distributions as one immutable release. Never reuse a published version.
- You have explicit standing maintainer authorization to commit, tag, push, create GitHub releases, publish to PyPI, run tests, and check, approve, or close PRs when requested.
- Never bypass a failing check. Report the failure and preserve its output.
- Strict Release Gate: For any release leading up to 1.0.0, you MUST run
uv run python benchmarks/evaluate_quality.py --ml --samples 500and empirically prove that the F1, Precision, or Recall metrics maintain our >60% baseline (the strict, non-cheating limit for the highly synthetic ai4privacy dataset). You are forbidden from continuing the release if the scores are degraded significantly. - Publish only from the exact commit that passed CI, using tag
vX.Y.Z.
At handoff, state files changed, checks run, checks not run, and any external actions still requiring maintainer approval.
- Dense Placeholders: The engine uses short 3-letter entity codes natively (e.g.
<PER_1>,<EML_1>,[REDACTED_LOC]) to tightly format tabular and PDF datasets. - PDF Rendering Engine: Re-engineered to explicitly draw pseudo-text over erased bounding boxes with proportional font-size down-scaling and exact
origin_ybaseline calculation (raised by 15% to perfectly center fallback Helvetica overlays). - Zero ML Dependencies: The
llama-cpp-pythonbackend and its associated scripts/tests have been completely and permanently eradicated in favor of lightweight ONNX logic.