Skip to content

Latest commit

 

History

History
62 lines (49 loc) · 5.74 KB

File metadata and controls

62 lines (49 loc) · 5.74 KB

pseudonymize maintainer instructions

Project

This repository is a Python 3.11+ library and CLI for local-first PII pseudonymization for text, structured data, and LLM payloads. uv owns dependencies, packaging, scripts, and the lock file. The public package is pseudonymize.

Read the relevant implementation, tests, and documentation before editing. Keep changes focused and preserve existing public interfaces, defaults, JSON keys, environment variables, console scripts, and tool signatures through the 0.1.x series as required by docs/compatibility.md.

  • Strict PII Scope: The sole purpose of this package is PII pseudonymization. Any ML capabilities (like LocalONNXPIIBackend) must be explicitly framed, designed, and tested exclusively for detecting and pseudonymizing PII. Never treat, use, or document ML detection as a general-purpose NLP or entity extraction tool.

Development

  • Use the existing style and the simplest working implementation.
  • Add a regression test for every bug fix and offline tests for new behavior.
  • Update user-facing documentation and CHANGELOG.md for behavior changes.
  • Do not edit uv.lock manually; use uv commands.
  • Never expose or commit .env contents, passwords, credentials, tokens, captured request headers, or local service-account files.
  • Preserve user changes in a dirty worktree. Do not reset, restore, or delete unrelated work.
  • File Deletion Pre-Approved: You are explicitly authorized to remove, overwrite, or delete files and directories inside this workspace (e.g., build artifacts, debug scripts, caches) without asking for my permission. You do not need to pause or seek confirmation before executing file deletion commands.
  • Never use dummy models, fake artifacts, or excessive mocking for integration tests. If an external model or binary is needed to test an inference pipeline, write a setup script or test fixture to dynamically download a real, lightweight version (e.g., a quantized BERT model) to a local cache directory excluded from version control. Tests must exercise actual logic against real weights.
  • Strict Benchmark Integrity & Zero-Tolerance for Cheating: You are absolutely forbidden from "cheating" the quality benchmarks. You must NEVER disable, bypass, or comment out validation logic (such as Luhn checks, checksums, or structural parsers) just to artificially inflate recall. You must NEVER hardcode strings, names, or regexes designed specifically to patch failures found only in the benchmark dataset. You must NEVER dynamically inflate the engine at runtime using the evaluation or train splits (e.g. injecting benchmark data into a dictionary/DAWG during the test run). You must NEVER fine-tune a model on the benchmark evaluation slice.
  • BLIND EVALUATION: You must never inspect, print, or analyze the evaluation/holdout dataset (validation split of ai4privacy/pii-masking-openpii-1.5m) to find missed edge cases. If you need to debug False Positives or False Negatives to build heuristics, you MUST use the train split exclusively. All improvements to precision, recall, and F1 must stem from generalized heuristics, better token alignment, and robust ML calibration that apply to unseen, real-world text.
  • Adversarial Testing: New tests must include extremely hard, adversarial edge cases (e.g., bidirectional overrides, deep nesting, escaped structures) to challenge the parsers and detectors.
  • LRU Cache Test Isolation: When mocking or monkeypatching configurations on a backend using fast-path LRU caching (e.g., LocalONNXPIIBackend._infer_text_cached), always call .cache_clear() on the cache method right before executing tests on the same text to prevent stale cache entries from contaminating test assertions.

Install and validate with:

uv sync --all-extras --all-groups --frozen
uv run ruff format --check .
uv run ruff check .
uv run mypy
uv run pytest

Releases

Follow docs/releasing.md exactly. Use /release:prepare X.Y.Z to prepare a candidate and /release:verify to validate it.

  • Keep version files unchanged during ordinary development.
  • For a release candidate, synchronize pyproject.toml, prepend CHANGELOG.md, and update README examples using only verified changes.
  • Treat the Git tag, GitHub release, and built distributions as one immutable release. Never reuse a published version.
  • You have explicit standing maintainer authorization to commit, tag, push, create GitHub releases, publish to PyPI, run tests, and check, approve, or close PRs when requested.
  • Never bypass a failing check. Report the failure and preserve its output.
  • Strict Release Gate: For any release leading up to 1.0.0, you MUST run uv run python benchmarks/evaluate_quality.py --ml --samples 500 and empirically prove that the F1, Precision, or Recall metrics maintain our >60% baseline (the strict, non-cheating limit for the highly synthetic ai4privacy dataset). You are forbidden from continuing the release if the scores are degraded significantly.
  • Publish only from the exact commit that passed CI, using tag vX.Y.Z.

At handoff, state files changed, checks run, checks not run, and any external actions still requiring maintainer approval.

Recent Architectural Changes

  • Dense Placeholders: The engine uses short 3-letter entity codes natively (e.g. <PER_1>, <EML_1>, [REDACTED_LOC]) to tightly format tabular and PDF datasets.
  • PDF Rendering Engine: Re-engineered to explicitly draw pseudo-text over erased bounding boxes with proportional font-size down-scaling and exact origin_y baseline calculation (raised by 15% to perfectly center fallback Helvetica overlays).
  • Zero ML Dependencies: The llama-cpp-python backend and its associated scripts/tests have been completely and permanently eradicated in favor of lightweight ONNX logic.