Pseudonymize is the small, local-first privacy boundary between sensitive application data and the systems that should not receive it unchanged.
It starts with Python strings and nested payloads, grows through format-aware documents and local models, and permits remote processing only through an explicit security contract. The core remains typed, dependency-free, offline-capable, and network-denied by default.
Local by default. Structure-aware. Backend-agnostic. Safe to observe.
Applications increasingly move text and structured payloads through LLMs, logs, queues, analytics, document pipelines, and third-party APIs. Sensitive values can cross those boundaries long before a traditional data-loss-prevention system sees them.
Most teams need a composable library rather than another service:
- local processing for common identifiers;
- stable aliases when repeated identity matters;
- redaction when identity does not matter;
- safe reports for audit and testing;
- format-aware locations without coupling detectors to file libraries;
- optional stronger detection without forcing models or network clients on every installation.
Pseudonymize should make that boundary explicit and testable.
- Keep the trusted core small. Text, nested data, policies, structured detectors, transformations, reports, and extension contracts remain standard-library only.
- Stay local unless the caller opts in twice.
NetworkPolicy.DENYis the default. Remote work requires both policy permission and backend-level consent. - Use native structure before inference. Traverse JSON paths directly, preserve document structure and coordinates, prefer native PDF text, and use OCR only when required.
- Separate format handling from privacy policy. Adapters extract and render. Backends detect. Policies filter. Transformers replace. No layer silently takes ownership of another.
- Never make observability another leak. Reports, statistics, warnings, exceptions, logs, CLI diagnostics, and representations must not copy matched source values.
- Make identity behavior explicit. Numbered, generic, deterministic, and redacted modes have different semantics. Reversible mappings are always opt-in and treated as sensitive data.
- Add formats only when they can be handled safely. Inspection precedes rewriting. A rendered PDF is accepted only when underlying content is removed, not visually covered.
- Keep optional dependencies narrow. Models, PDF, Office, OCR, Docling, and HTTP libraries live only in the extras that own them and never load during a base import.
- Earn compatibility after the architecture is proven. Alpha releases may replace public contracts without shims. Compatibility begins only at the documented stable milestone.
- Raise the verification bar with every release. Coverage never falls below the tagged baseline. New tests must add realistic workflows, interacting invariants, malformed inputs, and adversarial failure modes rather than merely execute new lines. Meaningful testing with real artifacts is strictly mandated. No dummy models, fake data generation that bypasses logic, or mock-based workarounds are permitted. External models required for testing must be downloaded dynamically and cached outside of version control.
- Strictly bounded to PII. The application of ML models is strictly constrained to PII pseudonymization. This package is not and will never be a general-purpose NLP or entity extraction framework.
source
-> input adapter
-> immutable document and content blocks
-> local or explicitly permitted detection backends
-> policy, overlap resolution, identity resolution, transformation
-> safe result and optional output adapter
Detectors operate on ContentBlock values, not on file formats. A JSON string, CSV cell, DOCX
paragraph, PDF text span, and OCR bounding box can therefore share detection logic while retaining
their own typed source locations.
The committed progression is:
- Text and nested JSON-compatible Python values.
- Dependency-free text, JSON, JSONL, and CSV files.
- Optional local ML for names, organizations, locations, and contextual addresses.
- Detection-only PDF and Office inspection.
- Format-preserving document output and secure PDF redaction.
- Local OCR for images, scanned PDFs, and mixed documents.
- Explicitly configured remote providers with bounded transport behavior.
Native structure always wins over a more expensive or lossy modality.
- Python developers building privacy-aware application and data pipelines
- Teams placing a local redaction boundary before hosted LLMs
- Security and privacy engineers who need inspectable, deterministic behavior
- Library authors implementing specialized document adapters or detection backends
- Researchers who need a small baseline with explicit limitations and reproducible tests
Pseudonymize does not compete on the number of extensions it claims to accept. It competes on:
- a dependency-free, typed core;
- deterministic, namespace-isolated HMAC aliases;
- nested-payload and LLM-friendly processing;
- typed source locations across modalities;
- reports designed not to contain raw detections;
- explicit remote consent with no vendor SDK in the base package;
- increasingly adversarial, cross-platform release gates.
Pseudonymization does not guarantee anonymization, complete PII detection, regulatory compliance, or immunity from contextual re-identification.
The committed roadmap excludes audio, video, reversible token vaults, database engines, Parquet, SQLite, framework-specific wrappers, and generic claims to process every file. These can be reconsidered only after the core roadmap is proven and should not distort the current architecture.
The project succeeds when an application can place a narrow local boundary around sensitive content, choose the appropriate transformation, inspect what happened without leaking what was found, and add only the format or detection dependencies it actually needs.
The following architectural and visionary proposals are under consideration for the long-term future:
- #34 Vision proposal: worldwide language and locale coverage
- #35 Vision proposal: maintained structured registry of privacy regulations per jurisdiction
- #36 Architecture and scalability review: current limits and proposed solutions
- #66 Proposal: training-corpus de-identification mode with erasure-by-key (crypto-shredding)
- #70 Proposal: residual re-identification risk report (the honest dual of detection)
- #71 Proposal: processing receipts (signed, value-free provenance manifests)
- #73 Proposal: pseudonym interchange spec with cross-language test vectors