Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,11 @@ All notable changes follow Keep a Changelog and Semantic Versioning.

## [Unreleased]

### Fixed

- Align README aliases with the emitted token format and execute its Python examples
in an integration test. Clarify overlap benchmark scoring and streaming limitations.

## [1.34.0] - 2026-09-30

### Added
Expand Down
24 changes: 24 additions & 0 deletions HANDOVER.md
Original file line number Diff line number Diff line change
Expand Up @@ -399,3 +399,27 @@ verification unless they are intentionally retained as a user-approved release a
- Push only when explicitly requested.
- A version is published only after its matching tag, successful release workflow, PyPI artifact,
and GitHub release exist. Until then, keep changes under `[Unreleased]`.


## Executable README review (2026-10-02)

The first Python quickstart failed its assertion against the current package:
README examples used long entity names while the implementation emits short codes.
Corrected aliases and added a test that executes all Python fences in order,
providing the documented JSON file in a temporary directory. The asynchronous
example is compiled and defined by that test; existing streaming tests exercise
actual asynchronous processing. No token format or library behavior changed.

The historical 1.20.0 score is now correctly labeled as an overlap result, and the
streaming introduction no longer promises arbitrary split-entity safety. These
clarifications reflect the existing scorer and heuristic segment splitter.

Observed verification with the frozen lockfile:

- README, streaming and session tests: 11 passed.
- Suite excluding ONNX/BIO modules with `--no-cov`: 490 passed; 1 Tesseract skip.
- Pre-commit, mypy (156 source files), strict docs build, package build and release
verifier passed. Full coverage and platform evidence remain CI responsibilities.

Reusable lesson assessment: executable onboarding assertions catch token-contract
drift that prose review misses. This is enforced by the new integration test.
37 changes: 20 additions & 17 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ payloads.
```text
Email paolo@example.com from 192.0.2.10.
↓
Email <EMAIL_1> from <IP_ADDRESS_1>.
Email <EML_1> from <IP_1>.
```

Pseudonymize detects structured sensitive values locally and transforms them into numbered,
Expand Down Expand Up @@ -55,13 +55,13 @@ no telemetry or model downloads, and denies remote-capable backends by default.

The engine is strictly gated on detection accuracy against the `ai4privacy/pii-masking-openpii-1.5m` dataset (validation split, 1000 randomly sampled rows).

**Current Baseline (1.20.0, 1000 rows, strict boundaries and type matching, ONNX backend enabled):**
**Historical baseline (1.20.0, 1,000 rows, one-to-one positive boundary overlap and exact entity-type matching, ONNX backend enabled):**
- **Precision:** 0.8587
- **Recall:** 0.8016
- **F1 Score:** 0.8292

Each detection is paired with at most one annotation and must agree with it on entity
type. Figures published before the scoring was strictly corrected are not comparable; see
type. These figures measure overlapping spans, not exact-boundary accuracy. Figures published before the scoring was strictly corrected are not comparable; see
[docs/benchmarks.md](docs/benchmarks.md).

## Installation
Expand Down Expand Up @@ -116,7 +116,7 @@ from pseudonymize import pseudonymize, redact
safe = pseudonymize("Email paolo@example.com")
hidden = redact("Email paolo@example.com")

assert safe == "Email <EMAIL_1>"
assert safe == "Email <EML_1>"
assert hidden == "Email [REDACTED]"
```

Expand All @@ -129,7 +129,7 @@ from pseudonymize import Pseudonymizer

result = Pseudonymizer().process_with_report("Email paolo@example.com from 192.0.2.10.")

assert result.output == "Email <EMAIL_1> from <IP_ADDRESS_1>."
assert result.output == "Email <EML_1> from <IP_1>."
assert result.statistics.detections_found == 2
assert result.detections[0].backend == "rules"
assert "paolo@example.com" not in repr(result)
Expand All @@ -155,16 +155,20 @@ payload = {

result = Pseudonymizer(policy=Policy.llm()).process_data_with_report(payload)

assert result.output["messages"][0]["content"] == "Email <EMAIL_1>"
assert result.output["messages"][1]["content"] == "Use <EMAIL_1> again"
assert result.output["messages"][0]["content"] == "Email <EML_1>"
assert result.output["messages"][1]["content"] == "Use <EML_1> again"
assert result.output["model"] == "example-model"
```

The input is not mutated. Dictionary keys and non-string values are preserved.

### Stream LLM responses

Real-time WebSocket chunks from OpenAI or Anthropic can be processed seamlessly without risking split-entity leakage across chunks. Both `process_stream` and `process_stream_async` are available:
Both `process_stream` and `process_stream_async` buffer incoming text chunks and retain
context between emitted segments. Segment boundaries are heuristic: punctuation and long
unbroken input can split an entity, so streaming does not guarantee the same detection as
processing the complete text. Use `process()` on a complete bounded message when that
equivalence is required.

```python
import asyncio
Expand All @@ -183,20 +187,19 @@ async def handle_stream(socket):

## Transformation modes

| Mode | Example | Identity behavior |
| --- | --- | --- |
| `numbered` | `<EMAIL_1>` | Stable inside one explicit scope |
| `generic` | `<EMAIL>` | Does not distinguish values of the same type |
| `deterministic` | `<EMAIL_K8M42PX7D3Q>` | Stable for the same key, namespace, type, and normalized value |
| `redacted` | `[REDACTED]` | Removes type and identity distinction |
- `numbered`: `<EML_1>`, stable inside one explicit scope.
- `generic`: `<EML>`, without distinguishing values of the same type.
- `deterministic`: an HMAC-derived `<EML_...>` token, stable for the same key,
namespace, entity type, and normalized value.
- `redacted`: `[REDACTED]`, without type or identity distinctions.

```python
from pseudonymize import Pseudonymizer

scope = Pseudonymizer().new_scope()

assert scope.process("paolo@example.com").text == "<EMAIL_1>"
assert scope.process("maria@example.com and paolo@example.com").text == ("<EMAIL_2> and <EMAIL_1>")
assert scope.process("paolo@example.com").text == "<EML_1>"
assert scope.process("maria@example.com and paolo@example.com").text == ("<EML_2> and <EML_1>")
```

Deterministic mode uses HMAC-SHA256 and requires a key of at least 32 bytes:
Expand Down Expand Up @@ -284,7 +287,7 @@ result = Pseudonymizer().process(
include_mapping=True,
)

assert result.restore("Reply to <EMAIL_1>.") == "Reply to paolo@example.com."
assert result.restore("Reply to <EML_1>.") == "Reply to paolo@example.com."
```

Mappings contain sensitive source values. They are hidden from `repr`, never persisted by the
Expand Down
19 changes: 19 additions & 0 deletions tests/integration/test_readme_examples.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
"""Execute the README quickstart against the installed library."""

import re
from pathlib import Path
from typing import Any

import pytest


def test_readme_python_examples(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> None:
readme = Path(__file__).resolve().parents[2] / "README.md"
examples = re.findall(r"^```python\n(.*?)^```", readme.read_text(encoding="utf-8"), re.M | re.S)
assert examples, "README must contain executable Python examples"
monkeypatch.chdir(tmp_path)
(tmp_path / "requests.json").write_text('{"email": "reader@example.com"}', encoding="utf-8")
namespace: dict[str, Any] = {}
for index, example in enumerate(examples, start=1):
exec(compile(example, f"README.md:python-example-{index}", "exec"), namespace) # noqa: S102
assert "reader@example.com" not in (tmp_path / "requests.safe.json").read_text(encoding="utf-8")
Loading