diff --git a/CHANGELOG.md b/CHANGELOG.md index 229028a..6df5fb3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -22,6 +22,9 @@ All notable changes follow Keep a Changelog and Semantic Versioning. - Invalidate Bloom filter membership caches after insertion so earlier misses cannot hide newly added entries. Validate filter parameters and allocate at least one bit. + +- Align README aliases with the emitted token format and execute its Python examples + in an integration test. Clarify overlap benchmark scoring and streaming limitations. - Include the hosted-service extra in release verification and isolated installation audits, while keeping service frameworks out of the dependency-free base import. diff --git a/HANDOVER.md b/HANDOVER.md index fc72397..cf5d7b3 100644 --- a/HANDOVER.md +++ b/HANDOVER.md @@ -477,3 +477,26 @@ Observed verification on Python 3.14.2 with the frozen lockfile: Reusable lesson assessment: mutation must invalidate cached negative membership results. The regression test records that invariant in its owning implementation. + +## Executable README review (2026-10-02) + +The first Python quickstart failed its assertion against the current package: +README examples used long entity names while the implementation emits short codes. +Corrected aliases and added a test that executes all Python fences in order, +providing the documented JSON file in a temporary directory. The asynchronous +example is compiled and defined by that test; existing streaming tests exercise +actual asynchronous processing. No token format or library behavior changed. + +The historical 1.20.0 score is now correctly labeled as an overlap result, and the +streaming introduction no longer promises arbitrary split-entity safety. These +clarifications reflect the existing scorer and heuristic segment splitter. + +Observed verification with the frozen lockfile: + +- README, streaming and session tests: 11 passed. +- Suite excluding ONNX/BIO modules with `--no-cov`: 490 passed; 1 Tesseract skip. +- Pre-commit, mypy (156 source files), strict docs build, package build and release + verifier passed. Full coverage and platform evidence remain CI responsibilities. + +Reusable lesson assessment: executable onboarding assertions catch token-contract + drift that prose review misses. This is enforced by the new integration test. diff --git a/README.md b/README.md index 51ddbf8..d5466e9 100644 --- a/README.md +++ b/README.md @@ -13,7 +13,7 @@ payloads. ```text Email paolo@example.com from 192.0.2.10. ↓ -Email from . +Email from . ``` Pseudonymize detects structured sensitive values locally and transforms them into numbered, @@ -55,13 +55,13 @@ no telemetry or model downloads, and denies remote-capable backends by default. The engine is strictly gated on detection accuracy against the `ai4privacy/pii-masking-openpii-1.5m` dataset (validation split, 1000 randomly sampled rows). -**Current Baseline (1.20.0, 1000 rows, strict boundaries and type matching, ONNX backend enabled):** +**Historical baseline (1.20.0, 1,000 rows, one-to-one positive boundary overlap and exact entity-type matching, ONNX backend enabled):** - **Precision:** 0.8587 - **Recall:** 0.8016 - **F1 Score:** 0.8292 Each detection is paired with at most one annotation and must agree with it on entity -type. Figures published before the scoring was strictly corrected are not comparable; see +type. These figures measure overlapping spans, not exact-boundary accuracy. Figures published before the scoring was strictly corrected are not comparable; see [docs/benchmarks.md](docs/benchmarks.md). ## Installation @@ -122,7 +122,7 @@ from pseudonymize import pseudonymize, redact safe = pseudonymize("Email paolo@example.com") hidden = redact("Email paolo@example.com") -assert safe == "Email " +assert safe == "Email " assert hidden == "Email [REDACTED]" ``` @@ -135,7 +135,7 @@ from pseudonymize import Pseudonymizer result = Pseudonymizer().process_with_report("Email paolo@example.com from 192.0.2.10.") -assert result.output == "Email from ." +assert result.output == "Email from ." assert result.statistics.detections_found == 2 assert result.detections[0].backend == "rules" assert "paolo@example.com" not in repr(result) @@ -161,8 +161,8 @@ payload = { result = Pseudonymizer(policy=Policy.llm()).process_data_with_report(payload) -assert result.output["messages"][0]["content"] == "Email " -assert result.output["messages"][1]["content"] == "Use again" +assert result.output["messages"][0]["content"] == "Email " +assert result.output["messages"][1]["content"] == "Use again" assert result.output["model"] == "example-model" ``` @@ -170,7 +170,11 @@ The input is not mutated. Dictionary keys and non-string values are preserved. ### Stream LLM responses -Real-time WebSocket chunks from OpenAI or Anthropic can be processed seamlessly without risking split-entity leakage across chunks. Both `process_stream` and `process_stream_async` are available: +Both `process_stream` and `process_stream_async` buffer incoming text chunks and retain +context between emitted segments. Segment boundaries are heuristic: punctuation and long +unbroken input can split an entity, so streaming does not guarantee the same detection as +processing the complete text. Use `process()` on a complete bounded message when that +equivalence is required. ```python import asyncio @@ -189,20 +193,19 @@ async def handle_stream(socket): ## Transformation modes -| Mode | Example | Identity behavior | -| --- | --- | --- | -| `numbered` | `` | Stable inside one explicit scope | -| `generic` | `` | Does not distinguish values of the same type | -| `deterministic` | `` | Stable for the same key, namespace, type, and normalized value | -| `redacted` | `[REDACTED]` | Removes type and identity distinction | +- `numbered`: ``, stable inside one explicit scope. +- `generic`: ``, without distinguishing values of the same type. +- `deterministic`: an HMAC-derived `` token, stable for the same key, + namespace, entity type, and normalized value. +- `redacted`: `[REDACTED]`, without type or identity distinctions. ```python from pseudonymize import Pseudonymizer scope = Pseudonymizer().new_scope() -assert scope.process("paolo@example.com").text == "" -assert scope.process("maria@example.com and paolo@example.com").text == (" and ") +assert scope.process("paolo@example.com").text == "" +assert scope.process("maria@example.com and paolo@example.com").text == (" and ") ``` Deterministic mode uses HMAC-SHA256 and requires a key of at least 32 bytes: @@ -290,7 +293,7 @@ result = Pseudonymizer().process( include_mapping=True, ) -assert result.restore("Reply to .") == "Reply to paolo@example.com." +assert result.restore("Reply to .") == "Reply to paolo@example.com." ``` Mappings contain sensitive source values. They are hidden from `repr`, never persisted by the diff --git a/tests/integration/test_readme_examples.py b/tests/integration/test_readme_examples.py new file mode 100644 index 0000000..e385683 --- /dev/null +++ b/tests/integration/test_readme_examples.py @@ -0,0 +1,19 @@ +"""Execute the README quickstart against the installed library.""" + +import re +from pathlib import Path +from typing import Any + +import pytest + + +def test_readme_python_examples(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> None: + readme = Path(__file__).resolve().parents[2] / "README.md" + examples = re.findall(r"^```python\n(.*?)^```", readme.read_text(encoding="utf-8"), re.M | re.S) + assert examples, "README must contain executable Python examples" + monkeypatch.chdir(tmp_path) + (tmp_path / "requests.json").write_text('{"email": "reader@example.com"}', encoding="utf-8") + namespace: dict[str, Any] = {} + for index, example in enumerate(examples, start=1): + exec(compile(example, f"README.md:python-example-{index}", "exec"), namespace) # noqa: S102 + assert "reader@example.com" not in (tmp_path / "requests.safe.json").read_text(encoding="utf-8")