You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
v1 ships the core, the CLI and the agent skills. This issue tracks the path to everything else, in the order that costs least and teaches most. The full document is misc/docs/roadmap.md; this is the tracking summary.
The claim each phase has to keep true: every addition is an ADD at a seam that already exists, and the core does not change. If a phase needs a core change, the core was coupled to its first surface and the core is what gets fixed.
The seams (decided in v1, not revisited)
Seam
v1 default
Replacement
segmenter=
"paragraph" (also "sentence", "document", or a callable)
imbed.fixed_step_chunker
detectors=
tells, forensic, rhetoric, rhythm — deterministic, zero extra deps
Landed in #2. The deterministic layer has high precision and low recall: 7 signals across the 12 known-mixed fixture documents at sentence granularity (the "8" above was stale — corrected in CLAUDE.md). Recall is the problem this phase existed to fix, and it is still the problem.
fast_detect_gpt(text, span, *, model=...) — conditional probability curvature, one sampling pass
binoculars(text, span, *, observer=..., performer=...) — cross-perplexity of a paired observer/performer
Deferred imports; import ductus loads no torch, asserted in a subprocess test
Default proxy models CPU-sized (gpt2, distilgpt2/gpt2), with EleutherAI/gpt-neo-2.7B and the Falcon-7B pair documented as the upgrade
Gate: measured, NOT cleared. Machine recall 1/40 → 3/40, machine precision 1/1 → 3/8. DEFAULT_DETECTORS unchanged, and is now a separate list from the DETECTORS registry so registering can never silently promote. Numbers: misc/docs/phase-1-results.md; re-run with python misc/measure_detectors.py
Decision written down before implementation: a continuous score becomes a banded signal with a dead zone, scored relative to the document, never a linear weight — misc/docs/curvature-as-evidence.md
Carried into Phase 2: the model detectors found 10 human-leaning segments the deterministic set misses entirely, all 10 correct, and every false machine flag was plain human prose (casual email, wire-service sport) — the Liang et al. failure mode behaving as predicted. Cheapest untried lever is the upgrade models; the fixture's fill_gaps shape is close to a worst case for these methods.
Phase 2 — calibration and a second fixture set — done (#3)
RoFT slice vendored (tests/fixtures/roft_boundary.json, MIT) — 18 documents, two at each of the nine boundary positions. It earned its keep immediately: the model-based detectors clear the Phase 1 gate on it outright (12/75 at 86% precision vs 0/75 for the deterministic set), confirming that LLMTrace's fill_gaps shape was close to their worst case
FPR on human-written text measured — the headline. 350 texts from BEA-2019 W&I+LOCNESS: 34.3% of human documents falsely accused, now 20.6%. A flagged sentence is far better evidence than a flagged document (2–3% per sentence)
FPR by proficiency band, against a native control. W&I+LOCNESS carries CEFR levels, which is sharper than one aggregate. Corpus is non-redistributable, so it is measured, not vendored — script downloads it, numbers are committed
aggregate= widened to take n_chars (the seam could not previously express a rate) and density_aggregate is the new default
Fit per-detector weights — deliberately not done. Seven deterministic signals across twelve documents is not an evidence base; fitting on it would produce the shape of a result with none of the content
Still no percentage. And the reassurance in this issue was wrong: lean = 2p − 1 is invertible, so relabelling a fitted probability does not stop it being one. Replaced by calibrate the instrument, not the verdict, plus a sufficiency test
Two findings that were not expected:
Most of the old false-positive rate was a length artefact. Pooled, it ran from 11% (under 750 chars) to 74% (over 2500) — a human document was ~7× more likely to be accused for being long.
The bias runs the opposite way from the literature. Length-matched, FPR rises with proficiency (A 31%, B 40%, C 47%), and the native control was the most-accused group. The deterministic rules fire on triad, not-x-but-y, discourse-opener — formal argumentative register. This package is biased against good writing, not against non-native writing. The shipped agent skills said otherwise and were corrected.
ductus/mcp.py — mk_mcp() over py2mcp.mk_mcp_from_refs, string refs, core never imports MCP. pip install 'ductus[mcp]' then ductus-mcp. The core did not change, as predicted
One list, and better than a parity test: TOOL_REFS is derived from tools._dispatch_funcs, so there is nothing that could drift and nothing for a parity test to catch. A verb that changes the host declares it at its own definition (@host_mutating) — so install_skills is unreachable over MCP without anyone hand-writing a second list
middleware= and auth= pass straight through for a deployed connector
The server's instructions carry the limits, since an MCP client reads that and nothing else
The first time a surface told us anything. Under from __future__ import annotations — every module here — the schema layer beneath fastmcp reads annotations as strings and drops every keyword-only default, so gauge(source=...) failed with five "missing required argument" errors. This package's convention is keyword-only from the 2nd/3rd argument, so every verb was uncallable. Narrowed to a minimal repro and filed upstream as i2mint/py2mcp#12; worked around in ductus.mcp._resolve_annotations with a test pinning it. The tool list and JSON schema both looked perfectly correct — it only appeared on an actual call, which is why this surface is tested by driving a real client rather than by inspecting a schema.
Taken before Phases 4 and 5: a UI on top of a one-in-five false-accusation rate ships the problem wider. Decision and all three gates fixed before the numbers: document-verdict-decision.md. Results: reducing-false-accusations.md.
20.6% → 6.0% human documents falsely accused (sentence); 12.0% → 6.3% on the default paragraph path; per-segment 2–3% → 0.4–2%
Nothing traded away — not one correctly-flagged machine segment lost on either fixture (LLMTrace 1/40 at 1/1, RoFT 0/75), and RoFT lost its one false flag
Per-rule table (misc/measure_rules.py): three rules were finding nothing — not-x-but-y (42 human docs, 0 TP), summary-closer (25, 0), contrastive-negation (19, 0). Two were broken, matching ordinary negation ("not eat from the tree but") and the "not only X but also Y" correlative
Tier and weight separated. A per-rule weight: in tells.yaml changes ductus only — acquaint reads tier, never weight, verified in its source. summary-closer keeps tier E, drops to weight 0.15
Document verdict is now a function of the segment verdicts (score.roll_up). Pooling gave 1−0.97ⁿ — a coin toss by 30 segments — treating accumulation as corroboration. lean unchanged; strength judged by rate and density, believed at its weakest
Cross-package landed, not left drifted: the pattern narrowing reaches acquaint, whose suite was run (2 failures, both encoding the old over-broad behaviour), fixed properly plus a regression test — 555 pass, ductus>=0.0.8 pinned
Two wrong turns, both caught by measuring rather than reasoning, both now pinned by tests: rate-only scoring made the default segmenter worse (12.0% → 13.7%), and computing lean from segment labels silently discarded sub-threshold human-leaning evidence.
Deliberately not done.triad (96 human docs, 2 TP), discourse-opener (88, 4) and exclamation (54, 8) are the dominant remaining cost with terrible cost-to-benefit — but they survive the pre-registered gate, and inventing a second criterion after seeing which rules the first one missed is the move this project has avoided three times. Evidence recorded for whoever sets one in advance. Also open, and flagged rather than settled: whether tier E is right for summary-closer in acquaint — that is a question about style advice, and the measurement is about authorship.
Built once the frontend became the remote consumer this phase was waiting for.
qh.mk_app over the same callables. ROUTED_FUNCS is derived from tools._dispatch_funcs, as TOOL_REFS is — three surfaces, one list, still no parity test to write
qh.export_ts_client for the frontend's typed client, plus export_types() for the report shape
The core did not change, as predicted. import ductus pulls in neither FastAPI nor qh, asserted in a subprocess test
Checked for the Phase 3 defect first, by driving a real client. qh does not have it — keyword-only defaults survive and every verb is callable with only its required arguments
What this surface found, and the other two could not. A verb list that is safe at a CLI is not automatically safe when the caller is a stranger. gauge(source=...) reads a file when the string names one, gauge(out=...) writes one, judgments= is a third — correct when you typed the command yourself, an arbitrary file read and an arbitrary file write when you did not. Neither the CLI nor a local stdio MCP host can see it, because on those surfaces it is not a bug.
Fixed the way host_mutating already works, not by inventing a mechanism: the verb declares it at its own definition (@host_paths(source="read", out="write")) and the adapter refuses by reading that declaration, knowing nothing about gauge. read is refused on exactly the condition under which _read_source would open a file, so the branch is unreachable rather than guessed at. mk_app(guard_host_paths=False) is the seam for a loopback service you run for yourself.
Two stale honesty claims corrected, both of which had survived the phase that made them wrong: render.py's footer still said detectors "over-flag non-native English" (this package's measured bias runs the other way — the native-speaker control was the most-accused group), and mcp.py still quoted the pre-#5 "one document in five" rather than the measured 6.0%. The footer now uses a natural frequency so the no-percentage guard there stays intact.
Upstream:qh.export_ts_client emitted TypeScript that tsc rejects outright — doubled braces, a stray } closing the client class after any zero-parameter endpoint, and every optional parameter emitted as required. Its tests passed because they counted braces, and two of the defects cancelled. Fixed and released as i2mint/qh#11 (qh 0.0.19), which also stopped Optional[str] collapsing to any.
frontend/, Vite + TypeScript + ProseMirror. Stack decision written before building, one house invariant at a time, each departure with its cost named: misc/docs/frontend-stack-decision.md.
Policy: invalidate, never silently re-anchor. Implemented two-level, which the research did not spell out and which matters: a finding is invalidated when the edit touches its own characters or anywhere in its segment, because segment lean and strength come from density_aggregate and are per unit of text — adding a sentence changes what every finding in that paragraph is worth without touching any of them. The document verdict goes on any edit at all
ProseMirror Mapping + DecorationSet.map, side-car, never in the document schema. No CRDT
Re-score recomputes against the current text; offsets correct by construction at that instant
Two channels, matching render.py so the report and the UI cannot disagree about what orange means
Paste, evaluate, edit, re-evaluate. Upload not built — paste covers the demonstration and PDF extraction is a different problem wearing the same button
Persistence and reload — deliberately not built. The fuzzy re-anchoring and orphan list are needed only here; more to the point, quietly keeping someone's writing in their browser is a privacy decision a tool like this should not make for them
The finding that was not anticipated: the editor must not normalise the text, and that is correctness rather than taste. A rich-text editor turns ' into ’, trims trailing spaces and collapses lone newlines. mixed-apostrophes (0.35) and trailing-whitespace (0.25) are human-leaning, and mid-sentence-newline is decisive — so a tidying editor would systematically delete the evidence that exonerates people, and do it invisibly. Hence a whitespace-preserving one-block schema, which also makes a plain offset o exactly position o + 1.
The TypeScript is generated from the Python, so the one-source-of-truth claim survives a second language: export_types() from the dataclasses, export_client() from the OpenAPI, both committed, and tests/test_generated_sources.py fails on drift (verified by renaming a field). CI now installs [http] so that guard actually runs there.
One bug only a browser could find: the position-mapping biases were inverted, so a finding at offset 0 swallowed text inserted before it — a finding about five characters claimed 74 it had never seen. Reasoning reads fine either way round. Pinned by a vitest case verified to fail against the old biases.
Gap: the frontend's vitest and tsc do not run in CI — the wads reusable workflow is Python-only with no Node support. Local-only, and worth a follow-up.
Left open after Phase 5
Frontend CI — #8. 10 vitest cases and the typecheck are local-only; wads has no Node job.
Persistence, reload, fuzzy re-anchoring, the orphan list. The redundant selectors on every Span are already what this would be built from. The orphan list is the part with real design in it.
Upload (md/txt/pdf), and the optional compare-with-last-scored view.
The frontend is not in the wheel.pip install ductus[http] serves the API and no UI unless pointed at one with DUCTUS_UI_DIR. Committing a minified bundle into a Python package is a trade not made here.
Not planned
Watermark detection (SynthID, Kirchenbauer). Only works when you control or trust the generator's key — a different deployment model. An explicitly-labelled verification mode at most, never the default path.
Anything that helps evade detection. This package describes text; acquaint's deslop makes writing good, not laundered.
Cross-package
The tells catalogue moved here from acquaint
acquaint.deslop imports it instead of carrying its own copy (contract pinned by tests/test_tells_contract.py)
v1 ships the core, the CLI and the agent skills. This issue tracks the path to everything else, in the order that costs least and teaches most. The full document is
misc/docs/roadmap.md; this is the tracking summary.The claim each phase has to keep true: every addition is an ADD at a seam that already exists, and the core does not change. If a phase needs a core change, the core was coupled to its first surface and the core is what gets fixed.
The seams (decided in v1, not revisited)
segmenter="paragraph"(also"sentence","document", or a callable)imbed.fixed_step_chunkerdetectors=tells,forensic,rhetoric,rhythm— deterministic, zero extra deps[local]); Sapling, Pangram ([api])aggregate=Phase 1 — model-based detectors (
[local]) — done; gate not cleared, shipped opt-inLanded in #2. The deterministic layer has high precision and low recall: 7 signals across the 12 known-mixed fixture documents at sentence granularity (the "8" above was stale — corrected in
CLAUDE.md). Recall is the problem this phase existed to fix, and it is still the problem.fast_detect_gpt(text, span, *, model=...)— conditional probability curvature, one sampling passbinoculars(text, span, *, observer=..., performer=...)— cross-perplexity of a paired observer/performerimport ductusloads notorch, asserted in a subprocess testgpt2,distilgpt2/gpt2), withEleutherAI/gpt-neo-2.7Band the Falcon-7B pair documented as the upgradeDEFAULT_DETECTORSunchanged, and is now a separate list from theDETECTORSregistry so registering can never silently promote. Numbers:misc/docs/phase-1-results.md; re-run withpython misc/measure_detectors.pymisc/docs/curvature-as-evidence.mdCarried into Phase 2: the model detectors found 10 human-leaning segments the deterministic set misses entirely, all 10 correct, and every false machine flag was plain human prose (casual email, wire-service sport) — the Liang et al. failure mode behaving as predicted. Cheapest untried lever is the upgrade models; the fixture's
fill_gapsshape is close to a worst case for these methods.Phase 2 — calibration and a second fixture set — done (#3)
Results:
phase-2-results.md. Decision first, as in Phase 1:what-calibration-means-here.md.tests/fixtures/roft_boundary.json, MIT) — 18 documents, two at each of the nine boundary positions. It earned its keep immediately: the model-based detectors clear the Phase 1 gate on it outright (12/75 at 86% precision vs 0/75 for the deterministic set), confirming that LLMTrace'sfill_gapsshape was close to their worst caseaggregate=widened to taken_chars(the seam could not previously express a rate) anddensity_aggregateis the new defaultFit per-detector weights— deliberately not done. Seven deterministic signals across twelve documents is not an evidence base; fitting on it would produce the shape of a result with none of the contentlean = 2p − 1is invertible, so relabelling a fitted probability does not stop it being one. Replaced by calibrate the instrument, not the verdict, plus a sufficiency testTwo findings that were not expected:
triad,not-x-but-y,discourse-opener— formal argumentative register. This package is biased against good writing, not against non-native writing. The shipped agent skills said otherwise and were corrected.Phase 3 — MCP — done (#4)
ductus/mcp.py—mk_mcp()overpy2mcp.mk_mcp_from_refs, string refs, core never imports MCP.pip install 'ductus[mcp]'thenductus-mcp. The core did not change, as predictedTOOL_REFSis derived fromtools._dispatch_funcs, so there is nothing that could drift and nothing for a parity test to catch. A verb that changes the host declares it at its own definition (@host_mutating) — soinstall_skillsis unreachable over MCP without anyone hand-writing a second listmiddleware=andauth=pass straight through for a deployed connectorinstructionscarry the limits, since an MCP client reads that and nothing elseThe first time a surface told us anything. Under
from __future__ import annotations— every module here — the schema layer beneathfastmcpreads annotations as strings and drops every keyword-only default, sogauge(source=...)failed with five "missing required argument" errors. This package's convention is keyword-only from the 2nd/3rd argument, so every verb was uncallable. Narrowed to a minimal repro and filed upstream as i2mint/py2mcp#12; worked around inductus.mcp._resolve_annotationswith a test pinning it. The tool list and JSON schema both looked perfectly correct — it only appeared on an actual call, which is why this surface is tested by driving a real client rather than by inspecting a schema.Reducing false accusations — done (#5, thorwhalen/acquaint#30)
Taken before Phases 4 and 5: a UI on top of a one-in-five false-accusation rate ships the problem wider. Decision and all three gates fixed before the numbers:
document-verdict-decision.md. Results:reducing-false-accusations.md.misc/measure_rules.py): three rules were finding nothing —not-x-but-y(42 human docs, 0 TP),summary-closer(25, 0),contrastive-negation(19, 0). Two were broken, matching ordinary negation ("not eat from the tree but") and the "not only X but also Y" correlativeweight:intells.yamlchangesductusonly —acquaintreadstier, neverweight, verified in its source.summary-closerkeeps tier E, drops to weight 0.15score.roll_up). Pooling gave1−0.97ⁿ— a coin toss by 30 segments — treating accumulation as corroboration.leanunchanged;strengthjudged by rate and density, believed at its weakestacquaint, whose suite was run (2 failures, both encoding the old over-broad behaviour), fixed properly plus a regression test — 555 pass,ductus>=0.0.8pinnedTwo wrong turns, both caught by measuring rather than reasoning, both now pinned by tests: rate-only scoring made the default segmenter worse (12.0% → 13.7%), and computing
leanfrom segment labels silently discarded sub-threshold human-leaning evidence.Deliberately not done.
triad(96 human docs, 2 TP),discourse-opener(88, 4) andexclamation(54, 8) are the dominant remaining cost with terrible cost-to-benefit — but they survive the pre-registered gate, and inventing a second criterion after seeing which rules the first one missed is the move this project has avoided three times. Evidence recorded for whoever sets one in advance. Also open, and flagged rather than settled: whether tier E is right forsummary-closerinacquaint— that is a question about style advice, and the measurement is about authorship.Phase 4 — HTTP — done (#6)
Built once the frontend became the remote consumer this phase was waiting for.
qh.mk_appover the same callables.ROUTED_FUNCSis derived fromtools._dispatch_funcs, asTOOL_REFSis — three surfaces, one list, still no parity test to writeqh.export_ts_clientfor the frontend's typed client, plusexport_types()for the report shapeimport ductuspulls in neither FastAPI nor qh, asserted in a subprocess testqhdoes not have it — keyword-only defaults survive and every verb is callable with only its required argumentsWhat this surface found, and the other two could not. A verb list that is safe at a CLI is not automatically safe when the caller is a stranger.
gauge(source=...)reads a file when the string names one,gauge(out=...)writes one,judgments=is a third — correct when you typed the command yourself, an arbitrary file read and an arbitrary file write when you did not. Neither the CLI nor a local stdio MCP host can see it, because on those surfaces it is not a bug.Fixed the way
host_mutatingalready works, not by inventing a mechanism: the verb declares it at its own definition (@host_paths(source="read", out="write")) and the adapter refuses by reading that declaration, knowing nothing aboutgauge.readis refused on exactly the condition under which_read_sourcewould open a file, so the branch is unreachable rather than guessed at.mk_app(guard_host_paths=False)is the seam for a loopback service you run for yourself.Two stale honesty claims corrected, both of which had survived the phase that made them wrong:
render.py's footer still said detectors "over-flag non-native English" (this package's measured bias runs the other way — the native-speaker control was the most-accused group), andmcp.pystill quoted the pre-#5 "one document in five" rather than the measured 6.0%. The footer now uses a natural frequency so the no-percentage guard there stays intact.Upstream:
qh.export_ts_clientemitted TypeScript thattscrejects outright — doubled braces, a stray}closing the client class after any zero-parameter endpoint, and every optional parameter emitted as required. Its tests passed because they counted braces, and two of the defects cancelled. Fixed and released as i2mint/qh#11 (qh 0.0.19), which also stoppedOptional[str]collapsing toany.Phase 5 — frontend (score → edit → re-score) — done (#7)
frontend/, Vite + TypeScript + ProseMirror. Stack decision written before building, one house invariant at a time, each departure with its cost named:misc/docs/frontend-stack-decision.md.density_aggregateand are per unit of text — adding a sentence changes what every finding in that paragraph is worth without touching any of them. The document verdict goes on any edit at allMapping+DecorationSet.map, side-car, never in the document schema. No CRDTrender.pyso the report and the UI cannot disagree about what orange meansPersistence and reload— deliberately not built. The fuzzy re-anchoring and orphan list are needed only here; more to the point, quietly keeping someone's writing in their browser is a privacy decision a tool like this should not make for themThe finding that was not anticipated: the editor must not normalise the text, and that is correctness rather than taste. A rich-text editor turns
'into’, trims trailing spaces and collapses lone newlines.mixed-apostrophes(0.35) andtrailing-whitespace(0.25) are human-leaning, andmid-sentence-newlineis decisive — so a tidying editor would systematically delete the evidence that exonerates people, and do it invisibly. Hence a whitespace-preserving one-block schema, which also makes a plain offsetoexactly positiono + 1.The TypeScript is generated from the Python, so the one-source-of-truth claim survives a second language:
export_types()from the dataclasses,export_client()from the OpenAPI, both committed, andtests/test_generated_sources.pyfails on drift (verified by renaming a field). CI now installs[http]so that guard actually runs there.One bug only a browser could find: the position-mapping biases were inverted, so a finding at offset 0 swallowed text inserted before it — a finding about five characters claimed 74 it had never seen. Reasoning reads fine either way round. Pinned by a vitest case verified to fail against the old biases.
Gap: the frontend's
vitestandtscdo not run in CI — the wads reusable workflow is Python-only with no Node support. Local-only, and worth a follow-up.Left open after Phase 5
Spanare already what this would be built from. The orphan list is the part with real design in it.pip install ductus[http]serves the API and no UI unless pointed at one withDUCTUS_UI_DIR. Committing a minified bundle into a Python package is a trade not made here.Not planned
acquaint'sdeslopmakes writing good, not laundered.Cross-package
acquaintacquaint.deslopimports it instead of carrying its own copy (contract pinned bytests/test_tells_contract.py)