Skip to content

Roadmap: from v1 (core + CLI + skills) to all surfaces #1

Description

@thorwhalen

v1 ships the core, the CLI and the agent skills. This issue tracks the path to everything else, in the order that costs least and teaches most. The full document is misc/docs/roadmap.md; this is the tracking summary.

The claim each phase has to keep true: every addition is an ADD at a seam that already exists, and the core does not change. If a phase needs a core change, the core was coupled to its first surface and the core is what gets fixed.

The seams (decided in v1, not revisited)

Seam v1 default Replacement
segmenter= "paragraph" (also "sentence", "document", or a callable) imbed.fixed_step_chunker
detectors= tells, forensic, rhetoric, rhythm — deterministic, zero extra deps Fast-DetectGPT, Binoculars ([local]); Sapling, Pangram ([api])
aggregate= transparent weighted sum a scorer calibrated on the fixtures

Phase 1 — model-based detectors ([local]) — done; gate not cleared, shipped opt-in

Landed in #2. The deterministic layer has high precision and low recall: 7 signals across the 12 known-mixed fixture documents at sentence granularity (the "8" above was stale — corrected in CLAUDE.md). Recall is the problem this phase existed to fix, and it is still the problem.

  • fast_detect_gpt(text, span, *, model=...) — conditional probability curvature, one sampling pass
  • binoculars(text, span, *, observer=..., performer=...) — cross-perplexity of a paired observer/performer
  • Deferred imports; import ductus loads no torch, asserted in a subprocess test
  • Default proxy models CPU-sized (gpt2, distilgpt2/gpt2), with EleutherAI/gpt-neo-2.7B and the Falcon-7B pair documented as the upgrade
  • Gate: measured, NOT cleared. Machine recall 1/40 → 3/40, machine precision 1/1 → 3/8. DEFAULT_DETECTORS unchanged, and is now a separate list from the DETECTORS registry so registering can never silently promote. Numbers: misc/docs/phase-1-results.md; re-run with python misc/measure_detectors.py
  • Decision written down before implementation: a continuous score becomes a banded signal with a dead zone, scored relative to the document, never a linear weight — misc/docs/curvature-as-evidence.md

Carried into Phase 2: the model detectors found 10 human-leaning segments the deterministic set misses entirely, all 10 correct, and every false machine flag was plain human prose (casual email, wire-service sport) — the Liang et al. failure mode behaving as predicted. Cheapest untried lever is the upgrade models; the fixture's fill_gaps shape is close to a worst case for these methods.

Phase 2 — calibration and a second fixture set — done (#3)

Results: phase-2-results.md. Decision first, as in Phase 1: what-calibration-means-here.md.

  • RoFT slice vendored (tests/fixtures/roft_boundary.json, MIT) — 18 documents, two at each of the nine boundary positions. It earned its keep immediately: the model-based detectors clear the Phase 1 gate on it outright (12/75 at 86% precision vs 0/75 for the deterministic set), confirming that LLMTrace's fill_gaps shape was close to their worst case
  • FPR on human-written text measured — the headline. 350 texts from BEA-2019 W&I+LOCNESS: 34.3% of human documents falsely accused, now 20.6%. A flagged sentence is far better evidence than a flagged document (2–3% per sentence)
  • FPR by proficiency band, against a native control. W&I+LOCNESS carries CEFR levels, which is sharper than one aggregate. Corpus is non-redistributable, so it is measured, not vendored — script downloads it, numbers are committed
  • aggregate= widened to take n_chars (the seam could not previously express a rate) and density_aggregate is the new default
  • Fit per-detector weights — deliberately not done. Seven deterministic signals across twelve documents is not an evidence base; fitting on it would produce the shape of a result with none of the content
  • Still no percentage. And the reassurance in this issue was wrong: lean = 2p − 1 is invertible, so relabelling a fitted probability does not stop it being one. Replaced by calibrate the instrument, not the verdict, plus a sufficiency test

Two findings that were not expected:

  1. Most of the old false-positive rate was a length artefact. Pooled, it ran from 11% (under 750 chars) to 74% (over 2500) — a human document was ~7× more likely to be accused for being long.
  2. The bias runs the opposite way from the literature. Length-matched, FPR rises with proficiency (A 31%, B 40%, C 47%), and the native control was the most-accused group. The deterministic rules fire on triad, not-x-but-y, discourse-opener — formal argumentative register. This package is biased against good writing, not against non-native writing. The shipped agent skills said otherwise and were corrected.

Phase 3 — MCP — done (#4)

  • ductus/mcp.py — mk_mcp() over py2mcp.mk_mcp_from_refs, string refs, core never imports MCP. pip install 'ductus[mcp]' then ductus-mcp. The core did not change, as predicted
  • One list, and better than a parity test: TOOL_REFS is derived from tools._dispatch_funcs, so there is nothing that could drift and nothing for a parity test to catch. A verb that changes the host declares it at its own definition (@host_mutating) — so install_skills is unreachable over MCP without anyone hand-writing a second list
  • middleware= and auth= pass straight through for a deployed connector
  • The server's instructions carry the limits, since an MCP client reads that and nothing else

The first time a surface told us anything. Under from __future__ import annotations — every module here — the schema layer beneath fastmcp reads annotations as strings and drops every keyword-only default, so gauge(source=...) failed with five "missing required argument" errors. This package's convention is keyword-only from the 2nd/3rd argument, so every verb was uncallable. Narrowed to a minimal repro and filed upstream as i2mint/py2mcp#12; worked around in ductus.mcp._resolve_annotations with a test pinning it. The tool list and JSON schema both looked perfectly correct — it only appeared on an actual call, which is why this surface is tested by driving a real client rather than by inspecting a schema.

Reducing false accusations — done (#5, thorwhalen/acquaint#30)

Taken before Phases 4 and 5: a UI on top of a one-in-five false-accusation rate ships the problem wider. Decision and all three gates fixed before the numbers: document-verdict-decision.md. Results: reducing-false-accusations.md.

  • 20.6% → 6.0% human documents falsely accused (sentence); 12.0% → 6.3% on the default paragraph path; per-segment 2–3% → 0.4–2%
  • Nothing traded away — not one correctly-flagged machine segment lost on either fixture (LLMTrace 1/40 at 1/1, RoFT 0/75), and RoFT lost its one false flag
  • Per-rule table (misc/measure_rules.py): three rules were finding nothing — not-x-but-y (42 human docs, 0 TP), summary-closer (25, 0), contrastive-negation (19, 0). Two were broken, matching ordinary negation ("not eat from the tree but") and the "not only X but also Y" correlative
  • Tier and weight separated. A per-rule weight: in tells.yaml changes ductus only — acquaint reads tier, never weight, verified in its source. summary-closer keeps tier E, drops to weight 0.15
  • Document verdict is now a function of the segment verdicts (score.roll_up). Pooling gave 1−0.97ⁿ — a coin toss by 30 segments — treating accumulation as corroboration. lean unchanged; strength judged by rate and density, believed at its weakest
  • Cross-package landed, not left drifted: the pattern narrowing reaches acquaint, whose suite was run (2 failures, both encoding the old over-broad behaviour), fixed properly plus a regression test — 555 pass, ductus>=0.0.8 pinned

Two wrong turns, both caught by measuring rather than reasoning, both now pinned by tests: rate-only scoring made the default segmenter worse (12.0% → 13.7%), and computing lean from segment labels silently discarded sub-threshold human-leaning evidence.

Deliberately not done. triad (96 human docs, 2 TP), discourse-opener (88, 4) and exclamation (54, 8) are the dominant remaining cost with terrible cost-to-benefit — but they survive the pre-registered gate, and inventing a second criterion after seeing which rules the first one missed is the move this project has avoided three times. Evidence recorded for whoever sets one in advance. Also open, and flagged rather than settled: whether tier E is right for summary-closer in acquaint — that is a question about style advice, and the measurement is about authorship.

Phase 4 — HTTP — done (#6)

Built once the frontend became the remote consumer this phase was waiting for.

  • qh.mk_app over the same callables. ROUTED_FUNCS is derived from tools._dispatch_funcs, as TOOL_REFS is — three surfaces, one list, still no parity test to write
  • qh.export_ts_client for the frontend's typed client, plus export_types() for the report shape
  • The core did not change, as predicted. import ductus pulls in neither FastAPI nor qh, asserted in a subprocess test
  • Checked for the Phase 3 defect first, by driving a real client. qh does not have it — keyword-only defaults survive and every verb is callable with only its required arguments

What this surface found, and the other two could not. A verb list that is safe at a CLI is not automatically safe when the caller is a stranger. gauge(source=...) reads a file when the string names one, gauge(out=...) writes one, judgments= is a third — correct when you typed the command yourself, an arbitrary file read and an arbitrary file write when you did not. Neither the CLI nor a local stdio MCP host can see it, because on those surfaces it is not a bug.

Fixed the way host_mutating already works, not by inventing a mechanism: the verb declares it at its own definition (@host_paths(source="read", out="write")) and the adapter refuses by reading that declaration, knowing nothing about gauge. read is refused on exactly the condition under which _read_source would open a file, so the branch is unreachable rather than guessed at. mk_app(guard_host_paths=False) is the seam for a loopback service you run for yourself.

Two stale honesty claims corrected, both of which had survived the phase that made them wrong: render.py's footer still said detectors "over-flag non-native English" (this package's measured bias runs the other way — the native-speaker control was the most-accused group), and mcp.py still quoted the pre-#5 "one document in five" rather than the measured 6.0%. The footer now uses a natural frequency so the no-percentage guard there stays intact.

Upstream: qh.export_ts_client emitted TypeScript that tsc rejects outright — doubled braces, a stray } closing the client class after any zero-parameter endpoint, and every optional parameter emitted as required. Its tests passed because they counted braces, and two of the defects cancelled. Fixed and released as i2mint/qh#11 (qh 0.0.19), which also stopped Optional[str] collapsing to any.

Phase 5 — frontend (score → edit → re-score) — done (#7)

frontend/, Vite + TypeScript + ProseMirror. Stack decision written before building, one house invariant at a time, each departure with its cost named: misc/docs/frontend-stack-decision.md.

  • Policy: invalidate, never silently re-anchor. Implemented two-level, which the research did not spell out and which matters: a finding is invalidated when the edit touches its own characters or anywhere in its segment, because segment lean and strength come from density_aggregate and are per unit of text — adding a sentence changes what every finding in that paragraph is worth without touching any of them. The document verdict goes on any edit at all
  • ProseMirror Mapping + DecorationSet.map, side-car, never in the document schema. No CRDT
  • Re-score recomputes against the current text; offsets correct by construction at that instant
  • Two channels, matching render.py so the report and the UI cannot disagree about what orange means
  • Paste, evaluate, edit, re-evaluate. Upload not built — paste covers the demonstration and PDF extraction is a different problem wearing the same button
  • Persistence and reload — deliberately not built. The fuzzy re-anchoring and orphan list are needed only here; more to the point, quietly keeping someone's writing in their browser is a privacy decision a tool like this should not make for them

The finding that was not anticipated: the editor must not normalise the text, and that is correctness rather than taste. A rich-text editor turns ' into ’, trims trailing spaces and collapses lone newlines. mixed-apostrophes (0.35) and trailing-whitespace (0.25) are human-leaning, and mid-sentence-newline is decisive — so a tidying editor would systematically delete the evidence that exonerates people, and do it invisibly. Hence a whitespace-preserving one-block schema, which also makes a plain offset o exactly position o + 1.

The TypeScript is generated from the Python, so the one-source-of-truth claim survives a second language: export_types() from the dataclasses, export_client() from the OpenAPI, both committed, and tests/test_generated_sources.py fails on drift (verified by renaming a field). CI now installs [http] so that guard actually runs there.

One bug only a browser could find: the position-mapping biases were inverted, so a finding at offset 0 swallowed text inserted before it — a finding about five characters claimed 74 it had never seen. Reasoning reads fine either way round. Pinned by a vitest case verified to fail against the old biases.

Gap: the frontend's vitest and tsc do not run in CI — the wads reusable workflow is Python-only with no Node support. Local-only, and worth a follow-up.

Left open after Phase 5

  • Frontend CI — #8. 10 vitest cases and the typecheck are local-only; wads has no Node job.
  • Persistence, reload, fuzzy re-anchoring, the orphan list. The redundant selectors on every Span are already what this would be built from. The orphan list is the part with real design in it.
  • Upload (md/txt/pdf), and the optional compare-with-last-scored view.
  • The frontend is not in the wheel. pip install ductus[http] serves the API and no UI unless pointed at one with DUCTUS_UI_DIR. Committing a minified bundle into a Python package is a trade not made here.

Not planned

  • Watermark detection (SynthID, Kirchenbauer). Only works when you control or trust the generator's key — a different deployment model. An explicitly-labelled verification mode at most, never the default path.
  • Anything that helps evade detection. This package describes text; acquaint's deslop makes writing good, not laundered.

Cross-package

  • The tells catalogue moved here from acquaint
  • acquaint.deslop imports it instead of carrying its own copy (contract pinned by tests/test_tells_contract.py)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions