You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Add two report-only measures of module shape to the repository tooling, so "name modules for the domain question they answer" (AGENTS.md) becomes a number we can watch instead of a taste judgement:
Placement legibility — can an outside reader file a module into its owning family from its imports, its test path, and its first-touch commit subject alone? Measured by asking the typesafe-ai/jev evaluation model through AI Gateway, one choice question per file, scored against where the file actually landed.
Change coupling — do the physical families match how the repo actually changes? Measured from git history alone (logical coupling / co-change), with per-family modularity.
Both were prototyped outside the repo and produced the findings in Evidence below. This issue is to rebuild that prototype as first-class repo tooling with the guards the prototype had to discover by hand.
Neither measure gates anything. They are advisory reports, like pnpm depgraph, and they must not fail CI.
Evidence (measured, current tree)
Corpus: 1,726 production files, 39 families (23 packages/* + 16 src/ zones including (root), vendor, platform-runtime).
Baselines and the model, same 1,726 files, 39-way choice (random = 2.6%):
accuracy
majority family
18.1%
neighbour-vote over import paths
28.3%
k-NN over import-path sets
42.8%
Jev, reader-view
49.7% (1,357,479 tok, $0.057, 47s, 44 batches of 40)
Test–retest agreement across two independent calls: 96.4% over 302 files (accuracy 53.6% vs 53.0%).
Ablations, 302-file sample — these are why leakage control is the hard part of this issue:
evidence condition
accuracy
bare file names, no directory names anywhere
37.4%
family-relative paths, family ids withheld
36.1%
real import paths, test basename, redacted subject (scored condition)
53.6%
+ full test path (mirror leak)
67.5%
+ verbatim commit subject (naming leak)
71.2%
The prototype's "scored condition" row above is what this issue now calls the reader condition, and its two ablation rows are the name-withheld condition; the naming is fixed in Required behaviour so a rebuild cannot score the wrong one.
Own-family leakage before redaction: commit subject 51.7% (893/1726), import evidence 21.8%, local test path 7.7%. One redaction pass scrubbed 1,684 field-level hits.
Coupling, from 1,044 usable commits of 1,979 (34 mass-migration commits excluded): 6,337 edges at support ≥ 3, of which 4,084 (64.4%) cross families. Coupling kept inside its own family: 36.4% observed vs 11.4% expected for that size distribution → modularity 0.250; last 120 days 39.3% / 0.278. Of commits spanning ≥ 2 families the median spans 4, and 40% span ≥ 5 (42% over the last 120 days).
The two axes are independent, which is the point: host-kit is legible (84.1%) with in-family coupling weight 0.2 and out/in ratio 177 — a toolbox, not a module — while cli has modularity 0.0126 (third strongest community) and legibility 13.8%, and maestro is the only family whose coupling is mostly internal (out/in 0.6) at legibility 57.1%.
Required behaviour
Layer A — deterministic history model and coupling (no network, no new dependency)
Create three module groups:
scripts/repo-history/ — the single rename-resolved history model both reports consume.
renames.ts: build an oldPath → newPath map from every R record in git log -M --name-status, and a resolve(path) that applies it transitively (cycle-guarded, memoised) so a historical path reaches today's path.
records.ts: per current production file, the first commit whose rename-resolved record adds or renames it: { id, sha, date, subject, status: 'A' | 'R' }.
commits.ts: per commit, the set of rename-resolved current production file ids, plus its date and subject.
scripts/coupling/ — logical coupling from that model.
Affinity. For a commit touching k ≥ 2 production files, add 1 / (k − 1) to each unordered pair; support = number of commits containing the pair. Skip commits with k > 60 and report how many were skipped and which. Keep edges with support >= 3.
Family modularity. Let W = total kept edge weight, in_f = weight of edges with both ends in family f, out_f = weight of edges with exactly one end in f, s_f = 2·in_f + out_f. Then intraShare = Σ_f in_f / W, expected = Σ_f (s_f / (2W))², and per family Q_f = in_f / W − (s_f / (2W))². Σ_f Q_f must equal intraShare − expected; assert that identity in a test.
Family span histogram. Per commit, the count of distinct families touched; report the histogram, the median over commits spanning ≥ 2 families, and the share of those commits spanning ≥ 5.
Compute everything twice: over all history and over the trailing window (--since-days, default 120), so drift is visible.
Layer B — legibility oracle (network, opt-in)
Families and ground truth come from the gate's own model. Import listSourceFiles from scripts/layering/check.ts and derive the zone partition through scripts/layering/model.ts (see scripts/depgraph/load.ts for the reuse pattern). Do not re-derive zones, file sets, or package prefixes anywhere else.
Evidence per file, in this order: value-imports resolved to real repo paths (the file's own path never shown), external package specifiers verbatim, the mirrored test file's basename, and the first-touch commit subject with the conventional-commit type(scope): prefix and trailing (#NNNN) stripped.
Value imports only: static import/export … from with at least one non-type binding. Exclude import type, type-only specifiers, and dynamic import().
Test mirror resolution, in priority order: same directory, sibling __tests__/, then <family root>/{test,__tests__}/ by shortest path; if the match lives in another family, mark it foreign and do not show that family's id.
Two conditions; only one is the score. The reader condition shows resolved import paths, the mirrored test's basename, and the subject with the type(scope): prefix and (#NNNN) stripped, hiding only the file's own path; family ids stay scrubbed out of the subject in both conditions, since a subject narrates a change rather than a layout. That is what someone opening the tree sees, and it is the only condition the three baselines — which read the physical layout of every other file — are comparable to. A name-withheld condition additionally replaces every family id in the answer space with a fixed «x» placeholder in the paths, subject, and test path, matched as a word sequence so hostKit, host/kit, and Host Kit go with host-kit. It runs behind --ablation, costs a second full pass, and is reported as the gap to the reader number: that gap is what names carry. It is never a score and never sits next to the baselines. Redaction stays mandatory where it applies: the withheld pass reports its own scrubber defect rate and the run refuses to continue above 1% of files. The reader condition prints the same audit as a name echo rate — files whose evidence already names their own family — which is a property of a legible tree, so it is reported and not capped.
Baselines are required output. Majority, neighbour-vote (modal family among import targets), and k-NN (Jaccard over import-path sets, self excluded) must be computed and printed next to the model result. A legibility number published without the k-NN column is not usable.
Score output: overall accuracy vs all three baselines; a per-family row with files, sampled n, model accuracy, k-NN accuracy, delta, median provider confidence, Q_f, and out/in flow; and a divergence list of files where the model and k-NN independently agree on a different family, annotated with whether the file arrived there via a rename.
Two named upper-bound conditions may be kept behind flags (--with-test-dir, --raw-subject) and must always be printed as leak references, never as the score.
CLI and output shape
Mirror the depgraph conventions exactly:
pnpm coupling # -> .tmp/coupling/report.json + text summary
pnpm coupling --since-days 30 --out /tmp/coupling.json
pnpm legibility # sample run, 300 files, seeded
pnpm legibility --all --out .tmp/legibility/report.json
pnpm legibility affected packages/capture-kit/src # future query; do not build now
Add pnpm legibility, pnpm legibility:test, pnpm coupling, pnpm coupling:test, pnpm repo-history:test as node --experimental-strip-types entries. Wire the three *:test scripts as model gates following the depgraph precedent: ids in the CheckId union and the ordered list in scripts/check-affected/model.ts, gate(...) entries in scripts/check-affected/checks.ts, run-gate steps in .github/workflows/ci.yml next to the existing gate: depgraph step, and the test scripts in check:tooling. The report commands themselves stay ungated.
Guards
Never leak the answer into the evidence. The file's own path is never shown and family ids stay scrubbed out of the subject under both conditions — a subject narrates a change, not a layout. Under the name-withheld condition every family id in the answer space is scrubbed out of the paths, subject, and test path, and that defect rate is capped at 1%; under the reader condition the same audit is the name echo, printed and not capped. The conditions that reintroduce the mirror-test path or the raw subject stack on the reader line and are labelled leak references at the point of output. Regression-proof this: plant a file whose subject contains its own family id and show the audit names it under both readings.
No second source of truth for zones, file sets, or edges. Everything physical comes from scripts/layering/. A locally reimplemented zone map or import resolver is a rejection, per the AGENTS.md rule about not reconstructing another source of truth.
Report-only. No threshold fails the build; no ratchet; no allowlist. These numbers move with history. Nothing here is added to check:layering, check:fallow, or any check:* gate beyond the model tests.
No network in unit tests. The evaluation call sits behind one injected seam; unit tests use a fake that returns canned answers. pnpm legibility:test, coupling:test, and repo-history:test must pass offline, in CI, with no key.
Model-call quirks are handled at the client, not papered over. With 39 options and provider rounding to two decimals, a tie at the top fails whole-call validation (did not select a highest-probability option) and would discard every answer in that batch — observed 4 times in a 44-batch full pass. Required behaviour: split the batch in half and retry, recursively; a file that still cannot be answered is recorded as unanswered with a typed reason, never silently dropped, and the run prints the split count and unanswered count.
Secrets and spend. Read the key from the environment only; never log or echo it. Fail with an actionable typed message when it is missing. Print token and cost totals. Default to a seeded 300-file sample; a full-tree pass requires an explicit --all. Cap requests per run and stop with a message at the cap.
Deterministic sampling. Seeded, stratified by family, at least one file per family, same seed ⇒ same file set, recorded in the report.
Never commit stamped output. Reports go to .tmp/; docs/agents/testing.md forbids run-stamped output under scripts/. Committed fixtures for tests are fine.
Nothing under src/, packages/, apple/, or android/ changes. This issue is tooling only. The refactors in Out of scope are not licence to touch production code.
Repository mechanics.pnpm only, never npm install artifacts; pnpm format repository-wide; OXC lint clean; keep every new file well under 1,000 lines; tests mirror source one-to-one (a module and its test are created together, no additions to an aggregate test file); no new barrels. Stage a new module before trusting a layering-scan result, since it reads tracked files only.
ai dependency. The pinned ai@7.0.68 predates experimental_evaluate; the declared range ^7.0.68 already admits a version that has it, so update the lockfile resolution only — do not widen the range, do not add a new dependency, and do not move ai out of devDependencies.
Data shape for the evaluation call
experimental_evaluate from ai, not a chat-completions endpoint — typesafe-ai/jev is an evaluation model and chat completions rejects it with ModelTypeMismatchError.
experimental_evaluate({model: 'typesafe-ai/jev',state: { task, families, evidenceLegend },// shared framing, sent once per callquestions: {f0: {type: 'choice',instructions: `<fixed task text>\nEvidence: ${line}`, criteria }},});
criteria maps all 39 family ids to null. Family descriptions must not be supplied: the metric is whether the name carries the question, so a gloss would measure the gloss instead.
One call carries ~40 questions. Measured: ~31K tokens per 40 files, ~1s per batch, 1.36M tokens and $0.057 for the full tree. Output tokens are priced at zero; input is $0.042/M.
Auth: AI_GATEWAY_API_KEY. Provider confidence, when present, is at providerMetadata.typesafe.confidence[questionId] — it is a separate statistic from the selected option's probability and must be reported as such.
Answers are rounded to two decimals; keep the raw choice, the selected probability, and the top-3 distribution.
Acceptance criteria
Layer A:
pnpm coupling writes .tmp/coupling/report.json and prints: commits used/skipped with the skip threshold, edge count and cross-family share, intraShare vs expected and modularity for all-time and the trailing window, the per-family table, the top coupling hubs with their out-family count, the heaviest family pairs, and the family-span histogram including the ≥ 5 share.
pnpm coupling:test and pnpm repo-history:test pass offline, and hard-assert on a committed synthetic fixture mini-repo (not on the live tree): affinity normalisation, support cut, the Σ Q_f = intraShare − expected identity, and transitive rename resolution across a two-hop rename.
A rename applied twice in history resolves to the current path, and a historical path that no longer exists is dropped rather than counted.
The mass-migration skip is proven: fabricate a commit touching more than 60 production files in the fixture and show it is excluded and counted.
Live-tree values land in the ranges recorded in Evidence (modularity 0.20–0.35, ≥ 60% of kept edges crossing families, ≥ 25% of multi-family commits spanning ≥ 5). A miss is a printed drift warning naming the metric, not a failure.
Layer B:
pnpm legibility --all reproduces, within ±3 points on the current tree, under the reader condition: majority ≈ 18%, neighbour-vote ≈ 28%, k-NN ≈ 43%, model ≥ 48%, and a per-family spread of at least 60 points between the best and worst family with n ≥ 20.
pnpm legibility --all --ablation prints the name-withheld accuracy and its delta to the reader number, counts both passes in the token and cost totals, and never places the withheld number beside the baselines.
The name-withheld scrubber defect rate is below 1% and is printed every run, next to the reader name-echo rate.
A per-family row is emitted for every family with n ≥ 1, and the report never averages a family with n < 3 into any headline number.
The split-and-retry path is exercised by a unit test with a fake that returns a tie-failing response, and the test proves: the batch halves, answers survive, and a permanently unanswerable file appears as unanswered with a typed reason.
Missing AI_GATEWAY_API_KEY exits non-zero with a message naming the variable and the model id, and performs no requests.
The three *:test scripts are registered as gates and owned by a workflow; pnpm check:gate-manifest, pnpm check:layering, pnpm lint, pnpm typecheck, pnpm format:check and pnpm check:tooling all pass.
A README.md in each module group states which question the report answers, what the numbers do not mean, and the measured reference values — written in the scripts/depgraph/README.md register, including an explicit "this is not a removability or correctness claim" caveat.
docs/agents/testing.md gains one pointer sentence for each report, next to the existing "Before editing a shared module" section. No other prose docs change.
Dependencies and sequencing
Two PR-sized layers; Layer A first and independently useful. Layer A adds no dependency and needs no key. Layer B depends on scripts/repo-history/ from Layer A for first-touch subjects and rename status.
AI_GATEWAY_API_KEY with access to typesafe-ai/jev is required only to run Layer B's report, never its tests.
A fresh worktree needs pnpm install --frozen-lockfile && pnpm build before scripts/layering imports resolve to this checkout.
pnpm legibility affected <path> is listed above as a future query only; do not build it here.
Out of scope (recorded findings for separate issues)
The prototype surfaced these. They are evidence, not work authorised by this issue:
The command surface is one module split across five families. Coupling hubs: cli-schema/command-schema.ts (96.7 w, reaches 22 families), client/client-types.ts (89.4 w, 26 families; 297 lines, 18 exports, all types), cli-schema/cli-help.ts (85.8 w, 26 families; 1,255 lines), cli.ts (72.2 w), agent-device-client.ts (62.7 w), cli/parser/args.ts (46.9 w, 14 families), daemon/handlers/session.ts (46.9 w, 26 families), command-registry/registry.ts (42.4 w, 24 families; 1,997 lines). Driven by ordinary feature commits, and ≥ 5-family commit share is flat at 42% recently.
src/daemon/replay/internal/session-replay-*.ts is the most common placement divergence (4 files, model and k-NN agreeing on ad-replay, which already exists) — an extraction left half-done. A reciprocal capture-kit/src/recording/output-path.ts ↔ host-kit/src/session-paths.ts swap sits next to it.
Naming/routing payoffs where cohesion is real and legibility is not: cli (Q 0.0126 / 13.8%), maestro (Q 0.0121, out/in 0.6 / 57.1%), screenshot-diff (28.6%).
contracts scores 12.1% legibility and this issue should not treat that as a defect: a vocabulary family's imports point outward at everything, so its legible surface is its entry names and its importers. Fixing the measure for rank-1 families needs in-edges as evidence; that is a follow-up, and any such change must keep the existing condition reproducible for comparison.
Purpose
Add two report-only measures of module shape to the repository tooling, so "name modules for the domain question they answer" (AGENTS.md) becomes a number we can watch instead of a taste judgement:
typesafe-ai/jevevaluation model through AI Gateway, one choice question per file, scored against where the file actually landed.Both were prototyped outside the repo and produced the findings in Evidence below. This issue is to rebuild that prototype as first-class repo tooling with the guards the prototype had to discover by hand.
Neither measure gates anything. They are advisory reports, like
pnpm depgraph, and they must not fail CI.Evidence (measured, current tree)
Corpus: 1,726 production files, 39 families (23
packages/*+ 16src/zones including(root),vendor,platform-runtime).Baselines and the model, same 1,726 files, 39-way choice (random = 2.6%):
Test–retest agreement across two independent calls: 96.4% over 302 files (accuracy 53.6% vs 53.0%).
Ablations, 302-file sample — these are why leakage control is the hard part of this issue:
The prototype's "scored condition" row above is what this issue now calls the reader condition, and its two ablation rows are the name-withheld condition; the naming is fixed in Required behaviour so a rebuild cannot score the wrong one.
Own-family leakage before redaction: commit subject 51.7% (893/1726), import evidence 21.8%, local test path 7.7%. One redaction pass scrubbed 1,684 field-level hits.
Per-family spread under the scored condition, best → worst:
selectors96.8,managed-allocation96.6,session-journal100 (n=7),host-kit84.1,platform-android78.1,provider-webdriver76.7,platform-apple74.3,capture-kit63.0,commands57.6,maestro57.1,daemon-server40.7,cli-schema37.5,mcp33.3,screenshot-diff28.6,kernel27.3,core14.3,cli13.8,contracts12.1,(root)1.3,sdk0.Coupling, from 1,044 usable commits of 1,979 (34 mass-migration commits excluded): 6,337 edges at support ≥ 3, of which 4,084 (64.4%) cross families. Coupling kept inside its own family: 36.4% observed vs 11.4% expected for that size distribution → modularity 0.250; last 120 days 39.3% / 0.278. Of commits spanning ≥ 2 families the median spans 4, and 40% span ≥ 5 (42% over the last 120 days).
The two axes are independent, which is the point:
host-kitis legible (84.1%) with in-family coupling weight 0.2 and out/in ratio 177 — a toolbox, not a module — whileclihas modularity 0.0126 (third strongest community) and legibility 13.8%, andmaestrois the only family whose coupling is mostly internal (out/in 0.6) at legibility 57.1%.Required behaviour
Layer A — deterministic history model and coupling (no network, no new dependency)
Create three module groups:
scripts/repo-history/— the single rename-resolved history model both reports consume.renames.ts: build anoldPath → newPathmap from everyRrecord ingit log -M --name-status, and aresolve(path)that applies it transitively (cycle-guarded, memoised) so a historical path reaches today's path.records.ts: per current production file, the first commit whose rename-resolved record adds or renames it:{ id, sha, date, subject, status: 'A' | 'R' }.commits.ts: per commit, the set of rename-resolved current production file ids, plus its date and subject.scripts/coupling/— logical coupling from that model.scripts/legibility/— corpus, redaction, baselines, evaluation client, scoring.Formulas, exactly:
k ≥ 2production files, add1 / (k − 1)to each unordered pair;support= number of commits containing the pair. Skip commits withk > 60and report how many were skipped and which. Keep edges withsupport >= 3.W= total kept edge weight,in_f= weight of edges with both ends in familyf,out_f= weight of edges with exactly one end inf,s_f = 2·in_f + out_f. ThenintraShare = Σ_f in_f / W,expected = Σ_f (s_f / (2W))², and per familyQ_f = in_f / W − (s_f / (2W))².Σ_f Q_fmust equalintraShare − expected; assert that identity in a test.--since-days, default 120), so drift is visible.Layer B — legibility oracle (network, opt-in)
listSourceFilesfromscripts/layering/check.tsand derive the zone partition throughscripts/layering/model.ts(seescripts/depgraph/load.tsfor the reuse pattern). Do not re-derive zones, file sets, or package prefixes anywhere else.type(scope):prefix and trailing(#NNNN)stripped.import/export … fromwith at least one non-type binding. Excludeimport type, type-only specifiers, and dynamicimport().__tests__/, then<family root>/{test,__tests__}/by shortest path; if the match lives in another family, mark it foreign and do not show that family's id.type(scope):prefix and(#NNNN)stripped, hiding only the file's own path; family ids stay scrubbed out of the subject in both conditions, since a subject narrates a change rather than a layout. That is what someone opening the tree sees, and it is the only condition the three baselines — which read the physical layout of every other file — are comparable to. A name-withheld condition additionally replaces every family id in the answer space with a fixed«x»placeholder in the paths, subject, and test path, matched as a word sequence sohostKit,host/kit, andHost Kitgo withhost-kit. It runs behind--ablation, costs a second full pass, and is reported as the gap to the reader number: that gap is what names carry. It is never a score and never sits next to the baselines. Redaction stays mandatory where it applies: the withheld pass reports its own scrubber defect rate and the run refuses to continue above 1% of files. The reader condition prints the same audit as a name echo rate — files whose evidence already names their own family — which is a property of a legible tree, so it is reported and not capped.files, sampledn, model accuracy, k-NN accuracy, delta, median provider confidence,Q_f, and out/in flow; and a divergence list of files where the model and k-NN independently agree on a different family, annotated with whether the file arrived there via a rename.--with-test-dir,--raw-subject) and must always be printed as leak references, never as the score.CLI and output shape
Mirror the
depgraphconventions exactly:Report JSON:
{ generated: { commit, date, files, families }, baselines, perFamily: [{ family, files, n, accuracy, knn, delta, medianConfidence, modularity, outInFlow }], divergences: [{ id, family, predicted, p, viaRename, subject }] }. Coupling JSON:{ generated, window: { since, allTime }, assortativity, perFamily: [{ family, files, inWeight, outWeight, outInFlow, partners, Q }], hubs, familyPairs, edges }.Add
pnpm legibility,pnpm legibility:test,pnpm coupling,pnpm coupling:test,pnpm repo-history:testasnode --experimental-strip-typesentries. Wire the three*:testscripts as model gates following thedepgraphprecedent: ids in theCheckIdunion and the ordered list inscripts/check-affected/model.ts,gate(...)entries inscripts/check-affected/checks.ts,run-gatesteps in.github/workflows/ci.ymlnext to the existinggate: depgraphstep, and the test scripts incheck:tooling. The report commands themselves stay ungated.Guards
scripts/layering/. A locally reimplemented zone map or import resolver is a rejection, per the AGENTS.md rule about not reconstructing another source of truth.check:layering,check:fallow, or anycheck:*gate beyond the model tests.pnpm legibility:test,coupling:test, andrepo-history:testmust pass offline, in CI, with no key.did not select a highest-probability option) and would discard every answer in that batch — observed 4 times in a 44-batch full pass. Required behaviour: split the batch in half and retry, recursively; a file that still cannot be answered is recorded as unanswered with a typed reason, never silently dropped, and the run prints the split count and unanswered count.--all. Cap requests per run and stop with a message at the cap..tmp/;docs/agents/testing.mdforbids run-stamped output underscripts/. Committed fixtures for tests are fine.src/,packages/,apple/, orandroid/changes. This issue is tooling only. The refactors in Out of scope are not licence to touch production code.pnpmonly, nevernpm installartifacts;pnpm formatrepository-wide; OXC lint clean; keep every new file well under 1,000 lines; tests mirror source one-to-one (a module and its test are created together, no additions to an aggregate test file); no new barrels. Stage a new module before trusting a layering-scan result, since it reads tracked files only.aidependency. The pinnedai@7.0.68predatesexperimental_evaluate; the declared range^7.0.68already admits a version that has it, so update the lockfile resolution only — do not widen the range, do not add a new dependency, and do not moveaiout ofdevDependencies.Data shape for the evaluation call
experimental_evaluatefromai, not a chat-completions endpoint —typesafe-ai/jevis an evaluation model and chat completions rejects it withModelTypeMismatchError.criteriamaps all 39 family ids tonull. Family descriptions must not be supplied: the metric is whether the name carries the question, so a gloss would measure the gloss instead.AI_GATEWAY_API_KEY. Provider confidence, when present, is atproviderMetadata.typesafe.confidence[questionId]— it is a separate statistic from the selected option's probability and must be reported as such.choice, the selected probability, and the top-3 distribution.Acceptance criteria
Layer A:
pnpm couplingwrites.tmp/coupling/report.jsonand prints: commits used/skipped with the skip threshold, edge count and cross-family share,intraSharevsexpectedand modularity for all-time and the trailing window, the per-family table, the top coupling hubs with their out-family count, the heaviest family pairs, and the family-span histogram including the ≥ 5 share.pnpm coupling:testandpnpm repo-history:testpass offline, and hard-assert on a committed synthetic fixture mini-repo (not on the live tree): affinity normalisation, support cut, theΣ Q_f = intraShare − expectedidentity, and transitive rename resolution across a two-hop rename.Layer B:
pnpm legibility --allreproduces, within ±3 points on the current tree, under the reader condition: majority ≈ 18%, neighbour-vote ≈ 28%, k-NN ≈ 43%, model ≥ 48%, and a per-family spread of at least 60 points between the best and worst family withn ≥ 20.pnpm legibility --all --ablationprints the name-withheld accuracy and its delta to the reader number, counts both passes in the token and cost totals, and never places the withheld number beside the baselines.n ≥ 1, and the report never averages a family withn < 3into any headline number.unansweredwith a typed reason.AI_GATEWAY_API_KEYexits non-zero with a message naming the variable and the model id, and performs no requests.*:testscripts are registered as gates and owned by a workflow;pnpm check:gate-manifest,pnpm check:layering,pnpm lint,pnpm typecheck,pnpm format:checkandpnpm check:toolingall pass.README.mdin each module group states which question the report answers, what the numbers do not mean, and the measured reference values — written in thescripts/depgraph/README.mdregister, including an explicit "this is not a removability or correctness claim" caveat.docs/agents/testing.mdgains one pointer sentence for each report, next to the existing "Before editing a shared module" section. No other prose docs change.Dependencies and sequencing
scripts/repo-history/from Layer A for first-touch subjects and rename status.AI_GATEWAY_API_KEYwith access totypesafe-ai/jevis required only to run Layer B's report, never its tests.pnpm install --frozen-lockfile && pnpm buildbeforescripts/layeringimports resolve to this checkout.pnpm legibility affected <path>is listed above as a future query only; do not build it here.Out of scope (recorded findings for separate issues)
The prototype surfaced these. They are evidence, not work authorised by this issue:
cli-schema/command-schema.ts(96.7 w, reaches 22 families),client/client-types.ts(89.4 w, 26 families; 297 lines, 18 exports, all types),cli-schema/cli-help.ts(85.8 w, 26 families; 1,255 lines),cli.ts(72.2 w),agent-device-client.ts(62.7 w),cli/parser/args.ts(46.9 w, 14 families),daemon/handlers/session.ts(46.9 w, 26 families),command-registry/registry.ts(42.4 w, 24 families; 1,997 lines). Driven by ordinary feature commits, and≥ 5-family commit share is flat at 42% recently.cli-schema: 8 files, out/in flow 30, touches 29 of 39 families,Q ≈ 0, legibility 37.5% — the clearest wrongly-a-family candidate.src/daemon/replay/internal/session-replay-*.tsis the most common placement divergence (4 files, model and k-NN agreeing onad-replay, which already exists) — an extraction left half-done. A reciprocalcapture-kit/src/recording/output-path.ts↔host-kit/src/session-paths.tsswap sits next to it.cli(Q 0.0126 / 13.8%),maestro(Q 0.0121, out/in 0.6 / 57.1%),screenshot-diff(28.6%).contractsscores 12.1% legibility and this issue should not treat that as a defect: a vocabulary family's imports point outward at everything, so its legible surface is its entry names and its importers. Fixing the measure for rank-1 families needs in-edges as evidence; that is a follow-up, and any such change must keep the existing condition reproducible for comparison.