Skip to content

Add efficiency table, RQ5 cost, figures and generated README results (step 5) - #18

Merged
drewOrc merged 5 commits into
mainfrom
feat/step5-report
Sep 29, 2026
Merged

drewOrc merged 5 commits into
mainfrom
feat/step5-report

Conversation

@drewOrc

@drewOrc drewOrc commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Step 5 (docs/PLAN.md RQ5, AC5, AC8). RQ1 to RQ5 are complete; Tier 1 acceptance is not. AC1 keeps its definition (renamed AC1b: a full run from a clean clone, including training and the Haiku run) and is not met; an AC1a (verify the Release artifacts and rebuild the analysis and README offline) is added as a faster check but does not replace AC1b. AC1 is not implemented in this PR. Nothing is trained, no API is called, no setting is chosen again, and the benchmark scope is unchanged.

What it adds

  • make bench-cpu writes results/efficiency/cpu_latency.json: CPU batch-1 latency for both encoders.
    • The curve weights were deleted, so the timed model is rebuilt with the same architecture. It is built by the training loader (train.load_model_and_tokenizer): the pinned pretrained backbone plus a seeded 151-way head. Latency depends on shapes, not weight values.
    • The run stops unless the model revision, max_length, torch and transformers versions, and the parameter count all equal the trained k=100 run's.
    • It records the attention implementation (sdpa for both), the hardware and the thread counts.
    • Method: torch.inference_mode(), 4 intra-op threads, 50 warm-up queries, then validation rows 0 to 499 in order. Reports p50, p95 and mean.
  • make llm-latency writes results/efficiency/haiku_latency.json: Haiku per-call latency from the journal. This is client-side time, so it includes the network round trip; it is not a like-for-like comparison with the encoders.
  • make cost writes results/cost/cost.json, with measured and assumed numbers kept in separate blocks.
    • Measured: M4 training wall-clock, CPU latency, Haiku tokens and spend, and each router's real Haiku spend.
    • Assumed: accelerator at US$0.5, 1 or 2 per hour; labelling at US$0.05, 0.2 or 1 per example (sensitivity only); US$0.05 per vCPU-hour for local inference. The Mac's purchase price is not used.
    • Break-even is computed for every scenario and is labelled as a scenario sensitivity.
  • make figures: learning curves, risk-coverage, router comparison and threshold transfer, using the Okabe-Ito palette plus distinct line styles and markers. Reruns are byte-identical. In the router comparison, accuracy is drawn as points (the axis starts at 75%, stated in the caption) and the call rate as bars from 0.
  • make report builds results/report.md and the README block between the generated markers from committed JSON. The tests check it in two ways:
    • The whole block must equal what the JSON gives.
    • A table-driven test maps each of the 12 first-screen numbers to its exact field in summary.json and formats it independently of report.py, so a number wired to the wrong field fails even after a regenerate.

README first screen

  • Two columns, "What transferred" and "What did not transfer", followed by an engineering conclusion: calibrate thresholds on labelled data that represents deployment traffic.
  • The 87% figure is described as the part of the gap closed by reweighting test to the validation OOS share, not as a cause.
  • Hybrid Haiku calls are given per seed: 1,235, 1,157 and 1,548 of 5,500 at k=10; 170, 9 and 31 at k=100. The k=100 hybrid is presented as a small safety and diagnostic lever.
  • The README also has a Limitations section and a privacy note on the journal design.

Checks

  • make lint, make test (575 passed) and make smoke are green.
  • Running make report, make figures and make cost twice gives no diff.
  • Mutations turn tests red. From the first round: the cost formula, break-even, local inference cost, percentiles, the parameter guard, a hand-edited README number, PNG metadata and the partial-feasibility filter. From review: X3 (the k=10 hybrid accuracy taken from the small-only field), X4 (the validation risk taken from the reweighted field), the k=10 call counts taken from k=100, the router accuracy drawn as bars, the setup check disabled, and version comparison without dropping the +cpu label of the Linux torch wheel.

Step 5 (docs/PLAN.md RQ5, AC5, AC8). No training and no API calls.

- make bench-cpu: CPU batch-1 latency of both encoders with the pinned
  pretrained backbone and a 151-way head; stops unless the parameter count
  equals the trained k=100 run's. make llm-latency: Haiku per-call latency.
- make cost: measured costs and assumed prices kept apart; break-even per
  scenario.
- make figures: learning curves, risk-coverage, routers, threshold transfer.
- make report: results/report.md and the README block between the
  generated markers; tests/test_report.py fails when either is stale.
…accuracy

Review fixes for step 5: table-driven tests map every first-screen number
to its field in summary.json; k=10 hybrid call counts per seed; router
accuracy drawn as points instead of bars on a truncated axis; latency
model built by the training loader and checked against the trained run.
…hange note

AC1 keeps its definition as AC1b (a full run from a clean clone) and
must pass before Tier 1 acceptance; AC1a is added as a faster artifact
check and does not replace it. The latency re-measurement changed one
cell of the training-only break-even table and six cells of the
labelling sensitivity table (all under 0.02 percent), not one cell.
@drewOrc
drewOrc merged commit 49da171 into main Sep 29, 2026
5 checks passed
@drewOrc
drewOrc deleted the feat/step5-report branch September 29, 2026 06:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant