Add efficiency table, RQ5 cost, figures and generated README results (step 5) - #18
Merged
Merged
Conversation
Step 5 (docs/PLAN.md RQ5, AC5, AC8). No training and no API calls. - make bench-cpu: CPU batch-1 latency of both encoders with the pinned pretrained backbone and a 151-way head; stops unless the parameter count equals the trained k=100 run's. make llm-latency: Haiku per-call latency. - make cost: measured costs and assumed prices kept apart; break-even per scenario. - make figures: learning curves, risk-coverage, routers, threshold transfer. - make report: results/report.md and the README block between the generated markers; tests/test_report.py fails when either is stale.
…accuracy Review fixes for step 5: table-driven tests map every first-screen number to its field in summary.json; k=10 hybrid call counts per seed; router accuracy drawn as points instead of bars on a truncated axis; latency model built by the training loader and checked against the trained run.
…hange note AC1 keeps its definition as AC1b (a full run from a clean clone) and must pass before Tier 1 acceptance; AC1a is added as a faster artifact check and does not replace it. The latency re-measurement changed one cell of the training-only break-even table and six cells of the labelling sensitivity table (all under 0.02 percent), not one cell.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Step 5 (docs/PLAN.md RQ5, AC5, AC8). RQ1 to RQ5 are complete; Tier 1 acceptance is not. AC1 keeps its definition (renamed AC1b: a full run from a clean clone, including training and the Haiku run) and is not met; an AC1a (verify the Release artifacts and rebuild the analysis and README offline) is added as a faster check but does not replace AC1b. AC1 is not implemented in this PR. Nothing is trained, no API is called, no setting is chosen again, and the benchmark scope is unchanged.
What it adds
make bench-cpuwritesresults/efficiency/cpu_latency.json: CPU batch-1 latency for both encoders.train.load_model_and_tokenizer): the pinned pretrained backbone plus a seeded 151-way head. Latency depends on shapes, not weight values.torch.inference_mode(), 4 intra-op threads, 50 warm-up queries, then validation rows 0 to 499 in order. Reports p50, p95 and mean.make llm-latencywritesresults/efficiency/haiku_latency.json: Haiku per-call latency from the journal. This is client-side time, so it includes the network round trip; it is not a like-for-like comparison with the encoders.make costwritesresults/cost/cost.json, with measured and assumed numbers kept in separate blocks.make figures: learning curves, risk-coverage, router comparison and threshold transfer, using the Okabe-Ito palette plus distinct line styles and markers. Reruns are byte-identical. In the router comparison, accuracy is drawn as points (the axis starts at 75%, stated in the caption) and the call rate as bars from 0.make reportbuildsresults/report.mdand the README block between the generated markers from committed JSON. The tests check it in two ways:summary.jsonand formats it independently ofreport.py, so a number wired to the wrong field fails even after a regenerate.README first screen
Checks
make lint,make test(575 passed) andmake smokeare green.make report,make figuresandmake costtwice gives no diff.+cpulabel of the Linux torch wheel.