Add Haiku 8-way runner with journal, resume and cost cap - #14
Merged
Merged
Conversation
make llm calls claude-haiku-4-5-20251001 on validation 3,100 + test 5,500 with the prompt of cost-aware-hybrid-router (byte-identical, SHA-256 pinned in tests) and stores one record per query: split, index, query SHA-256, gold label, raw reply, parsed agent, parse_failed, tokens, cost, latency, attempts and request id. - Journal per identity (model, prompt, temperature, max_tokens): a rerun calls only rows without a stored reply; a changed identity reuses nothing. - Cost cap (default US$5): each call reserves an upper bound before it starts, cumulative across reruns; actual cost from usage. - Retries in code, not the SDK: 429, 5xx, 408, 409 and connection errors back off; others fail at once. Failures are journaled and exit 1. - Completion: the predictions file holds every row once and its SHA-256 matches the summary and results/llm-manifest.json before "completed 8600/8600 llm predictions" is printed. - make llm-smoke: validation rows 0-19 into results/llm-smoke/, with the extrapolated full cost. - The API key is redacted from logs, journal and exceptions. - CI installs the llm group so tests use the SDK's exception classes; a fake client stands in for the API.
- One runner per target: flock on results/llm/<name>.lock, exit 2 when another process holds it (two runs could each spend up to the cap). - parse_agent counts "oos" as a word among the candidates; two or more candidates are ambiguous, so "oos (not travel_agent)" is oos with parse_failed instead of travel_agent. - Repair the journal tail before loading it: a whole last record missing its newline is kept (it was paid for), only a half line is cut. - count_tokens goes through the same retries and key redaction. - settle journals unexpected exceptions as failures and stops; calls in flight at Ctrl-C are waited for and journaled; a response above its cost bound stops the run with exit 1. - Summary records the parser's hash and the cap's scope (per identity and target; the smoke run is not included). - Tests for guards that had none: UTF-8 byte bound, another identity in the journal, changed gold label, manifest-only SHA change, a row recorded twice.
…lation on the last row
drewOrc
added a commit
that referenced
this pull request
Sep 29, 2026
…res (#15) The first `make llm-smoke` after #14 failed on every call before any request was sent: `TypeError: Messages.create() got an unexpected keyword argument 'temperature'`. Nothing reached the API, US$0 was spent, and no key appeared in the output. ## Cause anthropic 1.x removed `temperature`, `top_p` and `top_k` from `messages.create()`. The API did not remove them, and Haiku 4.5 still accepts them. The runner's tests use a fake client that accepts any keyword, so none of them could see the mismatch with the real SDK. ## Fix - `request_params` sends `extra_body={"temperature": 0.0}`. The SDK merges `extra_body` into the request JSON as is, which is the SDK upgrade guide's advice when the model still accepts the setting and the code depends on it. Temperature 0 is part of the comparison with cost-aware-hybrid-router, so it is moved, not dropped. - The identity still records temperature 0, so its hash is unchanged. The failed rows in the smoke journal are retried on the next run. - `count_tokens` only ever sent `model`, `system` and `messages`, so it was not affected. Its arguments now come from `count_tokens_params()`, the one place they are built, as `request_params()` is for `messages.create`. ## New tests (for the 2026-09-29 smoke incident) - `test_the_arguments_we_send_fit_the_installed_sdk_signatures` binds the kwargs from `request_params()` and `count_tokens_params()` to the installed SDK's `Messages.create` and `Messages.count_tokens` signatures. It needs no network and no key. - `test_the_fake_client_receives_exactly_the_built_arguments` checks that what the client receives is exactly those kwargs. The signature check therefore covers what is actually sent. - `test_request_matches_the_old_project_settings` now asserts `extra_body["temperature"] == 0.0` and that no top-level `temperature` is sent. Mutation check (backup and restore): 6 mutations, all caught: - temperature at the top level (also the original bug) - temperature dropped - temperature added to `count_tokens` - `classify` or `count_tokens` sending extra kwargs of their own `make lint`, `make test` (423 passed) and `make smoke` pass locally.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Step 4, first of two PRs: the Claude Haiku runner for the 8-way routing comparison (docs/PLAN.md section 4 "LLM 對照", RQ4, AC6). Nothing here has called the API; the analysis (uncertainty, risk-coverage, fallback, oracle) is the next PR.
What it does
make llmsends every validation (3,100) and test (5,500) query toclaude-haiku-4-5-20251001, temperature 0, max_tokens 20, with the system prompt of cost-aware-hybrid-router. The prompt is byte-identical tosrc/routers/llm_router.pythere (SHA-256560d22c5...5df574, pinned intests/test_llm.py).One record per query:
split,index,query_sha256,gold_intent,gold_agent,raw_text,agent,parse_failed, input/output/cache tokens,cost_usd,latency_ms,attempts,request_id,stop_reason,identity_sha256,created_at. The query text is not stored: CLINC150 is public and pinned, and the hash ties each record to its row (a rerun refuses a journal whose hash disagrees with the dataset).results/llm/haiku-8way.<id12>.journal.jsonl). A rerun calls only rows with no stored reply. The journal is keyed by an identity hash of model, prompt, temperature and max_tokens, so changing any of them reuses nothing. A half-written last line is trimmed and redone.MAX_USD, default 5): prints an upper-bound estimate first (prompt tokens from the token counting endpoint, one token per query byte, output at max_tokens). Each call reserves that bound before it starts; no call starts if spend so far for this identity (earlier runs included) plus in-flight reservations could pass the cap. Actual cost comes fromusage, at US$1 / US$5 per MTok (Haiku 4.5).max_retries=0. 429, 5xx, 408, 409 and connection errors back off 1, 2, 4, 8 s (orretry-after, capped at 60 s), up to 5 attempts. Other errors fail at once and stop new calls. A failed row is journaled and the run exits 1.results/llm/haiku-8way.jsonlis read back and must hold every (split, index) once with one identity, and its SHA-256 must matchresults/llm/haiku-8way.jsonandresults/llm-manifest.json. Only then does it printcompleted 8600/8600 llm predictions.make verify-llmreruns the check.make llm-smokecalls validation rows 0-19 intoresults/llm-smoke/, separate from the real results, and prints tokens, cost and the extrapolated cost for 8,600 rows.ANTHROPIC_API_KEYit exits 2 before loading data. The key is redacted from printed output, the journal and exception text, and exceptions are not chained to the SDK error.Changes outside the runner
classifyused to retry every exception. Now it retries only the errors listed above.testjob runsuv sync --locked --group llm, so the tests raise the SDK's real exception classes at a fake client. No key is set and nothing calls the API.tests/test_llm_deps.pyfails in CI if the group is missing, so the tests that skip without the SDK cannot all skip silently there.parse_failed. Raw replies are stored, so the old rule can be applied without new calls.Review fixes (second round)
flockonresults/llm/<name>.lock; a second process exits 2 instead of spending from the same cap.parse_agentcountsoosas a word among the candidates; two or more candidates are ambiguous, so"oos (not travel_agent)"is oos withparse_failed.count_tokensuses the same retries and key redaction as the calls.Tests
421 passed locally (73 more than main), all with a fake client.
make smokepasses.Mutation check: 43 mutations of
llm.pyandllm_run.py, all 43 caught, including the five the review found surviving (character bound, journal row of another identity, changed gold label, manifest-only SHA change, row recorded twice).