Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,10 @@ jobs:
enable-cache: true

- name: Install dependencies (fails if uv.lock is out of date)
run: uv sync --locked
# The llm group (anthropic SDK) is installed so the Haiku runner's
# retry and key-redaction tests use the SDK's real exception classes.
# They run against a fake client; no key is set and no API is called.
run: uv sync --locked --group llm

- name: make lint (ruff check, ruff format --check, em dash check)
run: make lint
Expand Down
6 changes: 6 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,12 @@ checkpoints/
# they are attached to a GitHub Release and checked with `make verify-logits`.
results/summary.md
results/logits/
# Haiku replies: the journal and the predictions file (about 3 MB) go to the
# GitHub Release like the logits; results/llm-manifest.json and
# results/llm/haiku-8way.json are committed. The smoke run is scratch.
results/llm/*.jsonl
results/llm/*.lock
results/llm-smoke/
*.npz
*.tmp

Expand Down
61 changes: 61 additions & 0 deletions DEVLOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,67 @@

---

## 2026-09-29(晚):PR #14 審查修正(4 medium、6 low)

### 本次工作 / 執行摘要
- 單一執行鎖:`results/llm/<name>.lock` 用 `flock(LOCK_EX|LOCK_NB)`,run 與 finalize 都要拿鎖;拿不到就 exit 2。審查實測兩個行程並行時各自花到上限,合計到 144%。
- parser:`oos` 以單字邊界也算候選,候選數 ≥ 2 就判 oos 並標 `parse_failed`。`"oos (not travel_agent)"` 原本會被判成 travel_agent,是 RQ2 最危險的靜默誤派。
- journal 尾行:先修尾再載入。完整但缺換行的尾行保留並補換行(那筆已付費),只有無法解析的半行才截掉。原本是先載入再截尾,會把記憶體裡算成功、磁碟上已刪掉的那筆再付一次費。
- `count_tokens` 走同一套重試與遮罩(529 會重試、錯誤字串遮 key)。
- `settle` 把非預期例外也記成 failed 並停跑;Ctrl-C 時先等在途呼叫回來並記帳再往外拋;回應的 input 或 output 超過上界就停跑(`bound violated`,exit 1)。
- summary 新增 `parser_sha256`(重新解析用的是當下的 parser,而 parser 不在 identity 裡)與 `cap_scope`。
- 補五個守門的測試:非 ASCII 的上界(bytes 不是字元數)、journal 混入別的身分、gold 被改、只改 manifest 的 SHA、同一列成功兩次。

### 核心發現 / 數據
- **上限的範圍是「每個身分 × 每個 target」**:smoke 與改 prompt 之前的花費不算在內。AC6 的總花費要手動把 `results/llm/haiku-8way.json` 與 smoke 的 summary 加總。
- 上限看不到的部分:client 端逾時但伺服器已計費的請求,重試時沿用同一份預留;每次這種逾時最多多出一個單筆上界,並行時一波約 workers 個上界。
- (無實跑數據)

### Blockers / 遇到的問題
- (無)

### Next
- [ ] 下一個分析 PR 同時報新舊兩種解析規則(舊:子字串比對取第一個命中),以及兩者判定不一致的列數

### Files / Budget
- `src/tinyrouter/llm.py`、`src/tinyrouter/llm_run.py`、`tests/test_llm.py`、`tests/test_llm_run.py`、`tests/llm_fakes.py`、`.gitignore`
- API 花費:US$0

---

## 2026-09-29:步驟 4 之一,Haiku 8 類執行器(尚未實跑)

### 本次工作 / 執行摘要
- 新增 `src/tinyrouter/llm_run.py` 與 `make llm`、`make llm-smoke`、`make verify-llm`:validation 3,100 + test 5,500 逐筆呼叫 Haiku,逐筆寫入 journal(split、index、query 的 SHA-256、gold intent 與 agent、原始回覆、解析後標籤、`parse_failed`、tokens、花費、延遲、嘗試次數、request id)。
- query 原文不存,只存 SHA-256:資料集公開且鎖定 revision,雜湊足以證明紀錄對應哪一列,續跑時也拿它比對資料有沒有變。
- 續跑:journal 檔名含身分雜湊(模型、system prompt 的 SHA-256、temperature、max_tokens、user 內容格式)。身分一變就換新檔,舊回覆不會被拿來用;同身分重跑只補沒有成功紀錄的列。最後一行寫到一半(當機)會被截掉重做。
- 成本上限:開跑前用 token counting 取 prompt 的基礎 token 數,印出上界估計;每筆開打前預留「基礎 + 每個 byte 算一個 token 的輸入、max_tokens 的輸出」的上界,累計花費(含之前幾次)加上在途預留會超過 `--max-usd` 就不開新呼叫。實際花費依回傳的 usage 計算。
- 重試改由程式自己做(SDK 的 `max_retries=0`):429、5xx、408、409、連線錯誤指數退避並尊重 retry-after,5 次用盡記為失敗、整體 exit 1;400、401、403、404 不重試,而且停止開新呼叫。
- 完成判定比照 `completeness.py`:預測檔從磁碟讀回,(split, index) 恰為預期集合且各一次、身分一致,SHA-256 在檔案、summary、`results/llm-manifest.json` 三處相同,才印 `completed 8600/8600 llm predictions`。預期列數寫成字面值。
- CI 的 test job 改裝 `--group llm`,讓重試與 key 遮蔽的測試用 SDK 真正的例外類別。仍不設 key、不打 API。
- `classify` 原本把任何例外都當可重試,改為依錯誤類型判斷。

### 核心發現 / 數據
- system prompt 與 cost-aware-hybrid-router `src/routers/llm_router.py` 逐位元組相同(SHA-256 `560d22c5...5df574`,測試釘住)。模型、temperature 0、max_tokens 20、query 原樣當唯一 user 訊息,都與舊專案相同。
- 解析規則與舊專案**不同**:舊版對回覆做子字串比對、取集合迭代到的第一個命中(順序不固定);這裡只接受完整標籤,或恰好命中一個 in-scope agent,其餘判為 oos 並標 `parse_failed`。原始回覆都有存,要用舊規則重算不必再打 API。
- 定價:Haiku 4.5 每百萬 tokens 輸入 US$1、輸出 US$5(claude-api skill 的模型表,快取日期 2026-06-24)。粗估全量約 US$2.5,上界約 US$3.5,低於 AC6 的 US$5。
- (無實跑數據)

### Blockers / 遇到的問題
- (無)

### Next
- [ ] Drew:`.env` 放 key 後 `make llm-smoke`,看 20 筆的實際 token 與推估全量花費
- [ ] `make llm`,把 `results/llm/haiku-8way.jsonl` 附到 Release
- [ ] 下一個 PR:不確定性、risk-coverage、fallback、oracle(RQ3、RQ4、AC6)

### Files / Budget
- 新增:`src/tinyrouter/llm_run.py`、`tests/test_llm_run.py`、`tests/test_llm_deps.py`、`tests/llm_fakes.py`
- 修改:`src/tinyrouter/llm.py`、`tests/test_llm.py`、`tests/test_makefile.py`、`Makefile`、`.github/workflows/ci.yml`、`.gitignore`、`pyproject.toml`(只改註解)、`README.md`、`docs/OPERATIONS.md`
- API 花費:US$0

---

## 2026-09-23(夜):PR #7 審查修正

### 本次工作 / 執行摘要
Expand Down
22 changes: 21 additions & 1 deletion Makefile
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
.PHONY: setup lint format test test-network smoke train evaluate ac2 pilot-lr pilot-steps baselines \
curve oos-ablation verify-logits report clean-checkpoints
curve oos-ablation verify-logits llm-smoke llm verify-llm report clean-checkpoints

CONFIG ?= configs/bert-base.yaml
SEED ?= 42
Expand Down Expand Up @@ -112,6 +112,26 @@ oos-ablation:
verify-logits:
uv run python -m tinyrouter.archive

# Claude Haiku over CLINC150 in the 8-way routing space (docs/PLAN.md section 4,
# AC6). Needs ANTHROPIC_API_KEY (in .env or exported); without it they exit 2.
# Every call is journaled; a rerun calls only rows with no stored reply, and a
# change of model, prompt, temperature or max_tokens starts a fresh journal.
# A call is not started if it could take this run's identity past MAX_USD
# (default 5). Exit 1 when stopped by the cap or when a call failed after
# retries. `make llm-smoke` calls validation rows 0-19, writes under
# results/llm-smoke/ and prints the extrapolated cost of all 8,600 rows.
# Done means the whole last line `completed 8600/8600 llm predictions`.
llm-smoke:
uv run --group llm $(UV_ENV) python -m tinyrouter.llm_run --smoke $(if $(MAX_USD),--max-usd $(MAX_USD),)

llm:
uv run --group llm $(UV_ENV) python -m tinyrouter.llm_run $(if $(MAX_USD),--max-usd $(MAX_USD),)

# results/llm/haiku-8way.jsonl has every row once and matches its SHA-256 in
# results/llm-manifest.json and results/llm/haiku-8way.json.
verify-llm:
uv run python -m tinyrouter.llm_run --verify

report:
uv run python -m tinyrouter.report

Expand Down
6 changes: 5 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,9 @@ Other targets:
| `make baselines` | majority-class and TF-IDF centroid baselines on every (k, seed) sample, archived like an encoder run |
| `make curve MODEL=bert\|modernbert` | baselines, then 6 values of k x 3 seeds with the lr and S_min from `configs/curve.yaml`; writes `results/curves/<model>.json`; keeps no weights |
| `make oos-ablation` | ModernBERT, k=100, no OOS training rows, 3 seeds; writes `results/curves/oos-ablation.json` |
| `make llm-smoke` | Claude Haiku on validation rows 0-19 (needs `ANTHROPIC_API_KEY`); writes `results/llm-smoke/`, prints tokens, cost and the extrapolated cost of all 8,600 rows |
| `make llm` | Claude Haiku on validation 3,100 + test 5,500 in the 8-way space, one stored record per query; resumes; stops before passing `MAX_USD` (default 5); writes `results/llm/haiku-8way.{jsonl,json}` and `results/llm-manifest.json` |
| `make verify-llm` | check `results/llm/haiku-8way.jsonl`: every row once, SHA-256 equal in the file, the summary and the manifest |
| `make verify-logits` | check every archive in `results/logits/` against `results/logits-manifest.json` |
| `make report` | build `results/summary.md` from `results/runs/*.json` |
| `make clean-checkpoints` | delete all trained weights |
Expand Down Expand Up @@ -83,7 +86,8 @@ src/tinyrouter/
curves.py learning curves and the OOS ablation (reuses AC2 at k=100 when equivalent)
baselines.py majority-class and TF-IDF centroid baselines per curve point
ac2.py AC2 run over three seeds and PASS/FAIL verdict
llm.py Claude Haiku zero-shot router (baseline and fallback; not run yet)
llm.py Claude Haiku zero-shot router: prompt, retries, pricing (baseline and fallback)
llm_run.py Haiku over validation + test: journal, resume, cost cap, completion check
report.py results/*.json -> results/summary.md
smoke.py end-to-end wiring check
tests/ pytest; `network` marker for Hub downloads
Expand Down
2 changes: 1 addition & 1 deletion docs/OPERATIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ TinyRouter is a research repository with no deployment target: nothing runs as a

| job | runs | required to merge |
|---|---|---|
| `test` | tracked-files guard, `uv sync --locked`, `make lint`, `make test` (offline) | yes |
| `test` | tracked-files guard, `uv sync --locked --group llm` (the SDK's exception classes for the Haiku runner tests; no key, no API call), `make lint`, `make test` (offline) | yes |
| `commit-hygiene` | rejects tool-attribution trailers in commit messages (patterns in `.github/disallowed-trailers.txt`) | yes |
| `network` | `make test-network` and `make smoke` against the Hugging Face Hub | no |
| `pr-text-hygiene` (`pr-text.yml`) | rejects the same patterns in the PR title and body, on open, edit, push and reopen | yes (added to the required checks once it is on `main`) |
Expand Down
6 changes: 4 additions & 2 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -28,8 +28,10 @@ packages = ["src/tinyrouter"]

[dependency-groups]
# The Claude Haiku baseline and cascade fallback. Not installed by
# `uv sync`; `uv sync --group llm` installs it. Nothing in the unit tests
# imports it (tinyrouter/llm.py defers the import to call time).
# `uv sync`; `uv sync --group llm` installs it, and `make llm` / `make
# llm-smoke` ask uv for it. tinyrouter/llm.py imports it only at call time.
# CI installs it: the retry and redaction tests raise the SDK's real
# exception classes at a fake client (they skip where it is missing).
llm = [
"anthropic==1.8.0",
]
Expand Down
Loading
Loading