Skip to content

Latest commit

 

History

History
25 lines (22 loc) · 2.25 KB

File metadata and controls

25 lines (22 loc) · 2.25 KB

Status

What is built, and how far each part has been exercised.

Component Evidence
Schema, prompt layouts, router, label codes done, tested
Calibration fitting, ECE, profile I/O done, tested
Contamination guard (all 13 eval subsets) done, tested
Training config + budget arithmetic done, verified
Benchmark harness done, tested (offline) + run against live Jev
Data pipeline — 29 sources, 3 splits, augmented mixture done, loads and splits verified
Held-out eval export done, round-trips into levbench
Collator, both readouts, objective done, tested
Training loop + checkpointing done — 30 steps on Qwen3.5-0.8B-Base, both modes, losses finite
Modal app, image, volumes image builds on Modal
Modal download / build_data / smoke run green on Modal
Modal train / calibrate / evaluate complete 4B run, calibrated and scored (ADR-018)
Modal serve runs; scored on S1Bench over HTTP
S1Bench harness done; all 13 subsets on S1Bench's pinned items, Jev within 0.8 pp of its published numbers on each
Decision engine (batched forward, optional fork, readout) runs; batched forward (ADR-023), skipped label codes and fitted temperatures (ADR-028)
Mode B head trains and serves; weak unseen-taxonomy transfer measured on massive-en-US (FINDINGS §12)
lev.load, release packaging, Hub publishing done; the release loads through both lev.load and the server on an H100 (check_release)

Three checkpoints have been trained, calibrated and scored on S1Bench against Jev through identical task files: 0.489, then 0.697, then 0.725 macro against Jev's 0.754 on the earlier six-subset definitions. The released checkpoint scores 0.689 against Jev's 0.761 on all 13 subsets as S1Bench pins them. What each run changed and why is in FINDINGS §12–17; the reader bug that corrupted the labels of one intermediate run is ADR-024.