Everything tried on the way to compass-0.2.0, in order, with the rule each attempt was held to and what happened. Every number is in dev/results/.
- No JevBench item, label or paraphrase in any training, calibration or selection file; every file passes an 8-gram overlap check against the public JevBench files.
- Design choices (readout, fusion weight, hyperparameters) are made on an internal selection split. Calibration is fitted on a separate calibration split after the design is frozen. A release split is evaluated once per configuration. A shadow suite of 280 drafted items in domains disjoint from every other split is run once per configuration and decides promotion.
- JevBench's public items are run once per frozen configuration and reported as they come.
This synthetic test contrasts a fixed typed decision with generated reasoning. Each item contains three shuffled, unrelated reporting chains and asks whether one person is exactly a given number of manager links above another. The label is balanced yes/no and computed from the generated graph.
The suite has 20 questions at each chain length: 1, 2, 3, 4, 6, 8, and 12 links, 140 in total, generated by build_tasks in scripts/manager_chain_eval.py with seed 20260927. A yes item names the person at the top of the main chain. A no item names the person one link short of it, or at length 1 the top of a distractor chain. The suite is outside Compass training, calibration, selection, release, and JevBench evaluation data. docs/diagrams/manager-chain-example.svg shows one item.
| System | How it answers | Runner | Result file |
|---|---|---|---|
| Compass 0.2.0 | The local /v1/systemone endpoint runs the fusion readout with the LoRA v2 adapter; yes when calibrated p(yes) ≥ 0.5. No generated tokens. |
scripts/run_compass_manager_chain.py |
dev/results/manager-chain-compass-4090-2026-09-28.json |
| Jev 1.13.0 | The same yes/no question sent to TypeSafe's hosted /v1/systemone endpoint. |
scripts/run_jev_manager_chain.py |
dev/results/manager-chain-jev-4090host-2026-09-28.json |
| Qwen3.5-9B, thinking on | Writes a reasoning trace, then FINAL: YES or FINAL: NO. The system prompt is qwen_prompt in scripts/manager_chain_eval.py. |
scripts/run_openrouter_manager_chain.py |
dev/results/manager-chain-qwen9b-2026-09-27.json |
Compass and Jev ran sequentially from the same RTX 4090 host. Compass served its release endpoint locally after five warm-up requests. Jev remained a remote TypeSafe API call. The reported latency is client-observed end-to-end time for those two routes.
Qwen ran through OpenRouter as qwen/qwen3.5-9b, pinned to Parasail's bf16 endpoint with fallbacks off, reasoning enabled, temperature 0, seed 20260927, a 32,768-token limit, and 12 concurrent requests. The answer is read only after </think>. A trace with no final answer is scored wrong. Reasoning tokens are counted with the Qwen/Qwen3.5-9B tokenizer.
| Length | Compass | Jev | Qwen3.5-9B | Qwen median reasoning tokens | Qwen gave no final answer |
|---|---|---|---|---|---|
| 1 | 20/20 | 20/20 | 16/20 | 842 | 4 |
| 2 | 10/20 | 14/20 | 17/20 | 987 | 3 |
| 3 | 10/20 | 12/20 | 17/20 | 981 | 3 |
| 4 | 10/20 | 13/20 | 19/20 | 988 | 1 |
| 6 | 10/20 | 12/20 | 20/20 | 1,139 | 0 |
| 8 | 12/20 | 10/20 | 20/20 | 1,242 | 0 |
| 12 | 10/20 | 10/20 | 20/20 | 2,260 | 0 |
| All | 82/140 (59 %) | 91/140 (65 %) | 129/140 (92 %) | 11 |
Compass and Jev are at chance from length 2 onward. Qwen answered every item it finished correctly. The 11 items without an answer are greedy-decoding loops, with the same line such as "Okay." repeated up to 200 times, which Parasail ended with finish_reason: "error" after 18,972 to 25,270 tokens, ten of them at 301 to 316 seconds. They occur only at lengths 1 to 4. At temperature 0 a rerun repeats the same loop, so they stand as wrong. The 140 recorded answers cost $0.056 for 444,748 output tokens, and the final pass took 16 minutes. The chart is docs/diagrams/manager-chain-results-qwen9b.svg.
Compass is built on Qwen3.5-4B, so this compares it with a model more than twice its size. Hosted temperature-0 decoding is not bit-for-bit reproducible.
| System | Median endpoint time (p95) | Generated answer tokens across 140 questions | Recorded dollar charge |
|---|---|---|---|
| Compass 0.2.0, local RTX 4090 | 0.287 s (0.324 s) | 0 | Self-hosted GPU |
| Jev 1.13.0, TypeSafe API | 0.334 s (0.399 s) | 2,800 | $0.0028 via TypeSafe |
| Qwen3.5-9B, thinking on | 23.54 s (301.72 s) | 444,748 | $0.056 via OpenRouter |
Compass finished 1.2× faster than Jev on the two client-observed endpoint routes. Qwen took 82× longer than Compass per request, generated 159× more answer tokens than Jev, and cost 20× more than Jev's API run. Compass uses a self-hosted GPU, so its cost does not belong in that API-charge comparison. Qwen requests ran with 12-way concurrency; its latency is not the suite wall time. docs/diagrams/manager-chain-runtime-qwen9b.svg renders the comparison.
- The first 4B run, sampled at temperature 0.6 with a 1,024-token cap, is the
qwen_thinkingblock indev/results/manager-chain-2026-09-27.json. 80 of 140 traces hit the cap, so it measures the cap, not the model. It is superseded. - An uncapped greedy 4B run on an RTX PRO 4500 through
transformersgeneratedecoded at about 31 tokens/s, one question at a time. Buildingcausal-conv1dfor CUDA 13 to matchtorch 2.14.0+cu130removed the fallback warnings but left speed unchanged (30.8 tokens/s before, 31.2 after). At roughly 2,000 tokens per trace the suite needed about 2.5 hours before loops, and each loop to 32,768 tokens added about 18 minutes. It was stopped during a two-question smoke test, whose second depth-1 question was still generating after 15 minutes.dev/results/manager-chain-qwen-2026-09-27.jsonholds one row from an earlier attempt and is not a result. - OpenRouter does not serve Qwen3.5-4B. The only hosted endpoint is Featherless AI through Hugging Face's inference providers, whose concurrency limits made its run time unpredictable.
uv run python scripts/run_openrouter_manager_chain.py \
--out dev/results/manager-chain-qwen9b-2026-09-27.json
uv run python scripts/render_manager_chain_chart.py \
--compass-input dev/results/manager-chain-compass-4090-2026-09-28.json \
--jev-input dev/results/manager-chain-jev-4090host-2026-09-28.json \
--qwen-input dev/results/manager-chain-qwen9b-2026-09-27.json \
--out docs/diagrams/manager-chain-results-qwen9b.svg
uv run python scripts/render_manager_chain_runtime_chart.py \
--compass-input dev/results/manager-chain-compass-4090-2026-09-28.json \
--jev-input dev/results/manager-chain-jev-4090host-2026-09-28.json \
--qwen-input dev/results/manager-chain-qwen9b-2026-09-27.json \
--out docs/diagrams/manager-chain-runtime-qwen9b.svgThe OpenRouter runner reads OPENROUTER_API_KEY from the environment or .env, checkpoints after every answer, and continues an interrupted run with --resume. Parasail rate-limits above about 12 concurrent requests.
| Readout | Accuracy | ECE |
|---|---|---|
| verification (per-candidate yes/no log-odds) | 62.0 % | 0.103 |
| direct (lettered options, letter logits) | 58.3 % | 0.101 |
| fusion, equal weight in log space | 65.2 % | 0.098 |
| fusion + content-free debiasing (λ 0.5 / 1.0) | 60.1 / 55.8 % | 0.109 / 0.205 |
| direct with its own system prompt | 50.4 % | 0.204 |
| fusion over that | 56.2 % | 0.133 |
Fusion was adopted. Debiasing (subtracting each candidate's log-odds against an empty document) and a dedicated prompt for the direct readout were both worse and were dropped.
| Variant | Mechanism | Data | Outcome |
|---|---|---|---|
| head-v1 | residual MLP on the hidden state at the readout position, frozen backbone | 2,700 generated items | +7 points on the internal release split, nothing on JevBench's public items, hard-tier ECE 0.105 → 0.231. Not promoted. |
lora-v2 (compass-0.2.0) |
LoRA r16 through the verification and direct readouts; listwise, per-readout, ordinal, permutation and opaque-label losses | 2,150 items (1,900 generated with contrast and held-out language forms, 250 drafted) | Shadow core families 58.8 → 68.2 %, ECE 0.103 → 0.046; public standard 80.6 → 87.5 %, hard 56.8 → 60.4 %, hard ECE 0.107 → 0.069. Promoted. |
| lora-v2b | same recipe, seed 1, completed data (2,170 items) | as above plus tradeoff | Ties on accuracy (core 68.8 %), ECE 0.097. Not promoted. |
| lora-v3 | same recipe plus 400 generated adequacy items and adequacy loss weight 2.0 | 2,470 items | Core 62.4 %, ECE 0.099; drafted adequacy unchanged at 57.5 %. Not promoted. |
| Configuration | Easy | Standard | Hard | Hard ECE | Hard TVD |
|---|---|---|---|---|---|
| verification only | 100 % | 80.6 % | 55.9 % | 0.089 | 0.250 |
| compass-0.1.0 (fusion) | 100 % | 80.6 % | 56.8 % | 0.105 | 0.248 |
| compass-0.1.1 (fusion, hard-like calibration) | 100 % | 80.6 % | 56.8 % | 0.107 | 0.250 |
| head-v1 | 100 % | 81.9 % | 55.9 % | 0.231 | 0.264 |
| compass-0.2.0 (fusion + LoRA) | 100 % | 87.5 % | 60.4 % | 0.069 | 0.216 |
| Configuration | Overall | Core (policy, multi-hop, adequacy, ambiguous) | Adequacy |
|---|---|---|---|
| compass-0.1.1 | 66.4–66.8 %, ECE 0.043–0.059 | 58.8–59.4 %, ECE 0.067–0.103 | 65.0 % |
| compass-0.2.0 | 76.4–76.8 %, ECE 0.053 | 68.2–68.8 %, ECE 0.046–0.050 | 52.5–57.5 % |
| lora-v2b | 77.1 %, ECE 0.060 | 68.8 %, ECE 0.097 | 57.5 % |
| lora-v3 | 73.6 %, ECE 0.039 | 62.4 %, ECE 0.099 | 57.5 % |
Ranges are the same configuration on two GPUs (bf16 kernel noise).
- Generated data with computed labels trains the model on the generator. It transferred to JevBench only when the training set also had drafted scenarios, contrast in every wrong option, and language forms held out of training (v2); it did not transfer as a hidden-state head (v1) or as more generated adequacy items (v3).
- Calibration must be fitted on items at the difficulty where it is measured. JevBench scores it on the hard tier only.
- Temporal arithmetic in one forward pass is not solvable by any system on the board without generation (20–33 % for all of them); it is not a training target.
- The remaining gap is answer-adequacy judging, which is what the private judge tier consists of. Only drafted or hand-written items in that style can be expected to move it.