Skip to content

Latest commit

 

History

History
129 lines (95 loc) · 10.7 KB

File metadata and controls

129 lines (95 loc) · 10.7 KB

Experiments

Everything tried on the way to compass-0.2.0, in order, with the rule each attempt was held to and what happened. Every number is in dev/results/.

Rules that applied throughout

  • No JevBench item, label or paraphrase in any training, calibration or selection file; every file passes an 8-gram overlap check against the public JevBench files.
  • Design choices (readout, fusion weight, hyperparameters) are made on an internal selection split. Calibration is fitted on a separate calibration split after the design is frozen. A release split is evaluated once per configuration. A shadow suite of 280 drafted items in domains disjoint from every other split is run once per configuration and decides promotion.
  • JevBench's public items are run once per frozen configuration and reported as they come.

Manager-chain stress test

This synthetic test contrasts a fixed typed decision with generated reasoning. Each item contains three shuffled, unrelated reporting chains and asks whether one person is exactly a given number of manager links above another. The label is balanced yes/no and computed from the generated graph.

The suite has 20 questions at each chain length: 1, 2, 3, 4, 6, 8, and 12 links, 140 in total, generated by build_tasks in scripts/manager_chain_eval.py with seed 20260927. A yes item names the person at the top of the main chain. A no item names the person one link short of it, or at length 1 the top of a distractor chain. The suite is outside Compass training, calibration, selection, release, and JevBench evaluation data. docs/diagrams/manager-chain-example.svg shows one item.

Setup

System How it answers Runner Result file
Compass 0.2.0 The local /v1/systemone endpoint runs the fusion readout with the LoRA v2 adapter; yes when calibrated p(yes) ≥ 0.5. No generated tokens. scripts/run_compass_manager_chain.py dev/results/manager-chain-compass-4090-2026-09-28.json
Jev 1.13.0 The same yes/no question sent to TypeSafe's hosted /v1/systemone endpoint. scripts/run_jev_manager_chain.py dev/results/manager-chain-jev-4090host-2026-09-28.json
Qwen3.5-9B, thinking on Writes a reasoning trace, then FINAL: YES or FINAL: NO. The system prompt is qwen_prompt in scripts/manager_chain_eval.py. scripts/run_openrouter_manager_chain.py dev/results/manager-chain-qwen9b-2026-09-27.json

Compass and Jev ran sequentially from the same RTX 4090 host. Compass served its release endpoint locally after five warm-up requests. Jev remained a remote TypeSafe API call. The reported latency is client-observed end-to-end time for those two routes.

Qwen ran through OpenRouter as qwen/qwen3.5-9b, pinned to Parasail's bf16 endpoint with fallbacks off, reasoning enabled, temperature 0, seed 20260927, a 32,768-token limit, and 12 concurrent requests. The answer is read only after </think>. A trace with no final answer is scored wrong. Reasoning tokens are counted with the Qwen/Qwen3.5-9B tokenizer.

Result

Length Compass Jev Qwen3.5-9B Qwen median reasoning tokens Qwen gave no final answer
1 20/20 20/20 16/20 842 4
2 10/20 14/20 17/20 987 3
3 10/20 12/20 17/20 981 3
4 10/20 13/20 19/20 988 1
6 10/20 12/20 20/20 1,139 0
8 12/20 10/20 20/20 1,242 0
12 10/20 10/20 20/20 2,260 0
All 82/140 (59 %) 91/140 (65 %) 129/140 (92 %) 11

Compass and Jev are at chance from length 2 onward. Qwen answered every item it finished correctly. The 11 items without an answer are greedy-decoding loops, with the same line such as "Okay." repeated up to 200 times, which Parasail ended with finish_reason: "error" after 18,972 to 25,270 tokens, ten of them at 301 to 316 seconds. They occur only at lengths 1 to 4. At temperature 0 a rerun repeats the same loop, so they stand as wrong. The 140 recorded answers cost $0.056 for 444,748 output tokens, and the final pass took 16 minutes. The chart is docs/diagrams/manager-chain-results-qwen9b.svg.

Compass is built on Qwen3.5-4B, so this compares it with a model more than twice its size. Hosted temperature-0 decoding is not bit-for-bit reproducible.

Speed and token work

System Median endpoint time (p95) Generated answer tokens across 140 questions Recorded dollar charge
Compass 0.2.0, local RTX 4090 0.287 s (0.324 s) 0 Self-hosted GPU
Jev 1.13.0, TypeSafe API 0.334 s (0.399 s) 2,800 $0.0028 via TypeSafe
Qwen3.5-9B, thinking on 23.54 s (301.72 s) 444,748 $0.056 via OpenRouter

Compass finished 1.2× faster than Jev on the two client-observed endpoint routes. Qwen took 82× longer than Compass per request, generated 159× more answer tokens than Jev, and cost 20× more than Jev's API run. Compass uses a self-hosted GPU, so its cost does not belong in that API-charge comparison. Qwen requests ran with 12-way concurrency; its latency is not the suite wall time. docs/diagrams/manager-chain-runtime-qwen9b.svg renders the comparison.

Why not Qwen3.5-4B

  • The first 4B run, sampled at temperature 0.6 with a 1,024-token cap, is the qwen_thinking block in dev/results/manager-chain-2026-09-27.json. 80 of 140 traces hit the cap, so it measures the cap, not the model. It is superseded.
  • An uncapped greedy 4B run on an RTX PRO 4500 through transformers generate decoded at about 31 tokens/s, one question at a time. Building causal-conv1d for CUDA 13 to match torch 2.14.0+cu130 removed the fallback warnings but left speed unchanged (30.8 tokens/s before, 31.2 after). At roughly 2,000 tokens per trace the suite needed about 2.5 hours before loops, and each loop to 32,768 tokens added about 18 minutes. It was stopped during a two-question smoke test, whose second depth-1 question was still generating after 15 minutes. dev/results/manager-chain-qwen-2026-09-27.json holds one row from an earlier attempt and is not a result.
  • OpenRouter does not serve Qwen3.5-4B. The only hosted endpoint is Featherless AI through Hugging Face's inference providers, whose concurrency limits made its run time unpredictable.

Reproduce

uv run python scripts/run_openrouter_manager_chain.py \
  --out dev/results/manager-chain-qwen9b-2026-09-27.json
uv run python scripts/render_manager_chain_chart.py \
  --compass-input dev/results/manager-chain-compass-4090-2026-09-28.json \
  --jev-input dev/results/manager-chain-jev-4090host-2026-09-28.json \
  --qwen-input dev/results/manager-chain-qwen9b-2026-09-27.json \
  --out docs/diagrams/manager-chain-results-qwen9b.svg
uv run python scripts/render_manager_chain_runtime_chart.py \
  --compass-input dev/results/manager-chain-compass-4090-2026-09-28.json \
  --jev-input dev/results/manager-chain-jev-4090host-2026-09-28.json \
  --qwen-input dev/results/manager-chain-qwen9b-2026-09-27.json \
  --out docs/diagrams/manager-chain-runtime-qwen9b.svg

The OpenRouter runner reads OPENROUTER_API_KEY from the environment or .env, checkpoints after every answer, and continues an interrupted run with --resume. Parasail rate-limits above about 12 concurrent requests.

Readout comparison (frozen Qwen3.5-4B, selection split, 276 items)

Readout Accuracy ECE
verification (per-candidate yes/no log-odds) 62.0 % 0.103
direct (lettered options, letter logits) 58.3 % 0.101
fusion, equal weight in log space 65.2 % 0.098
fusion + content-free debiasing (λ 0.5 / 1.0) 60.1 / 55.8 % 0.109 / 0.205
direct with its own system prompt 50.4 % 0.204
fusion over that 56.2 % 0.133

Fusion was adopted. Debiasing (subtracting each candidate's log-odds against an empty document) and a dedicated prompt for the direct readout were both worse and were dropped.

Trained variants

Variant Mechanism Data Outcome
head-v1 residual MLP on the hidden state at the readout position, frozen backbone 2,700 generated items +7 points on the internal release split, nothing on JevBench's public items, hard-tier ECE 0.105 → 0.231. Not promoted.
lora-v2 (compass-0.2.0) LoRA r16 through the verification and direct readouts; listwise, per-readout, ordinal, permutation and opaque-label losses 2,150 items (1,900 generated with contrast and held-out language forms, 250 drafted) Shadow core families 58.8 → 68.2 %, ECE 0.103 → 0.046; public standard 80.6 → 87.5 %, hard 56.8 → 60.4 %, hard ECE 0.107 → 0.069. Promoted.
lora-v2b same recipe, seed 1, completed data (2,170 items) as above plus tradeoff Ties on accuracy (core 68.8 %), ECE 0.097. Not promoted.
lora-v3 same recipe plus 400 generated adequacy items and adequacy loss weight 2.0 2,470 items Core 62.4 %, ECE 0.099; drafted adequacy unchanged at 57.5 %. Not promoted.

Public items, all configurations run

Configuration Easy Standard Hard Hard ECE Hard TVD
verification only 100 % 80.6 % 55.9 % 0.089 0.250
compass-0.1.0 (fusion) 100 % 80.6 % 56.8 % 0.105 0.248
compass-0.1.1 (fusion, hard-like calibration) 100 % 80.6 % 56.8 % 0.107 0.250
head-v1 100 % 81.9 % 55.9 % 0.231 0.264
compass-0.2.0 (fusion + LoRA) 100 % 87.5 % 60.4 % 0.069 0.216

Shadow suite, 280 drafted items

Configuration Overall Core (policy, multi-hop, adequacy, ambiguous) Adequacy
compass-0.1.1 66.4–66.8 %, ECE 0.043–0.059 58.8–59.4 %, ECE 0.067–0.103 65.0 %
compass-0.2.0 76.4–76.8 %, ECE 0.053 68.2–68.8 %, ECE 0.046–0.050 52.5–57.5 %
lora-v2b 77.1 %, ECE 0.060 68.8 %, ECE 0.097 57.5 %
lora-v3 73.6 %, ECE 0.039 62.4 %, ECE 0.099 57.5 %

Ranges are the same configuration on two GPUs (bf16 kernel noise).

What we learned

  • Generated data with computed labels trains the model on the generator. It transferred to JevBench only when the training set also had drafted scenarios, contrast in every wrong option, and language forms held out of training (v2); it did not transfer as a hidden-state head (v1) or as more generated adequacy items (v3).
  • Calibration must be fitted on items at the difficulty where it is measured. JevBench scores it on the hard tier only.
  • Temporal arithmetic in one forward pass is not solvable by any system on the board without generation (20–33 % for all of them); it is not a training target.
  • The remaining gap is answer-adequacy judging, which is what the private judge tier consists of. Only drafted or hand-written items in that style can be expected to move it.