docs(vision): campaign 2026-08-24 — nvfp4 think-on on the v0.32.15 sync (sync15nt) - #213
docs(vision): campaign 2026-08-24 — nvfp4 think-on on the v0.32.15 sync (sync15nt)#213glennneuber wants to merge 2 commits into
Conversation
…nc (sync15nt) The think-on half of the sync-0.32.15 parity campaign, re-run with the graph-cache thrashing check off (#211/#212), plus qwen3.6 nvfp4 in both modes. Standard tables, the ladder glossary (converged / capped / NOT CONVERGED), no-regression analysis against the pre-sync 31b repeats, the loop ranking with the 131072-rung timeout evidence, and the 26b clean-but-wrong open item. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Reviewing as consolidator. This is the payoff measurement for #211/#212 and it reports the result the way the week's corrections asked for. The before/after is the number I would lead with: the first attempt lost 17 of 27 cells to runner aborts; this run has zero server errors across ~14 h and converges 25 of 27, with the two remaining marked NOT CONVERGED rather than scored. That is the fix working, and it is stated as a campaign outcome rather than a claim about a patch. The glossary earns its placeDefining converged / capped / NOT CONVERGED in the document — and specifically saying "converged does not mean correct — a converged arm can still score 1/6" — is the direct answer to what #202 found: a capped cell rendered identically to a converged one, and a row of zeros reached a published page looking like a result. A reader now cannot take "converged 25/27" as a quality statement, which is exactly the confusion that produced the Recording NOT CONVERGED as a finding about the model at a stated ceiling rather than as a score is the H4a position stated where it will actually be read. Two things I checked from my sideNo sync regression, quantified. 31b within ±0.005 IoU of the three pre-sync repeats is the right comparison to make, and it is made against repeats rather than a single prior value — which matters, because #196 showed the pre-sync spread rendering was halving intervals, and #205 showed a 1-in-3 failure reading as flakiness. A ±0.005 claim only means something if the baseline's own spread is known. The 131072 rung yielding only HTTP_TIMEOUT errors is a genuinely useful negative: it bounds the ladder empirically rather than by assumption, and it is the sort of thing that would otherwise be re-attempted every campaign. Worth making sure that lands in the README rung table too — #209 already corrected it once for stopping at 65536, and a rung that exists but only produces timeouts is a different fact from a rung that exists. The open item is correctly scopedThe 26b Carrying ADR 0023's off-policy banner at the top is right, and worth keeping even though ADR 0029 superseded the sampling half — the banner is what stops a reader comparing these think-on cells against card-sampled campaigns without noticing they are different arms. No objection. This is what a campaign document looks like after a week of learning what they get wrong. |
…ions pointer Merged tables for both think modes across the sync15/sync15nt tags, the same-weights pair ratios (score parity; GGUF 1.5-3.6x req/h on this host; MLX prefill fastest, short-answer decode the gap), the cross-engine think-on hazards, and the explicit note that the serving decision, if taken, becomes an ADR citing this section. vision-model-recommendations.md gets a dated pointer. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The write-up for the sync15nt campaign (2026-08-23/24): the think-on half of the sync-0.32.15 parity series, re-run with MLX's graph-cache thrashing check off (#211, fixed at the runner in #212), plus the first valid qwen3.6:35b-a3b-nvfp4 cells in both think modes.
Contents: standard summarizer tables for both modes; a "Reading the ladder" glossary defining converged / capped / NOT CONVERGED and the per-cell rung outcomes; findings — no sync regression (31b within ±0.005 IoU of the three pre-sync repeats), the loop ranking (12b ≫ qwen3.6 > 26b) with the measured evidence that the 131072 rung yields only HTTP_TIMEOUT errors, qwen3.6's think=off recommendation, zero server errors across ~14 h with the check off, and the 26b clean-but-wrong
bboxm_free_*open item (n ≥ 3 repeat needed). ADR 0023's off-policy banner carried at the top.🤖 Generated with Claude Code