Skip to content

docs(vision): campaign 2026-08-24 — nvfp4 think-on on the v0.32.15 sync (sync15nt) - #213

Open
glennneuber wants to merge 2 commits into
mainfrom
docs/vision-campaign-2026-08-24-sync15nt
Open

docs(vision): campaign 2026-08-24 — nvfp4 think-on on the v0.32.15 sync (sync15nt)#213
glennneuber wants to merge 2 commits into
mainfrom
docs/vision-campaign-2026-08-24-sync15nt

Conversation

@glennneuber

Copy link
Copy Markdown

The write-up for the sync15nt campaign (2026-08-23/24): the think-on half of the sync-0.32.15 parity series, re-run with MLX's graph-cache thrashing check off (#211, fixed at the runner in #212), plus the first valid qwen3.6:35b-a3b-nvfp4 cells in both think modes.

Contents: standard summarizer tables for both modes; a "Reading the ladder" glossary defining converged / capped / NOT CONVERGED and the per-cell rung outcomes; findings — no sync regression (31b within ±0.005 IoU of the three pre-sync repeats), the loop ranking (12b ≫ qwen3.6 > 26b) with the measured evidence that the 131072 rung yields only HTTP_TIMEOUT errors, qwen3.6's think=off recommendation, zero server errors across ~14 h with the check off, and the 26b clean-but-wrong bboxm_free_* open item (n ≥ 3 repeat needed). ADR 0023's off-policy banner carried at the top.

🤖 Generated with Claude Code

…nc (sync15nt)

The think-on half of the sync-0.32.15 parity campaign, re-run with the
graph-cache thrashing check off (#211/#212), plus qwen3.6 nvfp4 in both
modes. Standard tables, the ladder glossary (converged / capped / NOT
CONVERGED), no-regression analysis against the pre-sync 31b repeats, the
loop ranking with the 131072-rung timeout evidence, and the 26b
clean-but-wrong open item.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@glennneuber

Copy link
Copy Markdown
Author

Reviewing as consolidator. This is the payoff measurement for #211/#212 and it reports the result the way the week's corrections asked for.

The before/after is the number I would lead with: the first attempt lost 17 of 27 cells to runner aborts; this run has zero server errors across ~14 h and converges 25 of 27, with the two remaining marked NOT CONVERGED rather than scored. That is the fix working, and it is stated as a campaign outcome rather than a claim about a patch.

The glossary earns its place

Defining converged / capped / NOT CONVERGED in the document — and specifically saying "converged does not mean correct — a converged arm can still score 1/6" — is the direct answer to what #202 found: a capped cell rendered identically to a converged one, and a row of zeros reached a published page looking like a result. A reader now cannot take "converged 25/27" as a quality statement, which is exactly the confusion that produced the nemotron3 retraction.

Recording NOT CONVERGED as a finding about the model at a stated ceiling rather than as a score is the H4a position stated where it will actually be read.

Two things I checked from my side

No sync regression, quantified. 31b within ±0.005 IoU of the three pre-sync repeats is the right comparison to make, and it is made against repeats rather than a single prior value — which matters, because #196 showed the pre-sync spread rendering was halving intervals, and #205 showed a 1-in-3 failure reading as flakiness. A ±0.005 claim only means something if the baseline's own spread is known.

The 131072 rung yielding only HTTP_TIMEOUT errors is a genuinely useful negative: it bounds the ladder empirically rather than by assumption, and it is the sort of thing that would otherwise be re-attempted every campaign. Worth making sure that lands in the README rung table too — #209 already corrected it once for stopping at 65536, and a rung that exists but only produces timeouts is a different fact from a rung that exists.

The open item is correctly scoped

The 26b bboxm_free_* being clean-but-wrong with an n ≥ 3 repeat named as the requirement, rather than explained now, is the discipline #205 established: a plausible mechanism plus one observation is not a finding. Leaving it open with the repeat count attached is better than a paragraph of hypothesis.

Carrying ADR 0023's off-policy banner at the top is right, and worth keeping even though ADR 0029 superseded the sampling half — the banner is what stops a reader comparing these think-on cells against card-sampled campaigns without noticing they are different arms.

No objection. This is what a campaign document looks like after a week of learning what they get wrong.

…ions pointer

Merged tables for both think modes across the sync15/sync15nt tags, the
same-weights pair ratios (score parity; GGUF 1.5-3.6x req/h on this host;
MLX prefill fastest, short-answer decode the gap), the cross-engine
think-on hazards, and the explicit note that the serving decision, if
taken, becomes an ADR citing this section. vision-model-recommendations.md
gets a dated pointer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant