Skip to content

SFM Safe Flow Expansion: confirmed fixed recipe (CR −.100, Validity +.133, clearance +.014 @ M100 paired CIs) + pre-registered MPC-distillation follow-up - #3

Draft
DHLeexpress wants to merge 31 commits into
masterfrom
agent/claude-sfm-mpc-distill-20260726
Draft

SFM Safe Flow Expansion: confirmed fixed recipe (CR −.100, Validity +.133, clearance +.014 @ M100 paired CIs) + pre-registered MPC-distillation follow-up#3
DHLeexpress wants to merge 31 commits into
masterfrom
agent/claude-sfm-mpc-distill-20260726

Conversation

@DHLeexpress

@DHLeexpress DHLeexpress commented Jul 26, 2026

Copy link
Copy Markdown
Owner

Recipe study (branch lineage: agent/claude-sfm-best-recipe-20260726, contained here)

Confirmed result (disjoint M100 bank ep0 330000, 700 CRN rollouts/method, temperature 1.0, no tilt/verifier/fallback; paired scenario-cluster 95% CIs vs r0):

SR CR Validity succ. clearance succ. time
r0 .643 .353 .603 .108 8.76 s
selected .723 .253 .736 .122 10.84 s
locked Kazuki (.3/.5) .827 .173 .353 .168 4.17 s

ΔCR −.100 [−.157,−.043] · ΔValidity +.133 [+.114,+.153] · Δclearance +.014 [+.005,+.023] · Δtime +2.08 s [1.84,2.32].

Selected recipe (round-invariant, from exact r0 SHA 1b5179c9…): margin execution selector, original replay ∪ exact-certified deterministic recovery positives (per-row SOCP certificate audit), α=.01, 100 exposure epochs, lr 1e-4, ESS .5; selected round 1 by the predeclared M50 rule on a fresh bank (only eligible round; rounds 2–4 collapse to timeout — reported plainly, no monotonic-learning claim).

Mechanism (figures in /data3/research1/claude_sfm_best_recipe_f06e8dd/mechanism/): collisions are decided in a 2–4-step closing certifiability window at congested episode starts — flow support crashes to 1–2/16 at NVP onset while deterministic certified escapes still exist; a few steps deeper nothing certifies. Recovery data moves flow mass into that window (NVP −25%, both diagnosed collision lineages → closed-loop successes).

Integrity: predeclared disjoint banks, frozen selection rules before every read, iteration-1 null result (cost/hard) fully reported with its own 11-round M50 curve and paper plot, Kazuki locked, temperature 1.0 everywhere, 33-artifact SHA-256 manifest, DELIVERY_COMPLETE.json. Additive modules only — flags off ⇒ bitwise-identical replay; existing 43-test suite green.

MPC distillation follow-up (this branch's tip; pre-registered, COMPLETE — fail-closed at M10)

Imports the Codex privileged candidate-pool logic read-only from agent/sfm-adhoc-controller-overlay-20260726@0e0eca2 (nothing on that branch touched). Separate per-round D_MPC+ (never in D+/GP/acquisition): privileged exact-SFM look-ahead ∧ exact full-H10 SOCP, ranked by native SafeMPPI cost, distilled in dedicated post-round gradient blocks with fresh Gaussian CFM bases; per-block before/after audits; every pre/post checkpoint kept.

RESULT: the dedicated blocks do visibly change the raw policy — SOCP-positive rate at MPC contexts up in 15/16 blocks (up to .106→.482), target-recovery RMSE down in 16/16, present even when the block's CFM loss rose — but no swept dose passes the pre-registered liveness gate on the fixed M10 bank (0/32 cells; post-block round-1 SR .243–.600 vs gate .637), the no-distill control collapses identically by rounds 2–3, and the M50/M100 confirmations are therefore not triggered (MPC_STUDY_DELIVERY.json, mpc_distill/MPC_M10_SELECTION.json).

Synthesis across both studies: hard-context-only distillation (deterministic escapes or privileged MPC pool alike) trades global liveness for local certifiable support; the confirmed recipe wins by embedding certified targets in the full replay mixture (~5% mass). Composition, not target quality, is the binding constraint.

Study root: /data3/research1/claude_sfm_best_recipe_f06e8dd (all artifacts + manifests).

🤖 Generated with Claude Code

https://claude.ai/code/session_015CpH3uJa2LFTQ6m8xvXNXa

DHLeexpress and others added 30 commits July 23, 2026 14:43
…s + exact-certified deterministic recovery C), extended trainer (lr/ESS/rounds/replay-mode knobs, immutable core reused), Stage-A branch diagnostics, locked-Kazuki fixed-bank evaluator with executed-window Validity; predeclared banks + research log; 6 new tests green, default-OFF bitwise equivalence asserted

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…full D+/D- plus appended certified recovery positives; 7 tests green

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…RULE + Stage-A/B results in log

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…luation) + mechanism snapshot/episode viz + M100 paired-difference analysis; 9 tests green

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d r7: CR -.006, V -.067 vs r0); paper-contract plot rendered; iteration-2 pre-registration (R_A=B9 knobs vs R_B=v2 family, freeze criterion declared before U10-12 reads; fresh M50 bank 340000; M100 330000 untouched)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…recovery); v2 documented as follow-up; M50 selection launched on fresh bank 340000

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… dSR +.091 vs r0 on fresh bank 340000; paper plot rendered; M100 confirmation launched on untouched 330000

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d-only @0e0eca2), separate per-round D_MPC+ (never in D+/GP/acquisition), intersection filter (privileged SFM look-ahead AND exact H10 SOCP) ranked by native SafeMPPI cost, dedicated distillation blocks with fresh CFM bases, before/after local+M10 audits; 3 tests green incl. exact prefix-replay reconstruction

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… 360000, M100 370000), selection rule, separation guarantees

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…+.133, dclr +.014, dt +2.08s vs r0), per-gamma tables, 33-artifact SHA manifest, DELIVERY_COMPLETE.json; MPC follow-up log

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…icy change confirmed (SOCP-rate up 15/16 blocks, target-recovery down 16/16, present even with rising loss) but 0/32 cells pass the pre-registered liveness gate; no-distill control collapses identically; M50/M100 not triggered; all pre/post checkpoints + audits kept

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…eeze + ode_times tuple already in); this is the smoke-validated version used by the sweep

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@DHLeexpress

Copy link
Copy Markdown
Owner Author

Follow-up study delivered on branch agent/claude-sfm-r1-stable-continuation-20260727 (descends from this PR's head): pre-registered stable-continuation search from the confirmed r1. Outcome: 0/16 round-1 candidates admissible under the hard gates (timeout binds at every dose; CR fell to .129 and Validity rose to .79 but liveness always paid), so the frozen winner degenerated to immutable r1 — declared before the final bank was read. Final M100 (ep0 410000, paired CIs): r1 vs locked Kazuki = Validity +.384 [+.354,+.412] ahead; CR +.089, clearance −.055, time +6.49 s behind → full Pareto dominance NOT achieved, reported metric-by-metric. Artifacts + 26-entry SHA manifest under /data3/research1/claude_sfm_best_recipe_f06e8dd/r1_continuation/.

@DHLeexpress

Copy link
Copy Markdown
Owner Author

Third follow-up delivered on agent/claude-sfm-unverified-teacher-run-20260728 (pinned tool @2a873327 used as-is + 3 additive commits): unverified-MPC-teacher continuation, pre-registered. Outcome: fail-closed at round 1 — the UNCHANGED ordinary e10 replay round from r1 already collapses liveness on the dev bank (SR .77→.44, timeout .04→.43, while CR .186→.129 and V .752→.836), the nine teacher doses (lr 1e-6..1e-5 × 1..4 ep) are epsilon on top, and the conflict probes show the teacher gradient is near-orthogonal to certified D+ (cos .03→−.07) while aligned with rejected D− (cos .38). 0/10 admissible ⇒ M50/M100 banks untouched, no Kazuki-win claim. Identical D_MPC across arms enforced by canonical content hash (torch.save bytes proven non-deterministic across equal payloads). Artifacts + manifest under /data3/research1/claude_sfm_best_recipe_f06e8dd/teacher_rounds/.

@DHLeexpress

Copy link
Copy Markdown
Owner Author

Fourth follow-up on agent/claude-sfm-corrected-privileged-distill-20260728 (frozen 736fa6c): corrected-replay privileged distillation study. Phase 1: the historical privileged deterministic SFM controller goes 70/70 on the current severe OOD (SR 1.000/CR 0.000, clearance .233, time 4.47 s; exact-SOCP window Validity .414, audit-only). Phase 3 fixes verified: original accumulated one-Adam-step-per-epoch replay (fail-closed W=2), exact L_T=ΣmᵢL_CFM(i) teacher weighting with a production-policy chunked-vs-direct gradient test (<1e-5), label-based baseline reader (prior r0-as-r1 read disclosed; no conclusion flips). Phase 4: the corrected teacher measurably moves r1 (drift up to 1.4e-3, RMSE .046) and the corrected ordinary one-step update is dramatically gentler than the old minibatch semantics — but the disjoint M50 confirms no arm beats r1 (best: corrected-ordinary V +.016/CR −.006, within noise; teacher M10 CR gains reverse). Gradient anti-alignment persists under the corrected objective: cos(teacher, D+)=−.07, cos(teacher, D−)=+.57 — even a perfect controller's successful executed windows point toward what the certificate rejects. Full Pareto tables, paper-contract plot, active-dodge PNG/MP4, 40+ artifact manifest under /data3/research1/claude_sfm_corrected_privileged_distill_736fa6c/.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant