Add paired neutral multiround correction study - #4
Draft
DHLeexpress wants to merge 56 commits into
Draft
Conversation
…s + exact-certified deterministic recovery C), extended trainer (lr/ESS/rounds/replay-mode knobs, immutable core reused), Stage-A branch diagnostics, locked-Kazuki fixed-bank evaluator with executed-window Validity; predeclared banks + research log; 6 new tests green, default-OFF bitwise equivalence asserted Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…full D+/D- plus appended certified recovery positives; 7 tests green Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…RULE + Stage-A/B results in log Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…luation) + mechanism snapshot/episode viz + M100 paired-difference analysis; 9 tests green Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d r7: CR -.006, V -.067 vs r0); paper-contract plot rendered; iteration-2 pre-registration (R_A=B9 knobs vs R_B=v2 family, freeze criterion declared before U10-12 reads; fresh M50 bank 340000; M100 330000 untouched) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…recovery); v2 documented as follow-up; M50 selection launched on fresh bank 340000 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… dSR +.091 vs r0 on fresh bank 340000; paper plot rendered; M100 confirmation launched on untouched 330000 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d-only @0e0eca2), separate per-round D_MPC+ (never in D+/GP/acquisition), intersection filter (privileged SFM look-ahead AND exact H10 SOCP) ranked by native SafeMPPI cost, dedicated distillation blocks with fresh CFM bases, before/after local+M10 audits; 3 tests green incl. exact prefix-replay reconstruction Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… 360000, M100 370000), selection rule, separation guarantees Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…+.133, dclr +.014, dt +2.08s vs r0), per-gamma tables, 33-artifact SHA manifest, DELIVERY_COMPLETE.json; MPC follow-up log Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…icy change confirmed (SOCP-rate up 15/16 blocks, target-recovery down 16/16, present even with rising loss) but 0/32 cells pass the pre-registered liveness gate; no-distill control collapses identically; M50/M100 not triggered; all pre/post checkpoints + audits kept Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…eeze + ode_times tuple already in); this is the smoke-validated version used by the sweep Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…90k dev / 400k M50 / 410k M100 untouched, hard admissibility, lexicographic acceptance, fixed-recipe rerun rule), r1 self-anchor collector, shared-shard 16-candidate continuation engine with declared mass composition (anchor+5% recovery+original) and full per-candidate logging; 5 tests green Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e (round-1 crash fix) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…missible at round 1, timeout gate binds at every dose); frozen winner = immutable r1, declared BEFORE final M100 read Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… at all doses); frozen winner = immutable r1; final M100 verdict = NOT a Kazuki win (Validity +.384 ahead; CR +.089, clearance -.055, time +6.49s behind, all CI-clean); delivery + manifest Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…430k), additive default-off --reuse-buffer (authenticated against round-shard SHA), round driver (ordinary gather+replay unchanged -> theta_half -> 9 pinned-tool dose arms with buffer-SHA equality assertion -> M10 gates -> lexicographic acceptance incl. no-teacher control) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
….json only) so cross-arm buffer SHA equality holds; teacher tests green Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
….save bytes are not deterministic across fresh-vs-reloaded object graphs; verified deep-equal payloads with differing file SHAs, canonical hash unifies all 9 crashed-run buffers) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e collapses liveness (SR .77->.44) before any teacher; 9 doses epsilon on top; gradient-level finding cos(teacher,D+)=.03->-.07 vs cos(teacher,D-)=.38 — unverified MPC teacher anti-aligned with certified behavior; 0/10 admissible, M50/M100 untouched, no Kazuki claim; delivery + manifest + branch viz Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… teacher / 390k M10 eval / 450k M50; r0-as-r1 baseline-reader bug disclosed, conclusions re-checked unchanged), corrected ordinary B1 (BS.signed_update on fail-closed W=2 holder, one accumulated Adam step per epoch), corrected teacher objective L_T=sum m_i L_i (chunk weights b*m_i, summed, per-chunk seeded), production-policy chunked-vs-direct gradient agreement test (<1e-5), label-based reader, Phase1+2 online qualification + faithful executed-action D_MPC with fail-closed context reconstruction, Phase-4 causal study driver (A/B/C/D arms, Pareto table, declared promotion) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…470k/480k, golden SOCP-only authority, O/F/B8 support benchmark, set-valued certified CFM U/G with tau_J-calibrated Gibbs, corrected W=2 whole-dataset replay, staged 8-arm screen -> 2x10-round expansion -> M50 -> M100); core module + driver (phase0+screening) + 5 tests green Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…fied replay, resumable), M10 evolution 0/1/2/5/10, per-recipe best, M50 both, winner -> untouched M100 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rt .938, closed-loop certified NVP .986, CR 0); winner ess0p25_U_d1 r1; M100 within-noise gains (dCR -.006, dV +.004); verifier-support gap confirmed binding Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… windows mixed at portion p into corrected low-dose replay, 4 arms x 3 rounds, goal-forgetting guarded by low dose + certified-gather liveness; labeled UNFAITHFUL throughout Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… round Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds a frozen two-scenario-per-round SFM expansion study that applies whole-population D+ and neutral D0 updates with identical doses and a persistent Adam optimizer.
The study freezes exact guidance-trigger contexts, pedestrian states, original K=16 latent banks, and B=4 acquisition IDs, then measures local verifier repair before training, after D+, and after D0. It also provides staged disjoint raw M20 evaluation and a four-arm round-50 launcher.
Validation: 46 focused/dependency tests passed; the full local analysis suite had 198 passes and one pre-existing environment failure caused by a missing Helios-only dataset path. Python compilation, diff checks, and launcher shell syntax passed.