null SURVIVES full identity removal → amnesic-probing objection defeated. Non-gating; verdict unchanged (strengthened). Litscan debt paid; DECISIONS #014 (+add-1, add-2).
- Litscan debt PAID first (CLAUDE.md rule 3): Ravfogel "Null It Out" (2004.07667) INLP Algorithm 1 + Elazar "Amnesic Probing" (2006.00995) protocol verified from ar5iv full text; the two code-load-bearing facts (I−BBᵀ rowspace-projection form; rank-matched "Rand" control) recorded in litscan/NOTES.md (2026-07-09).
- Built (src/metis/fit_inlp.py, inlp_arm.py): per node, iterate the exact #006 probe recipe on the residual (INLP) to remove the linearly-decodable "node-in-depth-1- frontier" subspace; measure the answer-branch SUBTRACT ΔT on train[:100] (same set, op, norm-preserving #005 as the probe-rung sanity) under single-direction / full-INLP- subspace / rank-matched-random deletions. Recipe + reading PINNED before results (#014).
- Fit facts: 23 nodes (same as #006). Identity is HIGHLY redundant — at k=8 residual decode AUC still ~0.82–0.93 (so #006's single direction badly under-deletes, exactly the amnesic objection); driving toward chance needs ~tens of directions (median k=40 reaches AUC ~0.55–0.65; node 22 hits 0.50 at k=2).
- Result (results/s1_seed0_win/inlp_arm_summary{,_full}.json), both configs: single-direction median ΔT −0.0 (reproduces basis_sanity_probes.json −0.0 → harness validated) INLP k≤8 median ΔT −0.0 (T_post 99.99) INLP k≈40 (~chance) median ΔT −0.0 (T_post 99.99) rand-matched median ΔT 0.0 at both ranks Removing the full linearly-decodable identity subspace leaves behavior unchanged, and no more than removing the same number of RANDOM directions. Per Elazar's own criterion (Rand ≈ INLP ⇒ property not causally important), the readable-but-causally-inert result is NOT a single-direction under-deletion artifact. Objection DEFEATED.
- Dynamic-range caveat (honest): baseline T=100 on training graphs (same as the probe- rung sanity), and k-dim removal + renorm is a modest perturbation. But the SAME measurement detects large causal effects from saturation — structural transplant moved T −91.7 (#010), COLLAPSE moves T (gate_summary) — so the identity-removal null is a real inertness, not a dead instrument. Apples-to-apples with frozen Gate A.
- Compute note: the pinned k=60/200-iter _full config was infeasible (~54 min, killed, no artifacts — LBFGS deep-round blowup); reduced to k=40/50-iter from a cached thoughts tensor (DECISIONS #014 addendum-2, user-approved), reading unchanged.
- Open menu now: S2 bounded build (optional v2) / write-up. Git push still PENDING (robustness episode + this INLP arm are all uncommitted).
graph short of ≥93%); #013 pre-registered fallback fires. Robustness = seed-1 + test-split. Readout replicates in all four seeds. No seed-4.
- seed-3 finished clean: 70 full-task epochs, val peaked 0.9767 (best of any seed), but test 389/419 = 0.9284 < 0.93 → FAIL by ONE graph. Honest fail; bar frozen, not moved. Readout ordering_ok all 4 steps (magnitudes ≈ the passing seeds).
- Recipe-v2 tally: seed-0 94.5% PASS, seed-1 95.7% PASS, seed-2 90.2% FAIL, seed-3 92.84% FAIL → 2/4 pass. Substrate accuracy is seed-variable in the 0.90–0.96 band and the ≥93% bar sits inside it. The BFS-wave readout (the core structural claim) replicates in ALL FOUR seeds regardless — structure robust, answer-conversion the seed-variable part.
- Executed #013 fallback (pre-approved branch, not a new fork; DECISIONS #013 addendum): robustness stands on seed-1 (PASS + Gate-A null) + held-out test-split (null). seed-2/seed-3 reported as honest misses, NO transplant on either (frozen rule: precondition fail → never gates). NO seed-4 — rolling seeds till one passes = selection bias, disallowed.
- Robustness episode now COMPLETE. Verdict untouched (robustness-only throughout). Next: commit + push the whole episode (pending user go); then back to the open menu (S2 bounded build / INLP arm) per earlier go/no-go.
(test 90.2% < 93%); readout replicates. Per DECISIONS #013 (user-approved): do NOT gate seed-2 — train fresh replacement seed-3 (launched). Non-gating.
- seed-2 finished (resumed run): reached stage 5, early-stopped after 42 full-task epochs on the patience-15 plateau, best full-task val 0.9455 (best.pt saved). Resume session epochs 95–141 ran ~68–73 s each (0.95 h) — Epic-kill + nighttime.
- Acceptance (results/s1_seed2_win/acceptance.json): test 378/419 = 90.2% → FAIL vs ≥93%. BUT the BFS-wave readout replicates cleanly: ordering_ok=True at all 4 steps, magnitudes ≈ seed-0/seed-1 (step-4 Optimal 8.71 vs 8.32 / 9.10). So the search structure is learned as well as the passing seeds; only answer-conversion on held-out test falls ~4–5 pts short. NOT the #004 starved-readout failure — it's a seed-specific generalization shortfall. Three-seed table: seed-0 94.5% PASS, seed-1 95.7% PASS, seed-2 90.2% FAIL (all readout-ok).
- Decision (DECISIONS #013, user-approved): frozen rule "precondition failure → fix training, never gates" ⇒ no transplant on seed-2; retain its artifacts and report the miss. Train a fresh qualifying replacement, s1_seed3_win (seed 3), same v2 recipe. If seed-3 also misses, robustness rests on seed-1 + test-split (both already replicate the Gate-A null) with seed-2/3 reported as honest misses.
- seed-3 LAUNCHED 22:40, fresh from epoch 0, detached (worker PID 35668), GPU free (Epic did not respawn), epoch 0 in 49.8 s → ~2 h ETA. Watcher armed; on completion run accept_s1 (--run-name s1_seed3_win), then the seed-3 Gate-A arm iff it PASSES.
at stage 3 / epoch 94; RESUMED from checkpoint (scientifically continuous) + GPU contention removed. Operational, not experimental — no gate data touched.
- Death: seed-2 (the chained run's 2nd leg) stopped ~21:24 at stage 3 / epoch 94, ~6 h in, curriculum incomplete. No traceback (stdout was block-buffered and lost; metrics.jsonl is per-epoch flushed so it is the true record → last good epoch 94). Windows sleep/hibernate confirmed still disabled (standby AC+DC = 0/never), so NOT a sleep event; most likely the chained run's host shell was torn down, or a CUDA OOM (see GPU note). Prereg §Seeds: two additional robustness seeds (non-gating); seed-1 done, seed-2 is the second — so completing it fulfils the registered plan.
- Resume, NOT restart: train_s1.py saves ckpts/{run}/latest_state.pt (model+optimizer
+py/torch RNG) every epoch → resumes at epoch+1 and appends metrics. ckpt (179 MB,
saved 21:24 w/ epoch 94) survived; relaunched 21:37 as a FULLY DETACHED OS process
(PowerShell Start-Process, not tied to any shell / to Claude Code),
-uunbuffered, → results/train_s1_seed2_resume.{out,err}. Real worker PID 34664 (uv python; venv shim PID 5940 is just the launcher). Printed[resume] from epoch 95and produced epochs 95–97 clean (stage-3 val 0.91). RNG restored ⇒ resumed run is continuous with epochs 0–94, not a fresh trajectory. - GPU contention removed (user-approved): RTX 2000 Ada Laptop, 8188 MiB, was 7933 MiB used (97%) with EpicGamesLauncher.exe (silent boot-autostart, Chromium renderer) co-resident on the GPU — a plausible OOM cause of the first death and an overnight risk. Killed the Epic family (PIDs 23864/29472/30340); training now sole GPU compute app. Reversible (re-autostarts on reboot).
- Speedup: post-resume epochs are ~68 s vs the ~250 s of the daytime run (the (midday) entry had flagged seed-2's slow epochs as daytime contention). ETA now ~1–2 h vs ~4 h.
- Next (auto, unattended): on train completion run accept_s1 (--run-name s1_seed2_win --device cuda); substrate accept = test acc ≥ 93% + BFS-wave readout, then the seed-2 M2 Gate-A arm (the actual robustness datapoint, mirroring seed-1). metrics diff stays uncommitted until seed-2 + acceptance complete (file is being appended).
Gate A null; S2 (GPT-2 COCONUT) blocked on Windows harness → needs a driver build
- Seed-1 replication (independently trained model, substrate PASS 95.7%; results/s1_seed1_win/m2_gateA_summary.json): candidate transplant median ΔT −0.0, fraction 0.0, self-transplant −0.0, placebo 0.0, zero −36.79 (instrument live). Gate A fails identically on a second seed → the structure/identity result is NOT a seed artifact.
- Test-split replication (held-out TEST graphs never trained on; registered non-gating, PREREG_M2 §4; m2_gates.py --split test, seed namespace gi+1e6; results/s1_seed0_win/m2_gateA_testsplit_summary.json): candidate ΔT −0.01, fraction 0.08, self −0.0, placebo 0.0, zero −25.43 (instrument weaker on harder held-out graphs but still clearly live). Same null → NOT memorized training graphs. (sample_size_rule flagged only on the zero-health margin, a non-gating instrument check; the candidate null is far from every threshold.)
- INLP arm still pending — requires verifying Ravfogel/Elazar against the actual papers first (litscan debt); deferred to that verification.
- S2 FEASIBILITY (found via smoke, as designed — DECISIONS #012): vendor/coconut run.py is hard-wired to Linux distributed (torchrun rendezvous needs libuv; dist.init_process_group("nccl") — NCCL is Linux-only). On this Windows single GPU the launcher fails at libuv and would next fail at nccl. Verdict: the vendor TRAINING HARNESS cannot run here; the vendor MODEL (Coconut class) is fine. S2 therefore needs a single-PROCESS driver — the same shape of task the S1 fast port was. Good news: FastCoconut (src/metis/fast_coconut.py) subclasses the vendor Coconut and is architecture-agnostic (wraps any HF causal LM), so it can wrap GPT-2 124M; the GPT-2 data pipeline comes from vendor/coconut dataset.py (BPE-tokenized NL ProsQA, curriculum staging). Because FastCoconut is OUR port, a forward-equivalence check vs vendor Coconut-GPT2 REACTIVATES before any S2 training counts (WINDOWS_SETUP discipline / #012 clause 1).
- Net: S2 is a bounded BUILD (driver + equivalence smoke, then multi-hour train, then S2 acceptance, then the S2 transplant arm), not the overnight launch it looked like. Surfacing the changed cost for a go/no-go before sinking the port time (reviewer holds S2 = optional v2; S1 preprint is complete now).
dissociation is now one monotone table. (Post-verdict; cannot flip Gate A.)
- tests/diag_m2_residual.py → results/s1_seed0_win/diag_m2_residual.json.
- (a) Δ = own − candidate_swap_donor (intermediate passes, n=256): median fraction of ‖Δ‖² inside span{u_T, u_D} = 0.151 vs random-2D baseline 0.0016 → 94× enrichment; |cos(Δ, u_T−u_D)| = 0.29. The two candidate node-embedding directions (2 of 768 dims) are ~94× over-represented in the label-swap residual: the readable identity readout is a dominant, concentrated component of what the swap changes. (Not 100% — ~85% is distributed relabeling knock-on — so the honest statement is "hugely enriched," not "exactly the wave.") Confirms the reviewer's correction quantitatively.
- (c) donor-vs-own cosine table, each paired with its already-measured causal ΔT — the whole dissociation in four rows: placebo (unreachable non-cands, wave≈0) cos 0.991 → ΔT 0.0 candidate (D≈0, T high) cos 0.998 → ΔT ~0 (Gate A) interior (path node, wave high) cos 0.771 → ΔT −0.08 different GRAPH (structure differs) cos 0.341 → ΔT −91.7 (#010) Reading: identity-perturbing swaps span cos 0.77–0.998 and are ALL causally inert — the interior row is the sharpest single point (a cos-0.77 change to readable path-node identity moves behavior by −0.08). Only the structural change (cos 0.34) is causal. Cosine-to-own tracks how much readable identity you perturb; causal ΔT tracks only structure. Readable ≠ causal, one table.
- Dropped as misspecified (honest): a global own-vs-swap linear-separability classifier — the discriminating direction u_T−u_D is GRAPH-SPECIFIC, so a single global boundary can't separate the pairs (scored below chance); (a) is the correct per-graph form of that question. Not reported as a result.
- Verdict UNCHANGED. This closes the mechanism quantitatively; the paper's spine is the unified readable-but-causally-inert-identity / causally-potent-structure dissociation, measured by M1 (probe deletion), #009 (final-step-only handle), #010 (structural transplant), and M2 (label-swap null + this table).
the state" OVERCLAIMS and is falsified by our own data. Correct claim: candidate identity is READABLE BUT CAUSALLY INERT mid-search. Verdict UNCHANGED.
- Flagged by the external reviewer — correct catch, conceded. The (night, +1)
phrase "candidate identity is simply NOT in the intermediate state" must not
reach the writeup. It is contradicted by:
- acceptance.json: the BFS-wave inner-product readout (⟨thought, node embedding⟩) orders NotReachable < Reachable < Frontier < Optimal at ALL steps 1–4. Decoy D is NotReachable (⟨t,u_D⟩ ≈ 0); target T is Frontier/ Optimal (⟨t,u_T⟩ high). That IS candidate identity, readably present in the intermediate thought.
- METIS-1: frontier membership decoded by node identity at AUC 0.998 from the same vectors.
- Arithmetic (reviewer's, verified): norm ≈ 27.4, rel-L2 under label swap 0.0702 → the swap moves the thought ≈ 1.92 in L2; wave components are ~4–5 in inner product. cos 0.998 with a SYSTEMATIC 7% residual (min cos 0.76; 2 graphs flipped) does NOT establish "identity absent" — it is exactly consistent with the readable identity readout riding in that small residual.
- Corrected claim (verdict unchanged; mechanism SHARPER): candidate identity is READABLY PRESENT but CAUSALLY INERT mid-search; the causally potent payload is structural. Not a new phenomenon — it is METIS-1's probe-inertness (delete the readable identity direction → ΔT 0), #009's final-step-only identity handle, and M2's transplant null, the SAME dissociation measured three ways. One phenomenon, three lenses. This is a STRONGER spine than "identity absent."
- The (night, +1) H-struct / H-heal dichotomy was INCOMPLETE. cos 0.998 rules out H-heal (nothing large enough to heal) but not "encoded-yet-ignored." The true third option — H-INERT-IDENTITY (present, readable, causally bypassed) — is the disciplined reading. Role-binding framing for the writeup: #010's redirect means the transplanted STRUCTURE decides which structural role wins; the recipient's PROMPT binds roles to tokens at the readout wire; M2's null is then entailed (same structure ⇒ same roles ⇒ same binding).
- Gate A FAIL stands exactly as committed (3cf2868); this corrects mechanism WORDING, not the gate result. Next: quantitative closure — Δ = own − swap should concentrate in span{u_T, u_D} (proving the 7% residual IS the readable identity component) — then the writeup around the unified dissociation.
- Also on record (reviewer credit): frozen R1/R2 (escape statistic 0.0001, fraction-flipped 0.02) did their job — they are what make this null unambiguous rather than arguable.
intermediate thought is invariant to candidate identity (cos 0.998 under label swap) yet genuinely different across graph structure (cos 0.34). Structure/ identity factorization is now airtight. (Post-verdict; cannot flip Gate A.)
- tests/diag_m2_structure.py → results/s1_seed0_win/diag_m2_structure.json,
n=100, intermediate passes only, all donors captured under the RECIPIENT's
pinned serialization (only graph CONTENT varies):
- candidate_swap donor (same structure, T↔D labels) vs own thought: median cosine 0.9975, min 0.7625, median rel-L2 0.070. Swapping the two candidate labels barely moves the mid-search thought.
- different-graph donor (the #010 donor, which carried −91.7) vs own: median cosine 0.3407, max 0.7251. Genuinely different vectors.
- Verdict on the two readings: H-STRUCT, not H-heal. The ~0 Gate A effect needs no healing story — the label-swap donor thought is a cos-0.998 near-copy of the recipient's own, so its transplant is ~a self-transplant. Candidate identity is simply NOT in the intermediate state. The #010 redirect came from structural difference (cos 0.34), and it is exactly that difference the mid- search thought carries.
- Full factorization, every claim on file: (1) thought ≠ decomposable branch-sum (METIS-1 KILL); (2) intermediate thought = transplantable STRUCTURAL search state, invariant to candidate labels (cos 0.998), varying with graph topology (cos 0.34), causal when structure differs (#010 −91.7); (3) candidate IDENTITY is bound at the readout step via prompt attention — the only linear identity handle anywhere is the final-step J-lens target↔decoy swap (#009 −99.97), never intermediate. Structure mid-search; identity at the wire.
- Minor honest thread for the writeup (not blocking): label-swap min cosine 0.76 means identity leaks slightly into the thought for a few graphs — a bounded leak, causally negligible (Gate A interior −0.08, candidate ~0). Footnote it.
- This is the mechanistic core of the paper. METIS-2 as a candidate-resolution claim is a pre-registered KILL; as a mechanism probe it delivered the sharpest positive characterization in the project. Next: writeup (structure/identity factorization), plus the registered non-gating robustness (seed 1 done at 95.7%; seed 2 training; test-split replication; INLP arm).
state is STRUCTURAL, not candidate-resolved. Gate B entailed-flat. Verdict locked.
- Gate A (results/s1_seed0_win/m2_gateA_summary.json, n=100): candidate-swap transplant median ΔT −0.0 (bar ≤ −30), median T_post 99.98 (bar ≤ 60), fraction ΔT≤−50 = 0.02 (bar ≥ 0.5), median escape 0.0001 → FAIL, and not marginally (sample-size rule not triggered; every margin enormous). Histogram: 78 graphs in [−10,0), 18 at ~0, only 2 fully flipped.
- The instrument is proven working by its own frozen controls: self-transplant 0.0 (plumbing exact), per-step zero −42.88 (sensitivity; 5th identical replication), placebo 0.0, noise floor 0.0, interior −0.08. A ~0 candidate effect sits between demonstrably-detectable extremes → the null is REAL.
- Donor validity confirmed (diagnostic, n=100): candidate-swap donors SOLVE their own label-swapped graphs — median T(answer=D) 99.82, 79% fraction correct. Genuine D-computing thoughts, transplanted into G, leave G's answer at T. The null is not a broken donor.
- FROZEN verdict semantics (§3) applied verbatim: Gate A fail = "#010's carriage was coarse (whole-graph state), not candidate-resolved — reported as the finding, no rescue, no post-hoc bars." Pre-registered reading; I called this exact outcome in the draft. METIS-1 KILL untouched; METIS-2 candidate- resolution claim is DEAD.
- Refined mechanistic picture (consistent across #009/#010/M2, each on file): the intermediate thought carries the STRUCTURE of the search (topology / positions / reachability) — a different-STRUCTURE donor redirects (#010, −91.7) but a same-structure / swapped-LABEL donor does not (M2, ~0). Token identity is bound LATE, at readout, via prompt attention — which is why the J-lens target↔decoy swap flips the answer only at the FINAL step (#009, −99.97) and never intermediate. Structure carried mid-search; identity bound at the wire.
- Gate B (m2_gateB_summary.json): entailed-flat — the candidate-swap effect is ~0, so mixing near-identical own/donor states gives a flat curve (range 0; 99/100 "unresponsive"; 1 "graded"). Not independently informative; Gate B cannot flip Gate A by design. Cross-check PASSED: α=1 endpoint reproduced Gate A(i) exactly (endpoint_vs_gateA 0.0) — runner cross-consistency.
- Next (post-verdict diagnostic, CANNOT flip): cosine(own, donor) on intermediate thoughts to distinguish H-struct (donor thoughts ≈ own; identity simply not encoded) from H-heal (donor thoughts differ but prompt attention heals identity downstream), with the #010 different-graph donor as the low-similarity anchor. Then the writeup — the structure/identity factorization is the mechanistic core.
DECISIONS #011); seed 1 precondition PASS; gates next
- Bar review held with the user; every bar approved; fraction-flipped GATES at ≥ 0.5 (user chose the stricter option). No bar weakened since draft — only constraints added. This commit is the freeze commit.
- Seed 1 (s1_seed1_win): substrate precondition PASS — test 95.70% (401/419, ≥ 93%), BFS-wave readout ordering intact at all steps 1–4 (results/s1_seed1_win/acceptance.json). Second seed to replicate the substrate. Seed 2 mid-training (stage 3; slower epochs from daytime machine contention; well under cap; chains into acceptance automatically).
- Next, strictly in order: m2_gates.py --phase gateA, then --phase gateB, on seed 0 (device cuda; runner now unlocked by this freeze).
DRAFT preregistration; date error corrected; still NO gate data
- External review received (user-supplied, from another Claude session; saved as review_20260708.md — advisory, non-governing, brief.txt precedent). Provenance caveat recorded in the file: its stated commit 1338fef does not exist here; every checkable content claim matched ea915b9.
- Verdict on the review: high quality; one genuine catch (R1) plus one real
error of ours it found (the misdated LOG heading, now corrected VISIBLY in
place). Integrated into PREREGISTRATION_M2.md — still DRAFT, still refusing
to run — as follows:
- R1 ADOPTED (necessary): Gate A(i) escape ceiling ≤ 0.05 — without it the gate could certify representation destruction as carriage (#010's own random-donor control is the demonstration: −30.4 ΔT at 0.98 escape). EXTENDED by us to Gate B: escape ceiling at the α=1 endpoint, per-α escape medians reported, pre-committed reading of elevated mid-curve escape (mixtures leaving the manifold — geometry finding, not gate failure).
- R2 ADOPTED as gating: Gate A(ii) fraction(ΔT ≤ −50) ≥ 0.5 + full per-graph histogram reported; escalation margin 0.05. (Review left gating ambiguous; we gate — a median can mask a large untouched minority. USER MAY DEMOTE to report-only at the bar review.)
- R3 ADOPTED: placebo-scope language in §2 (controls edit type and serialization, NOT state-difference magnitude; interior swap = informative middle condition, deliberately non-gating).
- R4 ADOPTED: test-split Gate-A replication registered in §4 (non-gating, n=100, own serialization-seed namespace).
- R5 ADOPTED: full transposed-depth-map invariant added to ALL THREE donor constructors (hard-fail); tests still 300/299/300, 20/20 prompt-diff.
- R6 ADOPTED: curve-shape taxonomy pre-committed (snap ≥80% of per-graph range in one increment / graded / unresponsive <30 range).
- Gate A(iv) self-transplant "= 0" reformulated to |median| ≤ 0.01 pp (fp-realistic; was already on the review agenda).
- Beyond the review: the amnesic-probing point is an EXPERIMENTAL gap, not just a citation debt — the kill-test's probe rung removed one direction per node; INLP-style subspace removal registered in §4 (non-gating; METIS-1 verdict scoped to its frozen ladder, not at stake). Verify-before-cite debts logged in litscan (Ravfogel, Elazar, Belinkov, Hewitt & Liang).
- m2_gates.py updated to compute every new clause; tests re-run PASS; both phases re-smoked on CPU (plumbing invariants hold; smoke slice remains effect-size-meaningless by design).
- Review's publication recommendation (stake the arXiv preprint on METIS-1 + arms NOW, in parallel) noted as consistent with the frozen plan; timing is the user's resource call.
- Freeze session (user, bar by bar) remains the next gate; runner still refuses.
findings for the bar review. NO gate data touched (runner refuses; verified). [Heading CORRECTED same day: originally misdated "2026-07-09 (early AM)" — a session-clock drift caught by the external review (review_20260708.md); actual time per seed-1 metrics timestamps was 2026-07-08 ~12:30. Content unchanged.]
- Plan approved (user): seeds overnight now; S2 deferred until after M2 gates; no venue-driven scheduling (rule 6 mechanical).
- Seeds: chained background run launched (train seed1 → accept → train seed2 → accept; ~7 h; results/train_s1_seed{1,2}.out).
- Built + verified pre-freeze:
- src/metis/minimal_pairs.py — label-surgery donors (candidate/placebo/ interior) with hard BFS validation. tests/test_minimal_pairs.py PASS: availability 300/299/300 on train[:300]; 20/20 candidate-swap prompts differ from recipients ONLY at transposed tokens under pinned serialization; invariants hold on all donors.
- src/metis/m2_gates.py — pre-registered runner; REFUSES while PREREGISTRATION_M2.md is DRAFT (verified: exit 1); --smoke = 5 graphs train[10000:10005], writes nothing. Both phases smoked on CPU: self- transplant ≡ 0.0, noise floor 0.0, gate B α=0 ≡ −0.0 and response monotone in α. Machinery sound.
- SMOKE DISCOVERY (pre-freeze, out-of-slice, diagnostic-tier): per-graph zero-ablation effects are strongly BIMODAL (≈0 or ≈−100; both devices, both slices) — small-n medians whipsaw (5-graph median −1.78 vs the stable −42.88 at n=100, 4 replications). Consequences for the review: (a) n=100 minimum for all gated medians is affirmed and should be stated as a frozen floor; (b) smoke checks are plumbing-only BY DESIGN — add explicit language; (c) cross-device per-graph values differ for saturated graphs under off-manifold inputs (zeroing) — gates stay single-device (cuda, pinned).
- Bar-review agenda for the freeze session (user, target 2026-07-10):
- Gate A(iii) "= 0" → fp tolerance |median| ≤ 0.01 pp (exactness is not device-realistic);
- zero health bar −30 at n=100: keep (4 replications at −42.9);
- n=100 floor language (bimodality);
- DECIDE: disclosed out-of-slice pilot of label-swap transplant before freeze (pro: the construction is novel, never causally tested; con: garden-of-forking-paths risk — if run, bars must be set before seeing it or the pilot must be published in the doc as prior data);
- Gate B mixing semantics: live-vector mixing, α=1 ≡ Gate A transplant, α=0 ≡ baseline — confirm as frozen text;
- escalation-set disjointness from the jlens fit slice (train[:300] vs train[500:2500]) — confirmed, note it in §2.
- Next: user bar review → freeze commit → gates (seed 0) → seed replications.
- PREREGISTRATION_M2.md written per #010's pinned next action. Core design: label-surgery minimal pairs (candidate-swap donor = same graph, T/D tokens transposed, prompts identical except swapped occurrences in edge tuples; placebo = distractor transposition; path-interior probe reported). Gate A = candidate-resolved carriage (transplant flips answer, placebo doesn't); Gate B = graded monotone dose-response under state mixing. Seeds 1–2 robustness retrains planned. S2 transfer registered as extension. Verdict semantics explicit: METIS-1 KILL is final either way.
- Bars are PROPOSALS awaiting user bar-by-bar review (v1.0 lesson: #001/#002 caught two invalid bars pre-freeze). Freeze target 2026-07-10.
- Next: user review → freeze commit → implement minimal-pair generation → gates. No experiment touches data before the freeze.
Intermediate thoughts ARE causal carriers of graph-specific search state — in a code no dictionary captures. All further runs STOPPED per the pinned rule; next action is drafting the METIS-2 preregistration with the user.
- Results (results/s1_seed0_win/explore_patch.json; n=100 recipients, 96 with
matched donors, pinned serializations, medians, pp):
- P1 plumbing: self_transplant 0.0 exactly; self_reserial_transplant −0.0 (thought state is not serialization-specific). Noise floor −0.0.
- P2 positive controls: donor_final_only −99.95, donor_all −99.97, escape ~0.0001 — clean directional flips, as #009 predicted.
- P3 DECISIVE — donor_intermediate (matched counterfactual donor's thoughts at all passes except last; final thought computed NATURALLY downstream): −91.69 with escape 0.0019. The recipient's own final readout, computed by the recipient's own forward over the recipient's own prompt, reads out the DONOR's answer. Contrast random_intermediate (wrong-graph thoughts, candidates disjoint): −30.37 with escape 0.9802 — non-matched content destroys the task representation entirely (98% of probability mass leaves the candidate pair), while matched counterfactual content cleanly REDIRECTS it. Directional, content-specific carriage — the exact P3 carrier signature pinned in #010.
- Dose/locality: donor_t0_only −2.91 — a single earliest-step transplant is mostly healed by prompt-driven recomputation; the redirect needs the accumulated multi-step state. (Connects to threat (e)/H-heal: single-point edits CAN be healed — which is why one-shot dictionary edits and single transplants do little — but the full mid-search state overrides the prompt.)
- Health check −42.88, fourth identical replication.
- Synthesis across #007/#009/#010 — the full picture, each piece on file: intermediate thoughts causally CARRY the search (transplant redirects, −92); the code is GRAPH-SPECIFIC and NON-DICTIONARY (every static per-node basis — geometric, decodability-optimal, causal-pullback — is inert mid-search; cross-graph Jacobian concentration only ~0.5 mid-search vs 0.93 at readout); the readout step alone has a clean global linear handle (J-lens swap 100→0). The pre-registered ontology (thought = branch-decomposable SUM) stays dead — the KILL verdict is untouched — but its replacement is now positively characterized: a causally potent, transplantable, non-decomposable state.
- Per #010's pinned decision rule: STOP. No further arms, no iteration. Next: (1) draft METIS-2 preregistration WITH user approval and freeze-by-commit — candidate core questions: what graph pair-structure the transplant effect respects (minimal pairs / single-edge counterfactuals), dose-response of state mixing (own vs donor interpolation — capacity question in its honest post-dictionary form), locality of the carried state (which steps carry what); (2) then the writeup, now a full positive+negative mechanism story.
- Publication posture per CLAUDE.md rule 6 unchanged: arXiv preprint after METIS-2's core result (or without it if the user prefers to stake now) → interp workshop → ICLR 2027 if earned. AAAI-27 remains moot.
handle even in a causally-defined dictionary; P1 positive control spectacular (clean 100→0 answer flip at the readout step). METIS-2 trigger NOT fired.
- Fit (results/s1_seed0_win/jlens_basis.pt + jlens_basis_report.json): 28 node tokens, K=4 per-step directions, 2,000 graphs, batched exact Jacobians, ~12 min. Early-tell diagnostics: the causal pullback found genuinely NEW directions — median |cos| to wte per step 0.41/0.55/0.77/0.78 (rotating toward the readout wire late, as tied weights predict), near-orthogonal to probe directions (0.03–0.13). Cross-graph concentration 0.50/0.49/0.71/0.93 — mid-search causal geometry is substantially GRAPH-DEPENDENT; only the readout direction is globally consistent. A global dictionary captures about half the mid-search gradient direction.
- Measurement (results/s1_seed0_win/explore_jlens.json; n=100, pinned
serializations identical to arm 1; medians, pp):
- P1 positive controls FIRED, harder than any effect yet seen: jswap_final (target<->decoy coordinate swap in the final thought) −99.97 with escape 0.0001 — a CLEAN total answer flip, mass moves entirely to the decoy; perstep_sub_path incl. final −99.49 (vs −65 for the wte basis in arm 1: the J-direction is a strictly better causal readout handle than the embedding).
- P2 decisive cells: perstep_sub_path_nolast −0.0; jswap_nolast −0.0. All same-v conditions −0.0. All random controls −0.0. Noise floor −0.0. Health check −42.88, identical third time running.
- Reading per the pinned #009 interpretation (no stretching): the negative SHARPENS — no per-node linear causal handle exists in intermediate thoughts under ANY direction-finding principle tried: geometric (wte), decodability- optimal (probes), or causal pullback (J-lens). The J-lens transfers to latent CoT beautifully — but only as readout-step control. METIS-2 trigger (strong selective intermediate effects) did NOT fire; per the pinned decision rule, proceed to #010: counterfactual patching / thought transplant (carrier-vs- cache) — the decomposition-agnostic test of whether intermediate thoughts causally carry ANYTHING per-graph, in any code.
- Doors explicitly left open (each its own entry if ever walked): LOCAL per-graph Jacobian directions (the concentration numbers hint the mid-search handle, if any, is graph-specific); J-space sparse-coding machinery; SAEs.
- Methodological bonus for the writeup: first application of the J-lens to latent CoT; the readout-step swap is a working intervention primitive (100→0, escape ~0) on a substrate where all static dictionaries fail mid-search.
the READOUT WIRE — final thought's target-embedding component; intermediate wave content causally bypassed even under continuous per-step pruning; probes inert everywhere. Verdict stands, mechanism sharpened.
- Ran DECISIONS #007 per-step re-application arm (pinned pre-run), then #008: the perstep_zero health check caught prompt-serialization noise (vendor expand_data shuffles edges via unseeded GLOBAL random, in place — every forward in every runner, frozen kill_test included, saw a fresh serialization; variance not bias, verdict unaffected, disclosed as measurement caveat). Arm re-run with per-graph pinned serialization + noise-floor condition; first-run artifacts kept (*unpinnedserial.json). Canonical numbers: results/s1_seed0_win/explore_perstep{inputemb,probes}.json.
- Pinned results (n=100 training graphs, medians, dT in pp; noise floor
baseline_reserial −0.0; health perstep_zero −42.88 IDENTICAL across bases,
matching #005's one-shot zeroing −42.9 — machinery validated):
- Input-embeddings basis: one-shot subtract −0.0; per-step same-v subtract −0.0; per-step PATH pruning (project out all valid-path node directions at every depth, incl. final) −65.08 with escape staying 0.033; path pruning WITHOUT the final step −0.01; random-span control −0.0.
- Probe basis (AUC 0.998 decoders): EVERY condition −0.0, including full path pruning with the final step.
- Reading (pinned interpretation, #007): H-INERT, not H-heal. Nothing upstream of the final step is causally read along branch-linear directions — continuous deletion at depths 1..d−1 does nothing, so there is nothing for later steps to "heal." The single causally potent linear component found anywhere on this substrate is u_wte(target) inside the FINAL thought — i.e., the answer readout wire itself (tied wte/lm_head), not search state. Deleting it flips the binary choice cleanly (T → ~35, mass moves to the decoy, escape tiny) — sharper and more targeted than zeroing (−42.9 with escape 0.76 = representation destroyed). And the probe direction that best DECODES the same information is causally inert even at the final step — a clean readable≠causal dissociation on latent CoT.
- Open (needs counterfactual patching / minimal pairs, next arm): whether intermediate thoughts matter through any NON-linear/nonlocal channel (zeroing all of them destroys T, so they carry something — but its causal route into the answer bypasses every per-branch linear component we can name).
- Writeup skeleton now: readout replicates (2505.12514) → readable (AUC .998) → load-bearing (zero −43) → per-branch linearly inert at every step and in every basis (KILL) → the one causal linear locus is the readout wire at the final step.
- Verified 2606.01243 ("Unlocking the Black Box of Latent Reasoning", Chang et al., v1 2026-05-31) against the actual full text (tables, figures, related work, limitations) — details in litscan/NOTES.md 2026-07-08 entry. Headline: NOT a collision. They steer WHOLE thoughts (slerp / gradient step / transplant / subspace projections) toward correctness on GSM8K/StrategyQA at 3–8B scale, gains +0.4–1.8 pts; no branch decomposition, no graph tasks, no capacity question, no negative results. They cite 2505.12514 but leave superposition unelaborated. Our kill verdict (per-branch linear surgery inert on the superposition substrate itself) is complementary — the writeup can use the contrast directly.
- Bonus corroboration: their edits are explicitly norm-preserving (slerp at ‖z‖, norm-constrained gradient) — independent support for the DECISIONS #005 correction.
- Writeup obligations recorded: cite as nearest neighbor; their transplant probe is precedent for our exploratory transplant arm (whole-sequence vs our carrier-vs-cache). Remaining skim-tier debt: 2604.04902.
- Next: post-verdict exploratory arms (adopt-list, LOG 2026-07-07 late), then negative-result writeup for arXiv.
causally-linearly editable on S1 (frozen §3 Gate A failure via basis exhaustion)
- Sequence (all artifacts in results/s1_seed0_win/): first sanity run (input-embeddings, frozen unit-norm ops) failed −0.72/0.95 (basis_sanity_unitnorm_v1.json) → mechanics diagnostics falsified §2's unit-norm premise (thought norms ≈ 27.4; the frozen ops' entire effect was an off-manifold scale crush) → DECISIONS #005: norm-preserving ops, bars untouched → corrected sanity failed CLEAN: −0.0/0.0 (basis_sanity_inputemb_normpreserving.json) → unembedding rung degenerate (tied weights, #005) → DECISIONS #006: probe directions, recipe pinned & committed pre-fit → probes decode frontier membership nearly perfectly (23 nodes, median holdout AUC 0.9979, min 0.63 — probe_basis_report.json) → probe-basis sanity failed IDENTICALLY: −0.0/0.0 (basis_sanity_probes.json). Ladder exhausted.
- Verdict per frozen §2/§3, no gate data ever collected (runner refused, correctly): continuous thoughts on S1 are not causally-linearly editable. No rescue, no post-hoc bars.
- The finding is sharper than a bare negative, and every piece is on file: (a) the substrate is healthy and replicates the paper (test 94.51%, BFS-wave readout at all steps — acceptance.json); (b) the step-1 thought is causally LOAD-BEARING (zeroing it: ΔT −42.9; random content at natural norm: −13.0 — DECISIONS #005 diagnostics, tests/diag_thought_norms.py); (c) frontier membership is almost perfectly linearly READABLE from the same vector (AUC 0.998); (d) yet deleting the readout direction — input-embedding OR probe, answer coefficient ≈ 32% of thought norm — produces median ΔT of exactly 0. Readable, load-bearing, and causally inert under linear per-branch surgery: the superposition READOUT (2505.12514) replicates; the superposition ONTOLOGY (thought = editable sum of branch components) does not survive causal test on its own best substrate.
- Redundancy (frozen threat (e)) is the leading mechanism candidate: the prompt stays in context, so later steps may reconstruct whatever the edit removed — distinguishable post-verdict via the adopted exploratory arms (per-step re-application, counterfactual patching / minimal pairs, transplant). These are EXPLORATORY: they cannot flip this verdict.
- Next, in order: (1) verify arXiv:2606.01243 (TOP-PRIORITY lit debt — nearest neighbor; required before any writeup); (2) post-verdict exploratory arms (LOG 2026-07-07 late adopt-list) as budget allows; (3) negative-result writeup targeting arXiv preprint. AAAI-27 is moot (bars not passed; CLAUDE.md rule 6).
- Training completed in 3.2 h wall-clock (160 epochs: fixed 25-epoch stages 0–3 per #004, then 60 full-task epochs, plateau early-stop patience 15 per #003/#004). Best full-task val 96.50%. Curve notes: stage 0 anchor hit (~77% at ep 20–24, matching Mac v2); stage 1's transition fired ~8 epochs later than Mac v2's (ended 47.9%, still climbing — investigated mid-run, train losses identical to Mac through ep 30, divergence = cross-device kernel numerics shifting transition timing, no stack fault); stages 2–3 transitioned early and strong (53.3%, 88.7%); full-task phase opened at 89.1% (run 1 best: 67.3%) and jumped above the bar at ep ~130.
- ACCEPTANCE (results/s1_seed0_win/acceptance.json): test accuracy 396/419 = 94.51% ≥ 93% PASS; BFS-wave inner-product readout ordering holds at ALL steps 1–4 (run 1 was broken at steps 2–3; here step 2: NR −0.04 < R +1.21 < F +2.48 < O +4.03; step 3: NR +0.20 < R +0.55 < F +2.91 < O +5.37). Both precondition clauses hold on the primary seed → gate data collection permitted.
- #004's diagnosis is confirmed by outcome: full 25-epoch stages 1–2 → intact mid-depth wave → generalization (94.5% test vs run 1's 60.9%).
- Next: kill_test.py --phase sanity (frozen basis sanity, node-embedding basis), then --phase gates. Bars per frozen §3; runner enforces order.
- Windows machine (i9-13950HX / 32 GB / RTX 2000 Ada 8 GB, driver 595.95) set up per WINDOWS_SETUP.md: python 3.12.13 (uv), torch 2.5.1+cu121, numpy 2.1.3, transformers 4.46.2, datasets 3.1.0, tqdm 4.67.0; vendor repos cloned.
- REQUIRED gate passed before any training: tests/test_fast_equivalence.py → EQUIVALENCE: PASS (tokenizer 200 samples identical; forward loss identical, max logit diff 0.00e+00; all parameter gradients identical). CPU-by-design (deterministic vendor-vs-fast comparison), unmodified.
- Windows sleep/hibernate disabled (standby-timeout AC+DC = never) before launch.
- s1_seed0_win LAUNCHED (v2 recipe, exactly per WINDOWS_SETUP.md: seed 0, --device cuda --fixed-early-stages --full-task-patience 15). Epoch 0: train_loss 2.263155, stage_val_acc 0.2879 (74/257), 56.4 s/epoch — vs Mac v2 epoch 0 loss 2.262385, 0.2996 (77/257), 274 s/epoch. Near-identical start, NOT bit-exact across CUDA/MPS kernels (expected; the bit-exact gate covers fast-vs-vendor code, not cross-device). ~4.9× Mac pace → projected full schedule ≈ 3–5 h, far under the 36 h cap. Canonical numbers: results/s1_seed0_win/metrics.jsonl.
- Sanity anchors from the stopped Mac v2 run to watch: stage 0 ≈ 77–79% by epochs 20–24; stage 1 crossing 50% around epoch 38. Next: acceptance → sanity → gates, strictly in order, when training ends.
- User supplied a planning brief from a parallel Claude session. Triage verdict: largely convergent with the frozen design (its prune/inject/collapse + topological placebos ≈ our Gates + controls; its Phase-1 readout gate ≈ our substrate precondition + basis sanity). The brief predates this repo's state and does NOT govern — PREREGISTRATION v1.0 does. Adopted as POST-VERDICT exploratory arms (cannot flip verdicts; DECISIONS entry when implemented): counterfactual patching with minimal pairs, full-thought transplant (carrier-vs-cache), ridge-decomposition subtraction arm, per-step re-application, wrong-step control, dose-response sweeps. Rejected: re-opening phases/substrate (S1 stays 2-layer per 2505.12514).
- Empirical check prompted by the brief: trained node embeddings (run-1 best.pt) have mean |cos| 0.063, max 0.50 — projection bleed is real but bounded; frozen Gate A-ii is the behavioral catch, ridge arm to be reported alongside.
- Two NEW unverified references from the brief logged in litscan/NOTES.md; arXiv:2606.01243 is potentially our nearest neighbor — verify before any writeup.
- Repo pushed to github.com/star2vec/metis; Windows/CUDA instance takes over training per WINDOWS_SETUP.md (run s1_seed0_win, v2 recipe).
- RACE CANCELLED: Mac v2 run stopped at user request (machine needed for other work) at epoch 38, mid-stage-1, stage_val_acc 50.2% — for the record, that is 2x run 1's entire stage-1 best (24.5%) with 11 epochs still to go in the stage, further confirming the #004 diagnosis. Partial v2 artifacts retained in results/s1_seed0_v2/. s1_seed0_win on the Windows/CUDA machine is now the SOLE primary-substrate run. brief.txt (external planning brief, triaged above) added to the repo root.
- Run 1 (s1_seed0): precondition MISS — test acc 60.9% vs ≥93%; full-task val plateaued at 67.3%. Diagnosis in DECISIONS #004: BFS-wave readout broken exactly at steps 2–3, the depths that patience-3-truncated stages 1–2 teach. Artifacts in results/s1_seed0/.
- v2 retrain (s1_seed0_v2, same seed, fixed 25-epoch curriculum stages per #004) live on the M1 and validating the diagnosis: stage 0 hit 77–79% (run 1 cut it at 44%, which turned out to be a pre-transition grind — the jump came at epochs 10–20); stage 1 passed run 1's all-time best within 7 epochs. Epoch-0 metrics reproduced run 1 exactly (determinism confirmed; only the schedule differs).
- Migration: user offers an i9/RTX-2000-Ada/32GB Windows laptop (~5–10× faster for this workload; latency-bound, so H100 assessed NOT worth it until/unless Huginn-scale transfer). Plan: push to a private GitHub remote, fresh Claude instance on the laptop runs the same recipe as s1_seed0_win (WINDOWS_SETUP.md), Mac v2 keeps racing — first substrate to pass acceptance is THE primary substrate; the other becomes a cross-platform robustness point. Equivalence test is a REQUIRED gate before any training on the new machine.
- Vendor training stack (Ber666/reasoning-by-superposition, run.py + coconut.py) is
hard-wired to 2-GPU NCCL/FSDP and pathologically slow in three independent ways:
(1) Coconut.forward builds latent index lists by per-element
.item()on a device tensor (~65k forced syncs/step) and rebuilds inputs_embeds through ~100k Python tensor ops per step — this, not model size, explains the paper's ~24 h × 2×A100 wall-time for a 15M-param model; (2) PreTrainedTokenizer.encode re-runs added-token machinery on 15k samples every epoch (~5 min/epoch); (3) on MPS, batch 128 stalls the allocator (26 s/step vs 2.3 s/step at batch 64). - Fix: src/metis/fast_coconut.py (FastCoconut, FastSTokenizer) — same math, vectorized implementation. Equivalence proven bit-exact by tests/test_fast_equivalence.py: identical logits (max diff 0.0), loss, recycled thought embeddings, and all parameter gradients vs vendor code on real stage-4 data. Not a DECISIONS entry: implementation, not experimental design; recipe (optimizer, curriculum, batching math) unchanged. Batch realized as 64 × grad-accum 4 = effective 256 = paper's 128 × 2 ranks.
- Driver fix before launch (pre-first-run, no data seen): per-stage early stopping now gates on STAGE-MATCHED val accuracy (predict a depth-k valid-path node, k = stage+1; neighbor_k membership), not full-task accuracy, which is ~0 during early stages regardless of learning and would have truncated every stage at patience epochs. From stage 4 on, stage accuracy IS full-task accuracy and selects best.pt.
- Data facts pinned while porting: neighbor_k[k] = depth-k nodes on valid root→target paths (strict subset of BFS frontier, 99/100 test graphs) = the paper's "Optimal" bucket; supervision samples a random valid-path node per epoch. Splits match the paper's Table 4 exactly (14,785 / 257 / 419).
- Measured (M1 MPS, fp32, memory freed by user): stage-4 epoch ≈ 9 min, val-gen 257 graphs ≈ 15 s. Launched primary seed (seed 0), 36 h cap, resumable; metrics stream to results/s1_seed0/metrics.jsonl; acceptance next per §2 precondition.
- User-guided walkthrough of DRAFT v0.1 §3, one bar at a time. Two internal-consistency bugs found in the draft (Gate A's |ΔS| clause unpassable by probability conservation; Gate B's +20-pt clause ceiling-locked) plus spec gaps (undefined basis-fallback sanity, unpinned target-branch/T definitions, readout-circular step selection). All revisions user-approved → DECISIONS #001.
- Both litscan house-rule debts paid (litscan/NOTES.md, "pre-freeze verification" section): 2505.12514 verified at figure level from the actual PDF; Ulterior Motives (2604.23460) read in full — detection-only, niche confirmed open.
- Verification caught two approved-bar invalidities BEFORE freeze: the task is binary forced choice with an unreachable decoy ⇒ 50% chance floor (the 50-pt drop bar exceeded the ~47-pt theoretical ceiling) and Gate B as approved was vacuous. Reformulated with user approval → DECISIONS #002. Compute reality also surfaced: paper runs ≈ 24 h on 2×A100-80GB (its C.3); early-stopping strategy approved.
- PREREGISTRATION.md rewritten and flipped to v1.0 FROZEN — the commit carrying this entry is the freeze commit. §2 + §3 frozen; changes hereafter = DECISIONS entries, post-hoc.
- Next: clone Ber666/reasoning-by-superposition (+ facebookresearch/coconut as reference), set up MPS env, train S1 primary seed (early stopping, ~36 h cap), substrate acceptance = test acc ≥ 93% + BFS-wave readout replication.
- Scaffolded from the mycelia session that generated and lit-scanned the idea. PREREGISTRATION.md is DRAFT v0.1 — nothing frozen, nothing run, no code exists.
- Lit scan first pass done same day (litscan/NOTES.md): niche NARROW BUT OPEN — theory and correlational evidence for thought-superposition exist (2505.12514 et al.); per-branch causal editing + capacity measurement absent; SPAR cohort active nearby, so speed matters. Two verification debts before freeze: 2505.12514's actual figures; Ulterior Motives (2604.23460) in full.
- Next, in order: (1) user reviews/edits the draft bars → freeze commit (v1.0); (2) clone facebookresearch/coconut, reproduce S1 substrate + sanity-check the inner-product readout replicates; (3) run the kill-test.
- Venue posture fixed in CLAUDE.md rule 6: AAAI-27 abstract (2026-07-21) may be registered as a free option; arXiv-then-workshop is the plan; bars never move for deadlines.