Skip to content

docs: context-closed-tasks experiment — brainstorms v2/v3, pilot protocol, run data, RESULTS (STRONG) - #124

Closed
freezscholte wants to merge 13 commits into
mainfrom
experiment/ccx-spikes
Closed

freezscholte wants to merge 13 commits into
mainfrom
experiment/ccx-spikes

Conversation

@freezscholte

@freezscholte freezscholte commented Jul 6, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

The complete record of the context-closed-tasks experiment (2026-07-06): can agent context be managed as contract revisions instead of transferred history?

  • Brainstorms: v2 (design + pre-registered pilot protocol), landscape/substrate research (Shepherd, CooperBench, code-level substrate map, origin lineage), v3 (post-pilot design state with earned schema commitments).
  • Pilot artifacts (experiments/ccx/): 8 frozen task contracts, scoring rubric, byte-stable brief.sh, blast-radius predicate, run harness, all run data (including the accidental negative control), two blinded scorers' verdicts, RESULTS.md.
  • Solution docs: shared-exclusion-contract walker rule, gate layering (acceptance ≠ merge-ready), stop-on-unknown gate.

Headline result (pre-registered STRONG reading met)

8 real tasks (NER-362 + NER-382), two arms, blinded scoring: brief-only fresh sessions — 0 implementation defects, 3 unlicensed decisions, 8/8 delivered-as-specified; status-quo continuous session — 21+ implementation defects, 40 unlicensed decisions, 1 silent task substitution. Two scorers, 8/8 directional agreement. Every brief-side defect was in the authored contracts — risk concentrated into the decomposer exactly as predicted. Full validity caveats inside RESULTS.md (read them before quoting the headline).

Side effects already shipped: PR #123 (NER-382 drift guard, found by this experiment, built with it, hardened by the review gate). Follow-ups: NER-383, thin harness, pilot 2 (neighbor ablation), pilot 3 (CooperBench — feasibility note included).

Verification

Docs-only + experiment artifacts (no crates/ changes). scripts/ci.sh unaffected; frontmatter of new solution docs validated.

🤖 Generated with Claude Code

https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju


View with Codesmith Autofix with Codesmith
Need help on this PR? Tag /codesmith with what you need. Autofix is disabled.

freezscholte and others added 13 commits July 6, 2026 09:51
…ate research

Adds the v2 context-closed-tasks brainstorm (contracts as enforced
substrate, pilot protocol) and the 2026-07-06 research delta: Shepherd
axis analysis, CooperBench fit assessment, code-grounded substrate map,
origin lineage, and the agreed spike protocol (T1-T3).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…enda

T1 deliverables: A4.1-schema contracts for forge-policy and
forge-evidence, byte-stable brief.sh, results log (brief ~2k tokens vs
~4k standing-instruction baseline; authoring-cost lower bound; three
schema-pressure observations). Brainstorm doc gains §7.1 agreed
sequencing, §7.2 measurement addendum, §7.3 byte-stable-brief/KV-cache
note; §8 decisions resolved.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…perBench feasible

T2: blast-check.py predicate over changed_paths in save/propose payloads;
pass + violation cases verified in a forge-dogfood temp clone; U3 resolved
(competing attempts isolated, attempt compare surfaces per-attempt paths);
friction F1 logged (pre-attach workspace edits silently discarded).
T3: CooperBench recon — --no-messaging is first-class, briefs injectable
via shadow dataset dir with zero code; Rust subset is one homogeneous
typst PR; smallest pilot-3 is ~10 pairs x 3 arms (~$30-75).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
… draft decomposition

Scope decided 2026-07-06 (Jan): intent-aware blame as the decomposition
subject + the attach drift-guard bug. Refactor slices eliminated after
verifying NER-366/NER-381 are fully done. Draft 8-task decomposition,
pinned Arm A/B/C definitions, measurement + rubric/inter-rater rules,
order of work.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
Contracts for NER-362 (5 tasks) and NER-382 (3 tasks) per PILOT.md §2,
plus the frozen rubric (defect classes, unlicensed-decision test with
calibration examples, unknown scoring, mechanical metrics). Per protocol:
frozen before any implementation; revisions from here are data.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…fs, randomized order

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…tract rev 2

Run data: arm A batch 1 (negative control, 8/8 unknown-stops with no
contract) archived as runs-invalid-01-nobrief; batch 2 = 3 impl + 5
typed unknown-stops incl. the rev-1 382-2 contract contradiction
(diff_working_vs_tree status-cache side effect vs read-only invariant).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…ay apply

Arm A final: 8 implementations, 7/8 zero blast violations; the one
violation (382-2 touching the store facade for re-export wiring) is
provisionally a too-narrow-contract defect, to be settled in scoring.
Contract YAML is not lint-clean (unquoted colons) — first earned datum
for the deferred contract-lint battery.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
Blinded scoring (8 tasks x 2 arms): Arm A 5 defects (all contract-class,
zero implementation) + 3 unlicensed vs Arm B 31 defects + 40 unlicensed;
A 11/11 gates and 8/8 tasks-as-specified vs B 10/11 and 7/8 (silent task
substitution at B-362-5). H-CORE 100% (threshold 70%), 1 contract
revision across 8 tasks. Decision rule -> thin harness + pilot 2,
gated on Jan's inter-rater pass. Full caveats in RESULTS.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…tatus updates

Lightweight ce-compound x3: shared-exclusion-contract walker rule
(architecture-patterns), gate layering acceptance!=merge-ready
(conventions), stop-on-unknown gate (design-patterns). Plus the v3
context-closed-tasks brainstorm (post-pilot design state, earned schema
commitments, updated roadmap) and unknown-status updates in the
landscape doc.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
@freezscholte
freezscholte deleted the experiment/ccx-spikes branch July 6, 2026 19:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant