Repository navigation
docs: context-closed-tasks experiment — brainstorms v2/v3, pilot protocol, run data, RESULTS (STRONG) - #124
Closed
freezscholte wants to merge 13 commits into
Closed
freezscholte wants to merge 13 commits into
freezscholte wants to merge 13 commits into
Conversation
…ate research Adds the v2 context-closed-tasks brainstorm (contracts as enforced substrate, pilot protocol) and the 2026-07-06 research delta: Shepherd axis analysis, CooperBench fit assessment, code-grounded substrate map, origin lineage, and the agreed spike protocol (T1-T3). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…enda T1 deliverables: A4.1-schema contracts for forge-policy and forge-evidence, byte-stable brief.sh, results log (brief ~2k tokens vs ~4k standing-instruction baseline; authoring-cost lower bound; three schema-pressure observations). Brainstorm doc gains §7.1 agreed sequencing, §7.2 measurement addendum, §7.3 byte-stable-brief/KV-cache note; §8 decisions resolved. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…perBench feasible T2: blast-check.py predicate over changed_paths in save/propose payloads; pass + violation cases verified in a forge-dogfood temp clone; U3 resolved (competing attempts isolated, attempt compare surfaces per-attempt paths); friction F1 logged (pre-attach workspace edits silently discarded). T3: CooperBench recon — --no-messaging is first-class, briefs injectable via shadow dataset dir with zero code; Rust subset is one homogeneous typst PR; smallest pilot-3 is ~10 pairs x 3 arms (~$30-75). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
… draft decomposition Scope decided 2026-07-06 (Jan): intent-aware blame as the decomposition subject + the attach drift-guard bug. Refactor slices eliminated after verifying NER-366/NER-381 are fully done. Draft 8-task decomposition, pinned Arm A/B/C definitions, measurement + rubric/inter-rater rules, order of work. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
Contracts for NER-362 (5 tasks) and NER-382 (3 tasks) per PILOT.md §2, plus the frozen rubric (defect classes, unlicensed-decision test with calibration examples, unknown scoring, mechanical metrics). Per protocol: frozen before any implementation; revisions from here are data. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…fs, randomized order Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…tract rev 2 Run data: arm A batch 1 (negative control, 8/8 unknown-stops with no contract) archived as runs-invalid-01-nobrief; batch 2 = 3 impl + 5 typed unknown-stops incl. the rev-1 382-2 contract contradiction (diff_working_vs_tree status-cache side effect vs read-only invariant). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…ay apply Arm A final: 8 implementations, 7/8 zero blast violations; the one violation (382-2 touching the store facade for re-export wiring) is provisionally a too-narrow-contract defect, to be settled in scoring. Contract YAML is not lint-clean (unquoted colons) — first earned datum for the deferred contract-lint battery. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
Blinded scoring (8 tasks x 2 arms): Arm A 5 defects (all contract-class, zero implementation) + 3 unlicensed vs Arm B 31 defects + 40 unlicensed; A 11/11 gates and 8/8 tasks-as-specified vs B 10/11 and 7/8 (silent task substitution at B-362-5). H-CORE 100% (threshold 70%), 1 contract revision across 8 tasks. Decision rule -> thin harness + pilot 2, gated on Jan's inter-rater pass. Full caveats in RESULTS.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…rer 1 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
…tatus updates Lightweight ce-compound x3: shared-exclusion-contract walker rule (architecture-patterns), gate layering acceptance!=merge-ready (conventions), stop-on-unknown gate (design-patterns). Plus the v3 context-closed-tasks brainstorm (post-pilot design state, earned schema commitments, updated roadmap) and unknown-status updates in the landscape doc. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The complete record of the context-closed-tasks experiment (2026-07-06): can agent context be managed as contract revisions instead of transferred history?
experiments/ccx/): 8 frozen task contracts, scoring rubric, byte-stablebrief.sh, blast-radius predicate, run harness, all run data (including the accidental negative control), two blinded scorers' verdicts, RESULTS.md.Headline result (pre-registered STRONG reading met)
8 real tasks (NER-362 + NER-382), two arms, blinded scoring: brief-only fresh sessions — 0 implementation defects, 3 unlicensed decisions, 8/8 delivered-as-specified; status-quo continuous session — 21+ implementation defects, 40 unlicensed decisions, 1 silent task substitution. Two scorers, 8/8 directional agreement. Every brief-side defect was in the authored contracts — risk concentrated into the decomposer exactly as predicted. Full validity caveats inside RESULTS.md (read them before quoting the headline).
Side effects already shipped: PR #123 (NER-382 drift guard, found by this experiment, built with it, hardened by the review gate). Follow-ups: NER-383, thin harness, pilot 2 (neighbor ablation), pilot 3 (CooperBench — feasibility note included).
Verification
Docs-only + experiment artifacts (no
crates/changes).scripts/ci.shunaffected; frontmatter of new solution docs validated.🤖 Generated with Claude Code
https://claude.ai/code/session_01GTeDdEu9Q4a96DdXK4tfju
Need help on this PR? Tag
/codesmithwith what you need. Autofix is disabled.