diff --git a/docs/BUILD_LOG.md b/docs/BUILD_LOG.md index 2110c6c..69d22f8 100644 --- a/docs/BUILD_LOG.md +++ b/docs/BUILD_LOG.md @@ -350,3 +350,36 @@ Evidence: `docs/v0.3/results/R7_DEVELOPMENT_AUDIT.md` and immutable JSON receipt - Generated all 18 blind rating packets before outcomes, with index `54d78382b3ddbe15cba1f8153275e8149d32ddaa5192163f99ca5f43d903e8fe`, and verified zero remaining audit containers/volumes. Evidence: `docs/v0.3/results/R7_HELD_OUT_AUDIT.md`, immutable execution receipts, and blind packet bundle. Full R7 remains blocked on two independent experienced TypeScript ratings and adjudication; R5/R6 have not started. + +## 2026-08-01 — R&D integration and Executable Operator Model proposal + +- Merged protected PR #18 into `codex/shadow-cockpit-rnd` after all five GitHub checks passed, then deleted the short-lived R7 head locally and from origin. +- Retargeted the stale Jules infrastructure PR #8 away from `main` and into R&D, updated it from the current base, passed all five checks, merged it, and deleted the short-lived remote head. `main` remained untouched; origin now has only `main` and the single persistent coding branch `codex/shadow-cockpit-rnd`. +- Audited the original product goal against current runtime evidence. The live agent plane, v0.3 cockpit, cold relay, delayed-transfer study, and skill-preservation claim remain incomplete; the automatic compiler gate is not substituted for product completion. +- Proposed ADR-007: every meaningful agent checkpoint can produce tested software, an executable `observe → actuate → recover` surface, and a local Executable Operator Model containing only behaviorally supported causal claims. +- Added the draft post-R7 specification with deterministic invalidation, shadow-control, context-starved relay, control-dividend, privacy, isolation, and ablation requirements. +- Ran the spec-driven strict validator after correcting its required heading and traceability format: 98/100, no errors. Its only warning expects an HTTP method/path; the spec explicitly keeps this boundary local and injected in the extension host instead of inventing a network API. +- Updated the PRD and v0.3 index so attention is selected from operator-model divergence and uncovered recovery routes rather than random functions, timer prompts, or question counts. + +Evidence: merged PRs [#18](https://github.com/yava-code/PureFlow/pull/18) and [#8](https://github.com/yava-code/PureFlow/pull/8); `docs/v0.3/ADR-007-EXECUTABLE-OPERATOR-MODEL.md`; `docs/v0.3/OPERATOR_MODEL_SPEC.md`; `docs/v0.3/GOAL_COMPLETION_AUDIT.md`. ADR-007 remains proposed and implementation remains gated by two independent R7 expert ratings. + +## 2026-08-01 — Evidence-Carrying Generation proposal + +- Closed the remaining architecture gap between critical-seam takeover evidence and accountability for the rest of a large generated change. +- Proposed ADR-008: a bidirectional Intent Ledger reconciles every changed line to a bounded semantic unit and labels its attribution `supported`, `claimed`, `unattributed`, `contradicted`, or `stale`. +- Kept agent-emitted intent references untrusted, required independent structural/evidence joins, and prohibited ledger coverage from updating human readiness. +- Added deterministic handling for generated artifacts, total changed-line reconciliation, negative trace-washing rules, a compact change-account UX, and a held-out ablation against raw diffs, AI summaries, and AST navigation. +- Distinguished the proposal from domain DSLs, requirements traceability, proof-carrying code, and generated explanations. No product-effectiveness or skill-retention result is claimed. + +Evidence: `docs/v0.3/ADR-008-EVIDENCE-CARRYING-GENERATION.md`; updated PRD, thesis, project state, completion audit, and v0.3 index. Implementation remains gated by the complete two-rater R7 expert audit. + +## 2026-08-01 — Decision Futures and Takeover Envelope proposal + +- Closed the remaining architecture gap between off-path takeover practice and real live engineering authority. +- Proposed ADR-009: agents speculatively implement viable alternatives and continue unrelated work while a bounded, precommitted Decision Future offers one high-leverage human choice at a low-cost breakpoint. +- Distinguished on-time live influence from late counterfactual practice, autonomous default, skip, and integrity failure. A late agreement can never be upgraded into production-decision evidence. +- Added false-fork abstention, immutable deadline/default authority, comparable-alternative rules, interruption-cost scheduling, and executable integration evidence. +- Defined separate Autonomy and Takeover Envelopes so full agentic coding remains available while uncovered or stale human-control seams stay explicit. +- Added a three-arm ablation against autonomous explanation and shadow-only prediction. No product-effectiveness, live-influence, or skill-retention result is claimed. + +Evidence: `docs/v0.3/ADR-009-DECISION-FUTURES.md`; updated PRD, thesis, research, project state, completion audit, and v0.3 index. Implementation remains gated by the complete two-rater R7 expert audit. diff --git a/docs/PROJECT_STATE.md b/docs/PROJECT_STATE.md index 3a34cb1..411d6d9 100644 --- a/docs/PROJECT_STATE.md +++ b/docs/PROJECT_STATE.md @@ -2,7 +2,7 @@ Last updated: 2026-08-01 -## Current branch milestone — R7 corpus frozen; compiler audit pending +## Current branch milestone — R7 automatic audit passed; expert gate pending Branch `codex/shadow-cockpit-rnd` resets the product R&D thesis around **Dual-Control Development**. @@ -34,15 +34,17 @@ Branch `codex/shadow-cockpit-rnd` resets the product R&D thesis around **Dual-Co - Protected PR #17 run `30674334938` passed `extension`, `extension-windows`, `contract`, `web`, and `jules-rnd-policy`. The Linux extension job explicitly provisioned the exact digest and passed the real Docker backend suite; Windows independently passed the deterministic contract suite. The R7 sandbox implementation gate is complete. - `CONCEPT_LAB_CONTROLLABILITY.md` records a post-R7 category extension: compile an executable `observe → actuate → recover` human control surface, select takeover cut sets, and let a context-starved agent continue writing code from human-selected evidence and directives. Dissent cases and control dividends remain hypotheses with explicit falsifiers, not implemented features. - The preregistered R7 collector froze 30 eligible patches from six repositories after evaluating 457 bounded eligibility records. Manifest `a4ef6cbfa48c66cb9d384bcc2834ecbfae8ff08810abfd1863b395b8aa47d149` contains 12 development and 18 held-out patches; `docs/v0.3/results/R7_CORPUS_COLLECTION.md` reports repository and first-match exclusion counts. No compiler or human outcome influenced selection. -- R7 has not passed. The recovery/probe compiler audit, three-run evidence, protected parity/adversarial runs, and two independent human ratings remain pending; R5/R6 stay gated. +- The frozen R7 automatic audit passed its preregistered automatic threshold: 17/18 held-out identities compiled and 16/18 were valid end-to-end. The frozen blind expert packet set and deterministic rating join exist, but two independent ratings and adjudication remain pending. Full R7 has not passed; R5/R6 stay gated. +- ADR-007 proposes an Executable Operator Model and shadow-control protocol. ADR-008 adds a bidirectional Intent Ledger for artifact accountability. ADR-009 adds Decision Futures and a Takeover Envelope so an on-time pre-reveal human commitment can determine a live integrated branch while agents retain implementation. Together they cover artifact accountability, demonstrated control, and real decision authority; none is implementation evidence. +- `GOAL_COMPLETION_AUDIT.md` maps the original product goal to current evidence. It explicitly records that the live agent plane, v0.3 cockpit, cold relay, delayed-transfer study, and skill-preservation claim remain incomplete. - The readiness ledger and v0.3 cockpit do not exist yet. R0–R4.5 remain a closed reviewed-fixture mechanism and do not execute arbitrary participant or workspace code. - No skill-retention or speed metric has been measured. Values in the PRD are predeclared R&D targets. - A new implementation audit found five R0 ambiguities: candidate-diff identity, pre-store fixture blobs, runtime identity, check IDs, and Git object format. The normative contract closes them with structured diffs, catalog-owned blobs, standalone Node `v22.17.0`, declared test IDs, and SHA-1 Git initialization; R0a/R0b now implement and verify that complete substrate. -- A guarded Jules dispatcher and PR policy are defined as a finite R0→R4 queue. They create at most one session after a successful preflight, stop after merged R4, remain inert unless dispatch is explicitly enabled, and keep plan approval on by default. Merges remain manual because the current project tests are not an independent immutable verifier. Full scheduled continuation still requires the dispatcher workflow to be reviewed into the default branch. -- The R&D branch is published at `origin/codex/shadow-cockpit-rnd`. Its first Jules workflow run was correctly skipped because `JULES_RND_LOOP_ENABLED` is not enabled; no Jules session was created. +- A guarded Jules dispatcher and PR policy are defined as a finite R0→R4 queue. They create at most one session after a successful preflight, stop after merged R4, remain inert unless dispatch is explicitly enabled, and keep plan approval on by default. Merges remain manual because the current project tests are not an independent immutable verifier. The workflow source is now present in R&D, but GitHub schedules and manual dispatch require the workflow file on the default branch; no automatic Jules loop is active. +- The R&D branch is published at `origin/codex/shadow-cockpit-rnd`. No Jules session was created by the guarded workflow. - `Protect main` is active: PR, conversation resolution, strict `extension`/`contract`/`web` checks, up-to-date base, deletion protection, and force-push protection are enforced with zero required approvals for the sole owner. - `Protect R&D integration` is active on exact branch `codex/shadow-cockpit-rnd`: PR-only updates, conversation resolution, strict `extension`, `extension-windows`, `contract`, `web`, and `jules-rnd-policy` GitHub Actions checks, up-to-date base, deletion protection, and force-push protection. Codex and Jules now integrate through short-lived heads. -- Superseded `codex/v0.2-ownership-compiler` and the obsolete Jules vibe-gate branch were preserved as dated archive tags and deleted as branches. `codex/shadow-cockpit-rnd` is the only persistent coding branch; draft PR [#8](https://github.com/yava-code/PureFlow/pull/8) temporarily retains the default-branch scheduler until the owner explicitly authorizes its merge into `main`. +- Superseded `codex/v0.2-ownership-compiler` and the obsolete Jules vibe-gate branch were preserved as dated archive tags and deleted as branches. `codex/shadow-cockpit-rnd` is the only persistent coding branch. PR [#8](https://github.com/yava-code/PureFlow/pull/8) was retargeted from `main` to R&D, merged, and its short-lived head was deleted; `main` remains untouched. The v0.1 runtime below remains released evidence and a reusable IDE shell. Its Mentor, Quiz, and Focus behavior is not the v0.3 product core. @@ -135,7 +137,7 @@ The repository contains no verified evidence that the owner submitted the final | The first live adapter is selected but no accessible Codex CLI is configured for this checkout | ADR-006 selects Codex App Server over local stdio, but the Microsoft Store packaged executable discovered here returns `Access denied` when launched from the repository shell | Keep replay R&D independent; the live spike must preflight a separately accessible, exact-version user-installed Codex CLI and fail closed when unavailable | | R7 expert audit is not complete | Frozen held-out automatic audit passed at 16/18 end-to-end, but independent causal-relevance ratings are not yet measured | Give `docs/v0.3/results/held-out-rater-packets/` to two experienced TypeScript raters using `docs/v0.3/R7_EXPERT_RATING.md`, then adjudicate and report agreement | | Human participants are not recruited | Takeover and delayed-transfer claims cannot be tested | Complete the technical gate, then recruit for the preregistered pilot | -| Default-branch Jules scheduler awaits explicit merge approval | Scheduled/manual continuation is not installed on `main`; draft PR #8 remains isolated and the enable variable stays off | Owner explicitly says `merge #8`; then merge through protected `main`, remove the temporary infrastructure branch, and run one guarded canary through the protected R&D branch | +| Default-branch Jules dispatcher awaits explicit authorization | The workflow exists on R&D, but GitHub will not schedule or manually dispatch it until a workflow file exists on `main`; the enable variable stays off | Create a new minimal workflow-only PR into protected `main` only after the owner explicitly authorizes that exact action, then run one guarded canary targeting R&D | No external input blocks the repository-owned fixture R0–R4.5 mechanism in `docs/v0.3/AGENT_EXECUTION.md`. It cannot execute or judge arbitrary human/agent code and must not be represented as a general-project takeover product. @@ -153,8 +155,9 @@ No external input blocks the repository-owned fixture R0–R4.5 mechanism in `do ## Next ordered actions 1. Obtain and adjudicate two independent blind expert ratings for packet index `54d78382b3ddbe15cba1f8153275e8149d32ddaa5192163f99ca5f43d903e8fe`. -2. Add the local readiness ledger and minimal cockpit only after the complete R7 expert gate passes. -3. Run the preregistered delayed-transfer pilot before making any skill-retention claim. +2. If full R7 passes, review ADR-007, ADR-008, and ADR-009 together, then freeze the Operator Model, Intent Ledger, Decision Future, and Takeover Envelope schemas, hash domains, projections, evidence authorities, timing policy, and selector ablations before implementation. +3. Add the local readiness ledger and minimal cockpit only after the complete R7 expert gate passes. +4. Run the preregistered delayed-transfer pilot before making any skill-retention claim. ## Recent milestone commits diff --git a/docs/v0.3/ADR-001-DUAL-CONTROL.md b/docs/v0.3/ADR-001-DUAL-CONTROL.md index 6f668cf..a1849ef 100644 --- a/docs/v0.3/ADR-001-DUAL-CONTROL.md +++ b/docs/v0.3/ADR-001-DUAL-CONTROL.md @@ -141,6 +141,8 @@ The router accepts the user's attention budget and readiness evidence. In v0.3 i After recovery episodes validate the mechanism, the next router experiment is speculative live steering: fork two viable designs, collect comparable executable evidence, and let the human's choice determine the integrated branch without requiring them to type the implementation. +ADR-009 formalizes that later experiment as a Decision Future with a precommitted autonomous default, bounded integration deadline, false-fork abstention, live-influence evidence, and a Takeover Envelope. It remains gated and does not change the implemented R0–R4.5 path. + ## Options considered ### Option 1: Extend the v0.1 Mentor and Focus Rep diff --git a/docs/v0.3/ADR-007-EXECUTABLE-OPERATOR-MODEL.md b/docs/v0.3/ADR-007-EXECUTABLE-OPERATOR-MODEL.md new file mode 100644 index 0000000..ada28d3 --- /dev/null +++ b/docs/v0.3/ADR-007-EXECUTABLE-OPERATOR-MODEL.md @@ -0,0 +1,272 @@ +# ADR-007 — Executable Operator Model and shadow control + +- **Status:** Proposed; implementation gated by the complete R7 expert audit +- **Date:** 2026-08-01 +- **Deciders:** repository owner after R7 adjudication +- **Supersedes:** no earlier ADR; extends ADR-001 and ADR-004 + +## Context + +PureFlow's current Dual-Control thesis correctly rejects passive diff review, explanations, quizzes, and manual-code quotas as evidence of takeover readiness. The remaining weakness is architectural: a Takeover Twin is still described primarily as an experience produced beside the real system. The product has no durable representation of what the operator can actually predict, observe, change, and recover. + +That gap matters because the requested end state is not “the developer completed exercises.” It is: + +> autonomous agents may write essentially all implementation code, while the developer retains an accurate, current, project-specific model and can direct or perform a takeover when automation disappears. + +Research does not support the analogy that passive autopilot use preserves manual ability. Automation can improve routine performance while leaving operators with worse situation awareness and recovery performance. Simple explanations do not reliably prevent overreliance. Cognitive forcing can reduce overreliance, but its benefit costs attention and acceptability. The 2026 Anthropic coding study found lower immediate mastery under AI assistance on average, with the largest reported gap in debugging; interaction modes that used AI to build comprehension did not show the same qualitative pattern. + +The design implication is not to remove autonomy. It is to make operator readiness a maintained system artifact with executable evidence, drift, and invalidation semantics. + +## Decision + +PureFlow will treat each meaningful autonomous checkpoint as capable of producing three outputs: + +1. a tested software checkpoint; +2. an immutable control surface describing how to `observe → actuate → recover` one critical seam; +3. an **Executable Operator Model** containing only claims the developer has behaviorally demonstrated or explicitly left unverified. + +The Operator Model is not an AI summary of the repository. It is a local, project-scoped graph of bounded causal claims connected to executable observations and prior human actions. + +```text + ┌────────────────────────────┐ +developer intent ───────►│ autonomous build plane │────► software checkpoint N+1 + └─────────────┬──────────────┘ + │ observable evidence + ▼ + ┌────────────────────────────┐ + │ controllability compiler │ + │ observe / actuate / recover│ + └─────────────┬──────────────┘ + │ model delta + ▼ + ┌────────────────────────────┐ + │ operator-model compiler │ + │ fresh / stale / contradicted│ + └─────────────┬──────────────┘ + │ highest information gap + ▼ + ┌────────────────────────────┐ + │ shadow control protocol │ + │ predict / evidence / direct│ + └─────────────┬──────────────┘ + │ executable result + ▼ + test / probe / rollback / handle +``` + +### Shadow control + +At a stable checkpoint, the production agent continues its next task. PureFlow may open one bounded shadow-control opportunity for checkpoint N: + +1. commit a concrete prediction before the relevant result or finished implementation is revealed; +2. select the observation that would distinguish plausible causes; +3. optionally give a hypothesis and conditional directive to a context-starved agent; +4. let that agent write every line of the intervention; +5. execute the intervention in a disposable twin; +6. recover after a wrong first move; +7. convert the verified result into a durable control artifact. + +The human owns the causal model and control policy. The agent may still own implementation, syntax, command execution, and repair mechanics. + +### Model divergence, not question frequency + +PureFlow schedules attention from divergence between the current checkpoint and the operator model. A changed function is not prompt-worthy merely because it exists. A seam becomes eligible when one or more of these are true: + +- a previously demonstrated claim became stale after a meaningful change; +- a new critical route has no human-controlled observation or recovery evidence; +- the agent changed an assumption, plan, or invariant after new evidence; +- two candidate implementations predict different failure envelopes; +- a prior human prediction was contradicted; +- delayed transfer evidence is absent or decayed. + +Random function questions, line counts, and periodic timer prompts cannot update the Operator Model. + +### Every interaction pays a control dividend + +An interaction that produces only a score or explanation is incomplete. A passed shadow-control episode should create or verify at least one reusable project artifact: + +- regression probe; +- deterministic observation recipe; +- rollback recipe; +- bounded incident packet; +- control-surface handle; +- clarified invariant attached to executable evidence. + +This makes the developer's attention part of software production rather than homework beside it. + +## Proposed contracts + +These contracts are intentionally non-normative until the ADR is accepted. They must receive domain-separated RFC 8785 hashing, strict projections, and controller-owned authority before implementation. + +```ts +type OperatorClaimState = + | "unverified" + | "fresh" + | "stale" + | "contradicted"; + +interface OperatorClaim { + schemaVersion: 1; + id: string; + projectId: string; + sourceTreeHash: string; + seamId: string; + controlSurfaceHash: string; + behaviorClaim: string; + observationIds: string[]; + recoveryId: string; + supportingCapabilityEvidenceHashes: string[]; + state: OperatorClaimState; + checkedAt?: string; + claimHash: string; +} + +interface OperatorModelSnapshot { + schemaVersion: 1; + projectId: string; + sourceTreeHash: string; + claims: OperatorClaim[]; + uncoveredSeamIds: string[]; + snapshotHash: string; +} + +interface ModelDelta { + schemaVersion: 1; + projectId: string; + beforeSnapshotHash: string; + sourceTreeHash: string; + freshClaimIds: string[]; + staleClaimIds: string[]; + contradictedClaimIds: string[]; + uncoveredSeamIds: string[]; + reasonCodes: string[]; + deltaHash: string; +} + +interface ShadowControlSession { + schemaVersion: 1; + id: string; + projectId: string; + sourceTreeHash: string; + modelDeltaHash: string; + controlSurfaceHash: string; + participantExperienceHash: string; + availableEvidenceIds: string[]; + coldExecutorProfileId?: string; + attentionBudgetSeconds: number; + judgeSpecHash: string; + sessionHash: string; +} + +type ShadowControlEvent = + | { type: "shadow.started"; sessionId: string; at: string } + | { type: "shadow.prediction.committed"; sessionId: string; attemptHash: string; at: string } + | { type: "shadow.evidence.selected"; sessionId: string; evidenceIds: string[]; at: string } + | { type: "shadow.directive.committed"; sessionId: string; directiveRef: string; at: string } + | { type: "shadow.patch.proposed"; sessionId: string; candidateDiffHash: string; at: string } + | { type: "shadow.judged"; sessionId: string; judgeResultHash: string; outcome: "passed" | "partial" | "failed" | "failed-integrity"; at: string } + | { type: "shadow.dividend.created"; sessionId: string; artifactHash: string; at: string }; +``` + +Free-form developer or model output must never become a command, path, environment value, mount, test, or judge. A small model may parse prose into an existing bounded choice or coach after the commitment. Only executable controller-owned evidence changes claim state. + +## Autonomy routing + +The router chooses among four policies for each eligible checkpoint: + +| Policy | When selected | Production behavior | +| --- | --- | --- | +| Silent autonomy | model coverage is fresh and risk is low | agents continue; no prompt | +| Shadow prediction | one high-information observation can test the model | agents continue; result is revealed after commitment | +| Cold relay | diagnosis and intervention ownership matter | fresh agent writes code from human-selected evidence and directive | +| Blackout replay | recovery readiness is weak or contradicted | human inherits a real pre-repair checkpoint in the twin | + +The router optimizes expected takeover information per active minute, not engagement. A `0` minute budget remains valid and produces no readiness claim. + +## User experience + +The native editor remains primary. The cockpit shows no course or generic score. Its central object is a live map of control coverage: + +```text +Agent checkpoint 42 passed + +Auth route: fresh control evidence +Tenant cache: model contradicted by new invalidation path +Billing webhook: no recovery handle + +Best 3-minute control opportunity: +Which observation separates stale-cache propagation from key collision? + +[Run shadow control] [Let agents decide] [Budget: 5 min] +``` + +Inside shadow control, the side chat behaves like an adversarial senior engineer. It asks for a prediction, an observation, or a conditional next move. It does not ask for prose about arbitrary code and does not require the user to manually type the patch. + +## Options considered + +### A. Diff and explanation review + +Low implementation cost, high familiarity, but poor scaling and no behavioral evidence of prediction, intervention, or recovery. + +### B. Random questions, streaks, and manual-code quotas + +Easy to demonstrate and gamify, but optimize completion proxies. They can be engagement surfaces only. + +### C. Takeover Twin without an Operator Model + +Preserves active episodes, but episodes remain isolated events. The product cannot compute what changed in the human-system relationship or why one intervention is worth the attention. + +### D. Executable Operator Model plus shadow control — selected + +Higher contract and experiment complexity, but it turns human readiness into a versioned, falsifiable product output while preserving full code-generation autonomy. + +## Consequences + +### Easier + +- explain why a developer is being interrupted; +- avoid beginner prompts on already-controlled seams; +- detect knowledge drift after code changes; +- connect attention to reusable tests and recovery infrastructure; +- compare human-selected control with an automatic policy; +- make the product valuable even when the agent's code is correct. + +### Harder + +- causal claims and control surfaces must be compiled without inventing facts; +- the product needs invalidation and provenance rules, not a mutable score; +- UI cannot hide uncertainty behind “understanding percentages”; +- effect must be tested on delayed adjacent tasks with real people; +- support will initially be narrow and test-backed. + +## Falsifiers + +Reject or narrow the decision if any preregistered study shows: + +- model-driven selection does not beat simple seam selection per active minute; +- the same cold executor succeeds equally with automatic evidence and directives; +- Operator Model state does not predict delayed takeover better than confidence or diff time; +- control dividends are not reused or create material maintenance cost; +- users set the budget to zero after novelty fades; +- production slowdown exceeds the existing non-inferiority margin; +- benefits occur only on the exact practiced mutation. + +## Entry and implementation gates + +This ADR does not authorize R5/R6 implementation. Before acceptance: + +1. complete and adjudicate the two-rater R7 expert audit; +2. freeze the exact Operator Model projection and hash domains in `CONTRACTS.md`; +3. add one model-delta fixture without a UI; +4. compare model-delta selection with random and weighted seam baselines; +5. compare human-directed cold relay with an automatic evidence policy; +6. only then expose model state in the cockpit. + +## Research anchors + +- Shen and Tamkin, [How AI Impacts Skill Formation](https://arxiv.org/abs/2601.20245), 2026. +- Buçinca, Malaya, and Gajos, [To Trust or to Think](https://doi.org/10.1145/3449287), 2021. +- NASA, [Developing a General Framework for Human-Autonomy Teaming](https://ntrs.nasa.gov/api/citations/20170003682/downloads/20170003682.pdf), 2017. +- Endsley and Kiris, [The Out-of-the-Loop Performance Problem and Level of Control in Automation](https://doi.org/10.1518/001872095779064555), 1995. + +These sources motivate the mechanism. They do not prove that an Executable Operator Model preserves programming skill. diff --git a/docs/v0.3/ADR-008-EVIDENCE-CARRYING-GENERATION.md b/docs/v0.3/ADR-008-EVIDENCE-CARRYING-GENERATION.md new file mode 100644 index 0000000..406d3f2 --- /dev/null +++ b/docs/v0.3/ADR-008-EVIDENCE-CARRYING-GENERATION.md @@ -0,0 +1,338 @@ +# ADR-008 — Evidence-Carrying Generation and the Intent Ledger + +- **Status:** Proposed; implementation gated by the complete R7 expert audit +- **Date:** 2026-08-01 +- **Deciders:** repository owner after R7 adjudication +- **Extends:** ADR-007 + +## Context + +ADR-007 defines the Executable Operator Model: a versioned record of the project behavior a developer has actually demonstrated through prediction, evidence selection, intervention, recovery, or delayed transfer. That solves the human-readiness side of the product, but it does not fully solve the artifact side. + +An autonomous run may change thousands of lines. Requiring the developer to read every diff recreates the review bottleneck PureFlow is meant to remove. Asking an LLM to summarize the diff is not enough: the same model that generated a change can produce a plausible but false explanation of it. Sampling only critical seams preserves takeover ability at those seams, but leaves the rest of the generated artifact opaque. + +The original product goal therefore needs two separate guarantees: + +1. **artifact accountability:** every changed line can be traced to a bounded semantic unit, an available reason, and the evidence or gap attached to that reason; +2. **operator readiness:** for selected critical seams, the developer has behaviorally demonstrated the ability to predict, observe, direct, and recover. + +PureFlow must provide the first without pretending it proves the second. + +## Decision + +PureFlow will add an **Intent Ledger** between the autonomous build plane and the Executable Operator Model. The ledger is a revision-bound, bidirectional intermediate representation of why the current change exists and what evidence supports that account. + +```text +human goal / committed decision + │ + ▼ + Intent Ledger N ───────────────► build-agent context + │ │ + │ forward obligations │ generated patch + claim refs + │ ▼ + ◄──────────── Flight Recorder / semantic compiler + │ reverse attribution + evidence join + ▼ + Intent Delta N→N+1 + supported / claimed / unattributed / contradicted / stale + │ + ├────────► navigable accountability for every changed line + └────────► Operator Model divergence and control selection +``` + +The build plane remains free to write all implementation code. It is asked to emit structured claim references while it works, but those references are never trusted as evidence by themselves. The Flight Recorder independently binds the resulting diff to syntax, build, test, trace, and repository evidence. Unsupported associations remain `claimed`; missing associations become `unattributed`. + +### Unit of accountability + +The ledger does not force one explanation per physical line. It compiles the smallest supported semantic unit available: + +- symbol, declaration, branch, call site, schema field, route, or configuration key when a language adapter supports it; +- deterministic generator invocation for derived files and lockfiles; +- bounded diff hunk when no safer structural unit exists. + +Every changed line must belong to exactly one primary unit for coverage accounting. A unit may have additional dependency and evidence edges. Unsupported files are still represented as bounded hunks, so the compiler cannot silently drop difficult changes. + +### Attribution states + +| State | Meaning | What it may claim | +| --- | --- | --- | +| `supported` | the structural mapping and at least one evidence reference were independently joined | this unit is connected to the stated intent and available evidence | +| `claimed` | an agent or imported artifact asserted the link, but independent support is incomplete | the build plane says this is why it changed | +| `unattributed` | no bounded intent link is available | the reason for this unit is an explicit gap | +| `contradicted` | executable evidence conflicts with the attached behavior claim | the current account is wrong or incomplete | +| `stale` | the source, intent, or evidence authority changed after the link was compiled | the previous account must be recomputed | + +`supported` means supported attribution, not proof of correctness, safety, or human understanding. + +### Bidirectional behavior + +The ledger operates in both directions: + +- **forward:** committed goals, constraints, decisions, and known invariants become obligations the build agent can reference while planning and generating; +- **reverse:** the actual patch, tests, traces, and agent events produce an `IntentDelta` that shows added behavior, changed assumptions, stale decisions, contradictions, and unattributed code. + +The reverse pass may create an `agent-claimed` intent node. It must never silently rewrite a user-committed node or present inferred prose as the developer's intent. + +### Generated and mechanical artifacts + +Formatting changes, generated sources, snapshots, migrations, and lockfiles can dominate a diff without deserving human attention. They are accounted for through a deterministic generation or transformation receipt that names: + +- the source-of-truth unit; +- the frozen command or controller-owned transform; +- input and output hashes; +- the evidence that regeneration is reproducible. + +This preserves total line coverage without turning mechanical output into a thousand fake decisions. + +### Relationship to the Operator Model + +The two artifacts have different authority: + +```text +Intent Ledger = what the software change is connected to +Operator Model = what the human has demonstrated they can control +``` + +The Operator Model may consume ledger gaps to select a shadow-control episode. A ledger link alone cannot make an operator claim `fresh`. Only the executable human evidence defined by ADR-007 can do that. + +## Proposed contracts + +These shapes are non-normative until the ADR is accepted. Final contracts require strict participant/internal projections, domain-separated RFC 8785 hashes, deterministic ordering, and controller-owned authorities. + +```ts +type IntentOrigin = + | "user-committed" + | "agent-claimed" + | "repository-imported" + | "executable-observed"; + +type AttributionState = + | "supported" + | "claimed" + | "unattributed" + | "contradicted" + | "stale"; + +interface IntentNode { + schemaVersion: 1; + id: string; + projectId: string; + sourceTreeHash: string; + kind: "goal" | "requirement" | "invariant" | "decision" | "constraint"; + origin: IntentOrigin; + statement: string; + parentIds: string[]; + evidenceRefs: string[]; + nodeHash: string; +} + +interface SemanticUnit { + schemaVersion: 1; + id: string; + projectId: string; + sourceTreeHash: string; + adapterId: string; + kind: "symbol" | "branch" | "call-site" | "config" | "generated" | "hunk"; + relativePath: string; + changedLineDigest: string; + sourceDigest: string; + unitHash: string; +} + +interface SemanticAttribution { + schemaVersion: 1; + projectId: string; + sourceTreeHash: string; + unitId: string; + intentNodeIds: string[]; + evidenceRefs: string[]; + controlSurfaceHash?: string; + state: AttributionState; + reasonCodes: string[]; + attributionHash: string; +} + +interface IntentLedgerSnapshot { + schemaVersion: 1; + projectId: string; + sourceTreeHash: string; + nodes: IntentNode[]; + units: SemanticUnit[]; + attributions: SemanticAttribution[]; + unattributedUnitIds: string[]; + snapshotHash: string; +} + +interface IntentDelta { + schemaVersion: 1; + projectId: string; + beforeSnapshotHash: string; + sourceTreeHash: string; + addedUnitIds: string[]; + changedUnitIds: string[]; + staleAttributionIds: string[]; + contradictedAttributionIds: string[]; + unattributedUnitIds: string[]; + deltaHash: string; +} +``` + +The canonical form must reject absolute paths, duplicate IDs, unknown evidence authorities, overlapping primary line ownership, omitted changed lines, hash drift, and cross-project references. + +## User experience + +The normal view is a compact change account, not a graph editor or mandatory review gate: + +```text +Agent checkpoint 42 + +2,143 changed lines → 37 semantic units +29 supported · 6 claimed · 2 unattributed + +New behavior: tenant cache invalidation +Changed decision: retry moved behind idempotency boundary +Evidence gap: rollback has no executable observation + +Best 4-minute action: +Test which signal distinguishes a stale key from a dropped invalidation. + +[Take control] [Open change account] [Keep agents running] +``` + +From any changed line, the developer can open a compact chain: + +```text +line → semantic unit → intent/claim → executable evidence → control surface +``` + +If a link is missing, the UI says so. It does not fill the gap with an LLM explanation. A small side-chat model may ask the developer to commit a decision, predict behavior, or choose evidence, but its assessment cannot change ledger or readiness authority. + +## Integrity rules + +### No trace-washing + +An agent cannot mark its own attribution `supported`. Agent-emitted references start as `claimed` until a deterministic structural join and an allowed evidence authority support them. + +### No test monoculture + +A test generated in the same action is useful evidence but not independent proof. The ledger records evidence origin so the selector can prefer external contracts, prior tests, runtime traces, or later human-created control dividends. + +### No giant semantic units + +Adapters must expose a configured maximum unit span and deterministic fallback. A repository-wide or file-wide unit cannot hide unrelated behavior merely to improve coverage. + +### No silent exclusions + +Ignored, binary, vendor, and unsupported artifacts receive explicit classification. Changed text lines must still reconcile against the Git diff. Generated artifacts may use one receipt-backed unit, but cannot disappear from totals. + +### No competence inference + +Opening the account, reading an explanation, accepting a node, or achieving high attribution coverage must not update the Executable Operator Model or a readiness claim. + +## Prior-art boundary + +This decision borrows useful ideas but defines a different product object: + +- Microsoft's Programming with Representations uses a domain-specific representation and guardrails to translate natural-language intent into robust code, including for people without coding expertise. PureFlow's ledger is revision-bound, accepts arbitrary existing code, exposes reverse-attribution failures, and exists to preserve professional ownership rather than decouple the user from coding expertise. +- Requirements-to-code traceability links requirements, implementation, and tests. The Intent Ledger adds live agent-event provenance, explicit unsupported states, control surfaces, and a separate behaviorally verified Operator Model. +- Proof-carrying code lets a consumer validate adherence to a safety policy. Evidence-carrying generation is deliberately weaker: it carries inspectable evidence and gaps, not a mathematical proof or a total-correctness claim. +- Code summaries and code-to-text representations can improve navigation, but generated prose has no authority unless joined to source structure and executable evidence. + +The intended novelty is the combination of total change reconciliation, untrusted agent claims, executable evidence, visible attribution debt, and a separate human-control model inside an autonomy-first IDE. + +## Validation plan + +The first post-R7 slice is a compiler and experiment, not a broad UI build. + +### Technical acceptance + +1. For every supported text diff, 100% of changed lines reconcile to exactly one primary semantic unit or the compile fails. +2. Reopening identical inputs yields byte-identical canonical snapshots and hashes on Windows and Linux. +3. No model-only claim can become `supported` in the negative fixture suite. +4. Unsupported languages fall back to bounded hunks without silently losing lines. +5. A material source, intent, evidence, or generator change deterministically marks affected links `stale`. +6. Clicking any changed line resolves to its unit and current attribution state without network access. +7. Deleting the project ledger removes all project-scoped attribution data without changing source code. + +### Mechanism experiment + +Freeze a corpus of real agent-produced changes before evaluation. Compare: + +- raw diff plus ordinary AI summary; +- semantic change account without executable evidence; +- Intent Ledger plus Operator Model routing. + +Measure supported-attribution precision and recall against two independent expert raters, false-provenance rate, unresolved-unit visibility, time to answer a code-navigation question, delayed adjacent takeover success, production critical-path time, active human minutes, and ledger maintenance cost. Freeze thresholds and the adjudication rule before opening the held-out changes. + +### Falsifiers + +Reject or narrow the decision if any preregistered study shows: + +- expert raters cannot reliably agree whether intent links are supported; +- false supported provenance survives the negative corpus; +- a plain AST/diff navigator performs as well on delayed takeover and causal navigation; +- users treat `claimed` links as truth despite the state distinction; +- incremental maintenance cost erases the production-speed advantage of agentic coding; +- reverse compilation produces mostly generic intent nodes that do not help control selection; +- Intent Ledger coverage does not improve selection beyond ADR-007's weighted semantic seams; +- benefits disappear on another supported language or real repository. + +## Options considered + +### A. Require full diff review + +Provides visibility but does not scale and restores the exact bottleneck the product rejects. + +### B. Generate a repository explanation or code graph + +Fast and useful for navigation, but cannot distinguish evidence from a plausible model story and does not maintain human control. + +### C. Use only the Executable Operator Model + +Strong for critical takeover seams, but does not reconcile the rest of a large generated change. + +### D. Make a domain DSL the source of truth + +Can improve reliability in bounded domains, but does not fit arbitrary existing repositories and shifts the product toward low-code generation. + +### E. Intent Ledger plus Executable Operator Model — selected + +Maintains autonomous generation, accounts for the whole change, and spends scarce human attention only where behavioral control evidence is valuable. + +## Consequences + +### Easier + +- navigate a 2,000-line agent change without pretending every line was manually reviewed; +- expose code the agent cannot honestly explain or support; +- detect drift between decisions, generated behavior, and tests; +- route control episodes from concrete accountability gaps; +- reuse human interventions as better evidence for future agents. + +### Harder + +- each language adapter needs deterministic unit boundaries and fallback behavior; +- attribution authority and evidence origin must survive rebases and regeneration; +- UI must communicate uncertainty without turning states into a vanity score; +- useful intent granularity is an empirical question; +- storage and incremental compilation need measurement on real repositories. + +## Entry and implementation gates + +This ADR does not authorize R5/R6 implementation. Before acceptance: + +1. complete and adjudicate the two-rater R7 expert audit; +2. review ADR-007 and ADR-008 together as one product architecture; +3. freeze unit-boundary, evidence-authority, projection, and hash rules in `CONTRACTS.md`; +4. build one offline Intent Ledger fixture with deliberately false agent claims and unsupported files; +5. pass the technical acceptance suite before any cockpit coverage UI; +6. run the held-out attribution study before using the ledger to make product claims. + +## Research anchors + +- Microsoft Research, [Programming with Representations](https://www.microsoft.com/en-us/research/project/pwr/), 2023–2026. +- YM et al., [PwR: Exploring the Role of Representations in Conversational Programming](https://www.microsoft.com/en-us/research/publication/pwr-exploring-the-role-of-representations-in-conversational-programming/), 2023. +- Schlathölter, [ReqToCode: Embedding Requirements Traceability as a Structural Property of the Codebase](https://arxiv.org/abs/2603.13999), 2026. +- Necula, [Proof-Carrying Code](https://doi.org/10.1145/263699.263712), 1997. + +These sources establish adjacent mechanisms. They do not validate PureFlow's combined architecture or its skill-preservation effect. diff --git a/docs/v0.3/ADR-009-DECISION-FUTURES.md b/docs/v0.3/ADR-009-DECISION-FUTURES.md new file mode 100644 index 0000000..fd9ec61 --- /dev/null +++ b/docs/v0.3/ADR-009-DECISION-FUTURES.md @@ -0,0 +1,425 @@ +# ADR-009 — Decision Futures and the Takeover Envelope + +- **Status:** Proposed; implementation gated by the complete R7 expert audit +- **Date:** 2026-08-01 +- **Deciders:** repository owner after R7 adjudication +- **Extends:** ADR-001, ADR-007, and ADR-008 + +## Context + +PureFlow already separates autonomous software production from human readiness work: + +- the Intent Ledger accounts for the generated artifact; +- the Executable Operator Model records behavior the developer has demonstrated; +- shadow control and Takeover Twins exercise prediction, diagnosis, intervention, and recovery without stopping the production swarm. + +One gap remains. Most shadow-control actions happen beside production. They can prove that the developer understood or controlled a seam, but they do not necessarily give the developer real authority over what the product becomes. If every consequential decision is still made by an agent and the human only practices on a copy, the system risks creating a competent simulator operator who no longer makes engineering decisions in the live project. + +ADR-001 names speculative live steering but does not define an executable protocol. It has no honest answer to these questions: + +- Was there a real architectural fork or a manufactured quiz? +- Did the developer's commitment actually determine the integrated path? +- How long could agents continue before a decision was needed? +- What happens if the human responds late or skips? +- How is the interruption timed and justified? +- How does one live decision change the future autonomy policy? + +The final product needs an autonomy mechanism in which agents retain implementation speed while the developer retains causal and architectural authority. + +## Decision + +PureFlow will represent selected high-leverage live forks as **Decision Futures**. A Decision Future is a bounded interval during which agents speculatively execute viable alternatives while unrelated production work continues. The developer may commit a prediction, choose discriminating evidence, and select or conditionally direct the path before a controller-owned integration deadline. + +```text +real unresolved fork + │ + ▼ +Decision Future opens + ├────────► agent branch A ── test / trace / benchmark + ├────────► agent branch B ── test / trace / benchmark + └────────► unrelated swarm work continues + │ + ▼ + low-cost human breakpoint + predict → inspect → decide + │ + ┌──────────┴──────────┐ + │ before deadline │ late / skipped + ▼ ▼ + choice controls merge autonomous policy merges + │ │ + ▼ ▼ + live decision evidence optional counterfactual twin + └──────────┬──────────┘ + ▼ + Operator Model + Intent Ledger + ▼ + Takeover Envelope +``` + +Agents may write every implementation line in every alternative. The human contribution is the scarce part automation otherwise removes: a pre-reveal causal model, evidence strategy, trade-off decision, and recovery policy. + +### Decision Future protocol + +1. **Detect.** The Flight Recorder finds a natural disagreement, plan revision, constraint conflict, near miss, or unresolved high-leverage trade-off. +2. **Qualify.** The compiler verifies that at least two alternatives are viable, behaviorally distinguishable, and safe to execute in isolated worktrees or twins. +3. **Precommit autonomy.** The controller records the default autonomous policy, integration deadline, reversible boundary, allowed evidence, and what unrelated work may continue. +4. **Execute speculatively.** Build agents implement and test alternatives without waiting for the human. +5. **Offer at a breakpoint.** The cockpit waits for an existing low-cost boundary such as an agent checkpoint, test completion, task switch, or explicit developer availability signal. +6. **Commit before reveal.** The developer predicts a consequence or chooses the observation that would separate alternatives before the decisive result is exposed. +7. **Exercise authority.** A valid on-time commitment selects an existing alternative or bounded conditional directive. The controller, not the model, performs integration. +8. **Join evidence.** The judge records whether the decision affected production, whether the prediction matched behavior, what recovery was needed, and which control dividend survived integration. + +### A real fork, not generated theatre + +A Decision Future is eligible only when: + +- the alternatives arose from the current task, agent disagreement, evidence mismatch, or a controller-derived adjacent design; +- both alternatives satisfy a frozen minimum validity suite before comparison; +- at least one observable consequence can discriminate between them; +- choosing between them would change a behavior, boundary, failure envelope, operating cost, or recovery path; +- the expected decision value justifies the human attention cost; +- final integration can be delayed within a declared bound while unrelated work continues. + +Two cosmetically different patches, an obviously inferior decoy, or alternatives that produce the same relevant behavior are not a Decision Future. The compiler must abstain rather than manufacture agency. + +### Timing modes + +| Outcome | Production behavior | Evidence meaning | +| --- | --- | --- | +| `live-steering` | on-time human commitment determines the integrated alternative | may create human decision evidence after executable judging | +| `counterfactual` | autonomous integration already occurred; the human decision runs in a twin | may create capability evidence, never retroactive production influence | +| `auto-default` | no valid commitment arrived; the precommitted policy chooses | no human decision evidence | +| `skipped` | user explicitly declines; the precommitted policy chooses | skip is recorded only for scheduling, not readiness | +| `failed-integrity` | authority, hash, deadline, or evidence binding is invalid | no integration or readiness claim from the session | + +The deadline is not a countdown dark pattern. It is the technical point after which keeping speculative branches alive would exceed the declared latency or compute budget. The UI shows the default action and consequence of waiting. + +### Low-cost initiative transfer + +PureFlow does not interrupt on a timer. The router estimates: + +```text +expected control value +− interruption and resumption cost +− speculative compute cost +− integration-delay risk +``` + +It offers a Decision Future only when the result is positive under the user's session attention budget. The first implementation remains deterministic and explainable. Learning a personal interruption policy is a later experiment and cannot use hidden surveillance signals. + +Useful breakpoints include: + +- the developer just completed or switched a task; +- the active agent is waiting for tests or external I/O; +- a stable checkpoint has landed and the next action has not started; +- the developer opened the cockpit or explicitly requested a control opportunity. + +Continuous popups, random functions, engagement streaks, and notifications during active typing are excluded from the core protocol. + +## Takeover Envelope + +The vehicle analogy needs two distinct operating domains: + +```text +Autonomy Envelope = where agents can currently execute with accepted software evidence +Takeover Envelope = where the human has current executable control and transfer evidence +``` + +PureFlow does not collapse either envelope into one percentage. It presents the critical seams and the relationship between them: + +| Seam state | Meaning | +| --- | --- | +| `autonomous-and-covered` | agents can operate and the human has fresh takeover evidence | +| `autonomous-human-stale` | agents can operate but prior human evidence no longer matches the source or control surface | +| `autonomous-uncovered` | agents can operate; no current human recovery evidence exists | +| `human-controlled-only` | the human has evidence but the autonomous path is unsupported or failing | +| `unsupported` | neither plane has accepted evidence | + +Full agentic coding is still allowed in `autonomous-uncovered` areas unless the user explicitly configures a high-risk team policy. PureFlow simply refuses to claim that takeover readiness exists there. + +Decision Futures shrink the highest-value divergence between the two envelopes by letting the developer make a real live decision. Shadow control and delayed transfer verify that the decision was not a one-time recognition event. + +## Proposed contracts + +These shapes remain non-normative until acceptance. Final contracts require strict internal/participant projections, controller-owned authorities, deterministic ordering, RFC 8785 canonicalization, and domain-separated hashes. + +```ts +type DecisionReason = + | "agent-disagreement" + | "plan-revision" + | "constraint-conflict" + | "evidence-mismatch" + | "near-miss" + | "uncovered-control-route"; + +interface DecisionAlternative { + id: string; + snapshotTreeHash: string; + behaviorClaimHash: string; + evidenceIds: string[]; + judgeSpecHash: string; + reversible: boolean; +} + +interface DecisionFuture { + schemaVersion: 1; + id: string; + projectId: string; + baseTreeHash: string; + intentDeltaHash: string; + operatorModelDeltaHash: string; + reason: DecisionReason; + alternatives: DecisionAlternative[]; + availableEvidenceIds: string[]; + defaultPolicyId: string; + openedAt: string; + integrationDeadline: string; + maxActiveSeconds: number; + authorityHash: string; + futureHash: string; +} + +interface DecisionCommitment { + schemaVersion: 1; + futureHash: string; + selectedAlternativeId?: string; + selectedExperimentId?: string; + predictionRef: string; + conditionalDirectiveRef?: string; + disclosedEvidenceIds: string[]; + committedAt: string; + commitmentHash: string; +} + +type DecisionOutcomeMode = + | "live-steering" + | "counterfactual" + | "auto-default" + | "skipped" + | "failed-integrity"; + +interface DecisionOutcome { + schemaVersion: 1; + futureHash: string; + commitmentHash?: string; + mode: DecisionOutcomeMode; + integratedAlternativeId: string; + humanSelectionAffectedIntegration: boolean; + judgeResultHash?: string; + controlDividendHash?: string; + outcomeHash: string; +} + +interface TakeoverEnvelopeSnapshot { + schemaVersion: 1; + projectId: string; + sourceTreeHash: string; + autonomyEvidenceHash: string; + operatorModelSnapshotHash: string; + seamStates: Array<{ + seamId: string; + state: + | "autonomous-and-covered" + | "autonomous-human-stale" + | "autonomous-uncovered" + | "human-controlled-only" + | "unsupported"; + reasonCodes: string[]; + }>; + snapshotHash: string; +} +``` + +Free-form text from a participant or model cannot become a branch name, command, path, environment value, judge, deadline, or integration instruction. A conditional directive is interpreted only by a context-starved executor inside the existing sandbox boundary; its patch must pass the same controller-owned judge before becoming a candidate. + +## User experience + +The cockpit explains why the decision is worth attention and what happens if the developer ignores it: + +```text +Agents found a real auth boundary fork + +A · retry outside idempotency lock + current tests pass · lower lock time · duplicate-write risk unknown + +B · retry inside idempotency lock + current tests pass · simpler recovery · higher contention + +Autopilot default in 7 min: A +Both branches are already being tested. Other agents keep working. + +Before the duplicate-delivery replay is revealed: +Which result would make you reject A? + +[Commit prediction] [Choose another observation] [Let autopilot decide] +``` + +After judging: + +```text +Your on-time decision selected B and determined the integrated path. +The replay confirmed your predicted duplicate-write boundary. + +Created control dividend: duplicate-delivery regression probe +Takeover envelope updated: auth/idempotency — immediate evidence only +Delayed adjacent recovery still required. +``` + +A late response uses different copy: + +```text +Autopilot already integrated A under the declared policy. +Your choice can still run as a counterfactual, but it cannot count as live influence. +``` + +## Interaction with existing architecture + +- **Flight Recorder** supplies observable disagreements, evidence changes, and timestamps. +- **Intent Ledger** binds alternatives to the same committed goal and exposes intent drift. +- **Controllability Compiler** requires each alternative to expose `observe → actuate → recover` handles. +- **Executable Operator Model** determines which decision would reduce a real human-control gap. +- **Decision Future** grants bounded live authority without requiring manual code. +- **Takeover Twin** preserves counterfactual and late-decision value off the production path. +- **Evidence Judge** decides consequences; a Side Coach never decides correctness. +- **Control Dividend** turns the human's attention into a reusable project artifact. +- **Takeover Envelope** makes the autonomy/readiness mismatch explicit for future routing. + +## Integrity rules + +### Precommitted default + +The autonomous policy and deadline are frozen before the participant projection opens. The controller cannot change the default after seeing the human's commitment. + +### Comparable alternatives + +All visible alternatives must pass the same declared minimum validity suite. Missing, failing, or incomparable evidence is visible and may force abstention. + +### No retroactive influence + +`humanSelectionAffectedIntegration` can be true only when a valid commitment preceded the deadline and the controller's integration record names that commitment. A later matching opinion cannot be upgraded. + +### No rubber-stamp evidence + +Selecting an alternative after decisive evidence is revealed, accepting the autonomous default, or approving an agent plan creates no decision evidence. The protocol requires a pre-reveal prediction or evidence strategy. + +### No production corruption + +Speculative alternatives use isolated worktrees. Only the controller may integrate a candidate after the required software checks. Abandonment and cleanup must preserve production refs, worktree state, and the prior stable checkpoint. + +### No hidden compulsion + +Budget `0`, explicit skip, timeout, or notification dismissal preserves full autonomous operation and creates no negative developer score. Required human approval exists only under a separately configured high-risk team policy. + +## Validation plan + +### Technical acceptance + +1. Replaying identical inputs produces byte-identical future, commitment, outcome, and envelope hashes on Windows and Linux. +2. The autonomous default, deadline, alternatives, and evidence catalog are immutable after the participant projection opens. +3. A late, replayed, cross-project, cross-tree, or post-reveal commitment cannot yield `live-steering`. +4. `humanSelectionAffectedIntegration` is true only when the merge record references the on-time commitment and selected alternative. +5. Skip and budget `0` leave production agents running and create no readiness evidence. +6. A false fork with equivalent behavior, an invalid alternative, or no discriminating observation produces abstention. +7. Every speculative worktree is removed and the production repository remains byte-for-byte invariant except for the controller-selected integration. +8. The Takeover Envelope changes only from accepted autonomy evidence and Operator Model evidence; interaction counts and confidence cannot alter it. + +### Mechanism experiment + +Compare three conditions on the same real agent-produced forks: + +1. autonomous default plus post-hoc explanation; +2. shadow prediction with no production influence; +3. Decision Future with bounded live influence. + +Primary outcome: delayed adjacent takeover success per active human minute. Secondary outcomes: production critical-path time, valid live-influence rate, prediction calibration, discriminating-experiment quality, recovery after a wrong first decision, control-dividend reuse, skip/late rate, interruption acceptance, speculative compute, and software quality. + +The automatic policy remains a comparator. A human decision is not presumed superior. The product hypothesis is that exercising real decision ownership improves later takeover without erasing agentic speed. + +Freeze the corpus, thresholds, default policy, timing rule, alternative-validity suite, adjudication procedure, and non-inferiority margin before revealing held-out outcomes. + +### Falsifiers + +Reject or narrow this decision if a preregistered study shows: + +- genuine behaviorally distinct forks are too rare to sustain the mechanism; +- users mostly rubber-stamp the default or respond only after decisive evidence; +- live influence does not improve delayed takeover over shadow prediction alone; +- interruption and branch cost erase the production-speed benefit; +- the automatic policy matches or beats the human condition on both software and delayed-transfer outcomes; +- human choices affect integration but produce no reusable control artifacts; +- developers reduce their attention budget to zero after novelty fades; +- the Takeover Envelope does not predict blackout-recovery performance; +- meaningful results require blocking every merge or manufacturing inferior decoys. + +## Prior-art boundary + +Mixed-initiative interaction, adjustable autonomy, speculative execution, and interruption management are established fields. NASA human-autonomy work specifically recommends negotiated decisions and greater interaction to reduce out-of-the-loop situation-awareness loss. Research on collaborative interruptions shows that timing and perceived benefit influence whether people accept an interruption. Recent professional-agency research also distinguishes automating execution from retaining human judgment. + +PureFlow does not claim novelty for any one of those ideas. The proposed contribution is their software-development combination: + +> agents speculatively implement real alternatives + the human commits before reveal + an on-time choice controls integration + executable consequences update a project-specific takeover envelope + delayed transfer tests whether control survives later + +No reviewed mainstream AI IDE documents that complete loop. This is a bounded landscape statement, not proof of global or patent novelty. + +## Options considered + +### A. Ask the developer to approve every plan + +High interruption cost, easy to rubber-stamp, and no evidence that the approval changed an outcome. + +### B. Ask random questions while the agent works + +May support recall, but does not exercise engineering authority and can become engagement theatre. + +### C. Run only off-path takeover rehearsals + +Useful for diagnosis and recovery, but all live product decisions remain delegated. + +### D. Block final integration until a human reviews the diff + +Restores the code-review bottleneck and makes throughput depend on passive inspection. + +### E. Decision Futures plus a Takeover Envelope — selected + +Preserves autonomous implementation and unrelated progress while giving a small number of human decisions real, measurable production authority. + +## Consequences + +### Easier + +- explain how the developer still makes consequential engineering decisions while agents write the code; +- use genuine agent disagreement as productive human work; +- separate live influence from practice and post-hoc agreement; +- time interactions around expected value instead of timers; +- show exactly where autopilot capability exceeds human takeover evidence. + +### Harder + +- alternatives must be genuinely comparable and safely isolated; +- speculative execution consumes compute and may delay one integration boundary; +- deadlines, defaults, and authority need tamper-evident contracts; +- some tasks have no honest design fork and must produce no interaction; +- long-term voluntary use requires human studies, not only compiler tests. + +## Entry and implementation gates + +This ADR does not authorize R5/R6 implementation. Before acceptance: + +1. complete and adjudicate the two-rater R7 expert audit; +2. review ADR-007, ADR-008, and ADR-009 as one architecture; +3. freeze Decision Future projections, hash domains, alternative validity, deadline semantics, and integration authority in `CONTRACTS.md`; +4. compile one offline natural disagreement fixture with a false-fork negative control; +5. compare shadow-only and live-influence modes before expanding the cockpit; +6. run the delayed-transfer human pilot before claiming skill preservation. + +## Research anchors + +- Horvitz, [Mixed-Initiative Interaction](https://www.microsoft.com/en-us/research/publication/mixed-initiative-interaction/), 1999. +- Kamar, Gal, and Grosz, [Modeling Information Exchange Opportunities for Effective Human-Computer Teamwork](https://www.microsoft.com/en-us/research/publication/modeling-information-exchange-opportunities-for-effective-human-computer-teamwork/), 2013. +- Iqbal and Horvitz, [Conversations Amidst Computing](https://www.microsoft.com/en-us/research/publication/conversations-amidst-computing-a-study-of-interruptions-and-recovery-of-task-activity/), 2007. +- NASA, [Understanding Human Autonomy Teaming Through Applications](https://ntrs.nasa.gov/api/citations/20170010163/downloads/20170010163.pdf), 2017. +- Nishal et al., [Helping Me Versus Doing It for Me](https://www.microsoft.com/en-us/research/publication/helping-me-versus-doing-it-for-me-designing-for-agency-in-llm-infused-writing-tools-for-science-journalism/), 2026. +- Buçinca, Malaya, and Gajos, [To Trust or to Think](https://doi.org/10.1145/3449287), 2021. + +These sources motivate negotiated authority, pre-reveal engagement, and interruption-aware scheduling. They do not validate Decision Futures or prove programming-skill preservation. diff --git a/docs/v0.3/CONCEPTS.md b/docs/v0.3/CONCEPTS.md index 4c7ff8b..4404477 100644 --- a/docs/v0.3/CONCEPTS.md +++ b/docs/v0.3/CONCEPTS.md @@ -100,6 +100,7 @@ This combines four supporting mechanisms: - evidence-carrying generation supplies attribution and seam candidates; - an attention scheduler chooses the smallest valuable human episode; +- Decision Futures let a bounded pre-reveal human choice determine a live integrated path while agents implement alternatives; - executable fault/counterfactual practice produces behavioral evidence; - a readiness-based router closes the loop with future delegation. @@ -154,7 +155,8 @@ Bug injection is only one possible experience backend. The architecture differs 2. production agents continue working in parallel; 3. the selector chooses the seam from readiness and future takeover value; 4. the episode may be prediction, counterfactual steering, diagnosis, intervention, or recovery—not only a seeded bug; -5. delayed evidence influences later task delegation. +5. selected genuine forks may give an on-time human commitment real integration authority without manual implementation; +6. delayed evidence influences later task delegation. Removing any two of those links risks collapsing the product into an existing category. diff --git a/docs/v0.3/GOAL_COMPLETION_AUDIT.md b/docs/v0.3/GOAL_COMPLETION_AUDIT.md new file mode 100644 index 0000000..69bca1e --- /dev/null +++ b/docs/v0.3/GOAL_COMPLETION_AUDIT.md @@ -0,0 +1,56 @@ +# Original-goal completion audit + +**Date:** 2026-08-01 +**Verdict:** active and incomplete + +The original goal is broader than a compiler benchmark or an explanatory extension. Completion requires a usable agentic IDE in which autonomous coding remains first-class and project-specific engineering capability is behaviorally preserved. This audit prevents completed infrastructure from being mistaken for the final product. + +| Requirement from the goal | Current evidence | Status | Evidence required for completion | +| --- | --- | --- | --- | +| Full agentic coding remains available | v0.1 is an IDE shell; v0.3 has a replay `AgentDriver`; ADR-006 selects a future live adapter | Incomplete | a live supported agent can plan, edit, test, and repair through the IDE on a real repository | +| Developer is not reduced to reading generated diffs | PRD, THESIS, ADR-007, ADR-008, and ADR-009 replace passive review with accountable generation, shadow control, and bounded live decisions | Architecture only | production UX demonstrates prediction, evidence selection, live influence, intervention, and recovery without requiring line-by-line review | +| Developer understands what the system does | R2 maps changed units to evidence; ADR-008 specifies total line reconciliation and explicit attribution debt | Partially represented, not proven | Intent Ledger exposes the current change account while the Executable Operator Model shows demonstrated causal claims, uncovered seams, contradictions, and delayed transfer | +| Developer continues making engineering decisions | ADR-009 specifies Decision Futures, bounded live influence, dissent, and a Takeover Envelope | Not implemented | an on-time pre-reveal commitment determines a real integrated path and later predicts adjacent takeover better than shadow-only control | +| AI can coach and question during work | Fixture-only precommitted Control Pulse exists; configured Side Coach remains bounded | Narrow fixture only | event-driven side interaction works during a live agent run and executable evidence, not an LLM score, adjudicates it | +| AI may still write almost all code | Context-starved relay architecture explicitly allows this | Not implemented | a fresh agent writes a complete intervention from human-selected evidence and directive in the sandbox | +| Skills remain available if AI disappears | Delayed AI-off adjacent takeover is the north-star metric | Unmeasured | preregistered controlled human study clears the delayed-transfer gate and later field evidence survives novelty | +| The product is more than a quiz/extension | Dual planes, SandboxRunner, Intent Ledger, compiler, judge, Operator Model, Decision Futures, and Takeover Envelope are designed | No v0.3 product runtime | readiness ledger, cockpit, live agent adapter, shadow/live-control flows, and local deletion/export work end-to-end | +| Autonomy speed is preserved | continuity rehearsal stays off-path; Decision Futures delay only one bounded integration boundary while agents execute alternatives and unrelated work | Unmeasured | production critical-path non-inferiority, speculative compute, and bounded attention pass the preregistered study | +| Every generated line is accountable | PRD and ADR-008 require deterministic reconciliation to `supported`, `claimed`, `unattributed`, `contradicted`, or `stale` units | Architecture plus narrow TypeScript evidence only | real agent changes pass total line reconciliation and held-out expert attribution checks without invented provenance | + +## What is complete + +- R0–R4.5 reviewed-fixture mechanism; +- digest-pinned Docker SandboxRunner contract and real integration tests; +- frozen 30-patch corpus and preregistered automatic R7 audit; +- automatic held-out result of 16/18 end-to-end valid episodes; +- blind packet generation and deterministic two-rater join tooling; +- falsifiable product thesis and the proposed Operator Model, Intent Ledger, Decision Futures, and Takeover Envelope extensions. + +## What is not complete + +- the two human R7 expert ratings and adjudication; +- post-gate readiness ledger and v0.3 cockpit; +- live Codex/Claude/OpenCode adapter accessible from this checkout; +- context-starved relay on arbitrary supported project code; +- Decision Futures with real bounded integration authority and an honest Takeover Envelope; +- delayed-transfer human experiment; +- longitudinal evidence of skill preservation; +- evidence that developers voluntarily keep a non-zero attention budget; +- any honest basis for saying PureFlow already preserves skills. + +## Completion rule + +The active goal may be marked complete only when all of these are true: + +1. a developer can use a real coding agent through PureFlow for ordinary project work; +2. the agent can write the implementation without a manual-code quota; +3. PureFlow compiles project-derived shadow control with no hidden-answer or production-worktree leak; +4. the developer can predict, choose evidence, direct a cold agent, recover, and make at least one pre-reveal decision that actually determines an integrated path; +5. local Operator Model state invalidates correctly as the code changes; +6. every changed line resolves through the local Intent Ledger to a bounded unit and an honest attribution state; +7. a delayed adjacent task demonstrates takeover without answer-generating AI; +8. a controlled study beats the active diff/explanation/question comparator while preserving production speed; +9. the runtime, deletion/export, packaging, and protected CI pass on supported platforms. + +Green unit tests for any individual component are necessary evidence, not completion of this goal. diff --git a/docs/v0.3/OPERATOR_MODEL_SPEC.md b/docs/v0.3/OPERATOR_MODEL_SPEC.md new file mode 100644 index 0000000..406dce4 --- /dev/null +++ b/docs/v0.3/OPERATOR_MODEL_SPEC.md @@ -0,0 +1,327 @@ +# Executable Operator Model vertical-slice specification + +- **Status:** Draft; implementation blocked by the complete R7 expert gate +- **Date:** 2026-08-01 +- **Source decision:** ADR-007 +- **Target:** first post-R7 mechanism ablation, before cockpit expansion + +## Context + +PureFlow needs to distinguish a developer who merely saw agent output from one who can control a changed project seam. Existing immediate evidence is event-scoped. It does not express which causal claims remain valid after later code changes, which critical seams have no human-controlled recovery path, or which small intervention would reduce that gap most efficiently. + +The first slice must prove that a deterministic operator-model delta can be derived from existing immutable project evidence and can select a more relevant shadow-control opportunity without using an LLM as the oracle. + +## Functional Requirements + +Requirement index: + +- FR-1: Immutable snapshot +- FR-2: Explicit claim states +- FR-3: Change invalidation +- FR-4: Executable evidence only +- FR-5: Deterministic delta +- FR-6: Bounded selection +- FR-7: Complete control surface +- FR-8: Pre-reveal commitment +- FR-9: Context-starved relay +- FR-10: Non-executable prose +- FR-11: Control dividend +- FR-12: Non-scalar evidence +- FR-13: Local ownership +- FR-14: Zero-budget autonomy + +### FR-1: Immutable snapshot + +The system MUST create an immutable `OperatorModelSnapshot` from a project ID, source tree, supported control surfaces, and existing capability evidence. + +### FR-2: Explicit claim states + +The system MUST label each claim `unverified`, `fresh`, `stale`, or `contradicted` using deterministic reason codes. + +### FR-3: Change invalidation + +A changed source tree MUST NOT preserve a `fresh` claim unless the compiler proves that the claim's bound semantic unit and control surface are unchanged. + +### FR-4: Executable evidence only + +Only prior executable capability evidence MAY advance a claim to `fresh`; explanations, quiz answers, confidence, and model annotations MUST NOT. + +### FR-5: Deterministic delta + +The system MUST emit a deterministic `ModelDelta` between a prior snapshot and a new source tree. + +### FR-6: Bounded selection + +The selector MUST choose at most one bounded shadow-control candidate from the delta for the fixture slice. + +### FR-7: Complete control surface + +The candidate MUST identify an existing controller-owned observation, actuator, recovery judge, and participant-safe prompt. + +### FR-8: Pre-reveal commitment + +The participant MUST commit a prediction before its bound observation is disclosed. + +### FR-9: Context-starved relay + +A context-starved relay MAY write the complete candidate patch, but it MUST receive only participant-selected evidence and a bounded committed directive. + +### FR-10: Non-executable prose + +Free-form model or participant text MUST NOT become executable input. + +### FR-11: Control dividend + +A terminal shadow-control result MUST create no more than one proposed control dividend, and that dividend MUST reference executable evidence. + +### FR-12: Non-scalar evidence + +The system MUST retain failures, assistance, abandonment, and contradiction; it MUST NOT collapse them into a scalar readiness score. + +### FR-13: Local ownership + +All data MUST remain local, project-scoped, exportable, and deletable. + +### FR-14: Zero-budget autonomy + +Budget `0` MUST produce no prompt and no readiness claim while leaving autonomous production unaffected. + +## Non-Functional Requirements + +- **NFR-1 Determinism:** Identical reopened inputs MUST yield byte-identical canonical snapshot, delta, and selection hashes on Windows and Linux. +- **NFR-2 Integrity:** Unknown fields, unsorted IDs, duplicate IDs, hash drift, cross-project evidence, and unknown authorities MUST fail before selection or execution. +- **NFR-3 Isolation:** Shadow execution MUST use the existing SandboxRunner boundary and MUST leave the production worktree unchanged. +- **NFR-4 Boundedness:** The fixture slice MUST cap one session at 600 active seconds, eight evidence references, and one control dividend. +- **NFR-5 Non-blocking production:** Skipping, abandoning, timing out, or failing shadow control MUST NOT block normal checkpoint integration. +- **NFR-6 Honesty:** UI and stored evidence MUST distinguish immediate capability evidence from delayed verified readiness. +- **NFR-7 Privacy:** Persisted records MUST contain no absolute workspace path, credential, raw terminal history, clipboard content, or hidden model reasoning. + +## Acceptance Criteria + +Traceability: (FR-1), (FR-2), (FR-3), (FR-4), (FR-5), (FR-6), (FR-7), (FR-8), (FR-9), (FR-10), (FR-11), (FR-12), (FR-13), and (FR-14). + +### AC-1: Cross-platform snapshot identity + +References FR-1, FR-4, and NFR-1. Given identical fixture evidence on Windows and Linux, when the snapshot compiler runs, then canonical JSON and `snapshotHash` are identical. + +### AC-2: Material change stales a claim + +References FR-2 and FR-3. Given a fresh claim and a materially changed bound unit, when a new snapshot is compiled, then the claim is `stale` with a deterministic reason and cannot remain `fresh`. + +### AC-3: Contradiction remains evidence + +References FR-2. Given an observation that contradicts a committed prediction, when the result is joined, then the claim becomes `contradicted` rather than failed or deleted. + +### AC-4: Prose cannot create freshness + +References FR-4 and FR-12. Given only an explanation, confidence value, quiz answer, or Side Coach annotation, when the snapshot is compiled, then no claim advances to `fresh`. + +### AC-5: Delta replay is deterministic + +References FR-5 and NFR-1. Given the same prior snapshot and new tree, when delta compilation repeats three times, then every output and hash is identical. + +### AC-6: Selection has a complete authority chain + +References FR-6 and FR-7. Given several gaps, when selection runs, then it returns one eligible candidate with a controller-owned observer, actuator, recovery judge, and transparent reason list. + +### AC-7: Observation requires commitment + +References FR-8. Given a shadow session, when observation execution is requested before the prediction commitment exists, then execution is rejected. + +### AC-8: Relay prose stays non-executable + +References FR-9 and FR-10. Given a relay directive containing code, command arguments, a path, environment value, or unknown evidence ID, when validation runs, then it fails before an executor starts. + +### AC-9: A passed session emits one bounded dividend + +References FR-11. Given a passed session, when its result is finalized, then at most one proposed dividend is emitted and it points to a verified test, probe, rollback, or control handle. + +### AC-10: Immediate evidence is not verified readiness + +References FR-12 and NFR-6. Given immediate recovery without delayed transfer, when evidence is rendered, then it is labeled capability evidence and never verified readiness. + +### AC-11: Project evidence is private and deletable + +References FR-13 and NFR-7. Given exported or persisted evidence, when scanned with absolute-path and secret fixtures, then none are present; deleting the project removes all Operator Model records. + +### AC-12: Zero budget leaves production autonomous + +References FR-14 and NFR-5. Given budget `0`, when an eligible delta exists, then agents continue, no shadow session is offered, and no capability event is fabricated. + +### AC-13: Every terminal path preserves production state + +References NFR-3. Given pass, fail, cancel, timeout, and tamper cases, when cleanup finishes, then source files, index, HEAD, refs, remotes, and worktree list match their captured state. + +## Edge Cases + +Edge-case index: + +- EC-1: Cross-project evidence +- EC-2: Unproven rename +- EC-3: Judge drift +- EC-4: Incompatible claims +- EC-5: Precommit skip +- EC-6: Protected-path relay patch +- EC-7: Partial timeout output +- EC-8: Missing dividend +- EC-9: Time decay +- EC-10: Attribution gap + +### EC-1: Cross-project evidence + +Evidence refers to a previous project identity. + +### EC-2: Unproven rename + +A semantic unit was renamed without behavioral evidence proving equivalence. + +### EC-3: Judge drift + +The control surface exists but its recovery judge changed. + +### EC-4: Incompatible claims + +Two claims bind to the same seam with incompatible predictions. + +### EC-5: Precommit skip + +The participant skips after seeing the prompt but before commitment. + +### EC-6: Protected-path relay patch + +The relay proposes a passing patch that edits a protected oracle or test. + +### EC-7: Partial timeout output + +The runner times out after producing partial output. + +### EC-8: Missing dividend + +A passed immediate episode has no valid reusable dividend. + +### EC-9: Time decay + +A previously fresh claim ages without a code change. + +### EC-10: Attribution gap + +The model compiler cannot attribute a changed line or invariant. + +Every edge case must produce an explicit state or reason code. None may silently become success. + +## API Contracts + +ADR-007 defines the proposed public shapes. Implementation must add exact JSON schemas and separate internal/participant projections before code. The participant projection may contain a bounded behavior claim, labels for approved observations, attention budget, and opaque session identity. It must exclude revisions, command definitions, controller authorities, oracle identities, hidden repairs, and source paths. + +HTTP endpoint: N/A — this slice is an injected extension-host TypeScript boundary with no network transport. Adding a fake `POST` route solely to satisfy a generic validator would widen the attack surface and misstate the architecture. + +Required service boundary: + +```ts +interface OperatorModelCompiler { + compileSnapshot(input: CompileOperatorSnapshotInput): Promise; + compileDelta(input: CompileOperatorDeltaInput): Promise; +} + +interface ShadowControlSelector { + select(input: SelectShadowControlInput): Promise; +} + +interface ShadowControlController { + commitPrediction(input: CommitShadowPredictionInput): Promise; + commitDirective(input: CommitRelayDirectiveInput): Promise; + execute(input: ExecuteShadowControlInput): Promise; + abandon(input: AbandonShadowControlInput): Promise; +} +``` + +## Data Models + +| Entity | Required identity | Persistence | Authority | +| --- | --- | --- | --- | +| `OperatorClaim` | project, source tree, seam, control surface | local append-only evidence plus derived snapshot | controller compiler | +| `OperatorModelSnapshot` | project and source tree | immutable local object | controller compiler | +| `ModelDelta` | before snapshot and new source tree | immutable experiment evidence | controller compiler | +| `ShadowControlSession` | project, delta, participant experience | session lifetime plus bounded result | controller | +| `ShadowControlResult` | session, committed prediction, judge result | append-only | executable judge | +| `ControlDividend` | terminal result and artifact hash | proposed until separately verified | controller-owned materializer | + +## Out of Scope + +Exclusion index: + +- OS-1: Live multi-agent adapter +- OS-2: Universal comprehension +- OS-3: Invented invariants +- OS-4: Management scoring +- OS-5: Manual-code quotas +- OS-6: Free-form execution +- OS-7: Immediate readiness claims +- OS-8: Stack expansion +- OS-9: Premature production UI +- OS-10: Skill-preservation claim + +### OS-1: Live multi-agent adapter + +Live multi-agent adapter implementation is excluded from this slice. + +### OS-2: Universal comprehension + +Universal program understanding is not attempted. + +### OS-3: Invented invariants + +Automatic natural-language invariant generation presented as fact is prohibited. + +### OS-4: Management scoring + +Manager dashboards and public developer rankings are excluded. + +### OS-5: Manual-code quotas + +No requirement is based on manually typed lines. + +### OS-6: Free-form execution + +Arbitrary shell commands from a model or participant are excluded. + +### OS-7: Immediate readiness claims + +Immediate performance cannot create verified readiness. + +### OS-8: Stack expansion + +More than one language or the preregistered TypeScript support envelope is excluded. + +### OS-9: Premature production UI + +Production UI is excluded until the selector and relay ablations pass. + +### OS-10: Skill-preservation claim + +This slice cannot establish that PureFlow preserves long-term skill. + +## Required Ablations + +The slice is not accepted from unit tests alone. Under the same attention and executor budget, compare: + +1. random changed function; +2. existing weighted semantic seam; +3. Operator Model delta selection. + +For cold relay, compare: + +1. automatic evidence and directive policy; +2. human-selected evidence plus committed conditional directive. + +Primary technical outcomes are valid executable selection, recovery after a wrong first action, fault-route coverage, and control-dividend validity. The later human outcome remains delayed regression-free adjacent-task success. + +## Implementation Entry Gate + +Do not generate tests or implementation until: + +- two independent R7 expert bundles are joined and adjudicated; +- R7 passes its complete preregistered gate; +- ADR-007 is accepted; +- exact schemas, hash domains, projections, and reason codes are added to `CONTRACTS.md`; +- the specification status changes from Draft to Approved. diff --git a/docs/v0.3/README.md b/docs/v0.3/README.md index 8a914a0..fd1cc1e 100644 --- a/docs/v0.3/README.md +++ b/docs/v0.3/README.md @@ -19,28 +19,39 @@ PureFlow v0.3 asks whether an AI IDE can keep autonomous coding fast while behav 13. [`R7_PREREGISTRATION.md`](R7_PREREGISTRATION.md) — frozen 30-patch sampling, metrics, and analysis protocol. 14. [`SANDBOX_RUNNER_SPEC.md`](SANDBOX_RUNNER_SPEC.md) — exact Phase-B runner acceptance and negative contract. 15. [`CONCEPT_LAB_CONTROLLABILITY.md`](CONCEPT_LAB_CONTROLLABILITY.md) — post-R7 control-surface, cold-relay, and dissent hypotheses. -16. [`AGENT_EXECUTION.md`](AGENT_EXECUTION.md) — ordered implementation workstreams and acceptance gates. -17. [`JULES_LOOP.md`](JULES_LOOP.md) — guarded server-side execution queue for the audited R0–R4 slice. +16. [`ADR-007-EXECUTABLE-OPERATOR-MODEL.md`](ADR-007-EXECUTABLE-OPERATOR-MODEL.md) — proposed maintained model of demonstrated human control and shadow-control protocol. +17. [`OPERATOR_MODEL_SPEC.md`](OPERATOR_MODEL_SPEC.md) — draft post-R7 vertical-slice requirements, contracts, and ablations. +18. [`ADR-008-EVIDENCE-CARRYING-GENERATION.md`](ADR-008-EVIDENCE-CARRYING-GENERATION.md) — proposed bidirectional Intent Ledger for total change accountability without passive full-diff review. +19. [`ADR-009-DECISION-FUTURES.md`](ADR-009-DECISION-FUTURES.md) — proposed speculative live-steering protocol and project-scoped Takeover Envelope. +20. [`GOAL_COMPLETION_AUDIT.md`](GOAL_COMPLETION_AUDIT.md) — requirement-by-requirement evidence separating infrastructure from the requested final product. +21. [`AGENT_EXECUTION.md`](AGENT_EXECUTION.md) — ordered implementation workstreams and acceptance gates. +22. [`JULES_LOOP.md`](JULES_LOOP.md) — guarded server-side execution queue for the audited R0–R4 slice. ## Current truth - The released v0.1 VSCodium IDE exists and remains the runtime baseline. -- The complete v0.3 Dual-Control product runtime is not implemented. R0–R4.5 implement and protect the reviewed-fixture path from canonical evidence through recovery judging and a precommitted Control Pulse. The digest-pinned Docker `SandboxRunner` has passed local Windows and protected Linux/Windows gates. The readiness ledger, cockpit, corpus audit, and human pilot remain gated. +- The complete v0.3 Dual-Control product runtime is not implemented. R0–R4.5 implement and protect the reviewed-fixture path from canonical evidence through recovery judging and a precommitted Control Pulse. The digest-pinned Docker `SandboxRunner` has passed local Windows and protected Linux/Windows gates. +- The frozen automatic R7 audit passed its preregistered automatic threshold at 16/18 end-to-end held-out episodes. Full R7 is still incomplete until two independent TypeScript raters submit the frozen blind ratings and disagreements are adjudicated. The readiness ledger, cockpit, live adapter, and human pilot remain gated. - No retention, takeover, productivity, or usability target has been measured. - The first valid build is one test-backed vertical slice, not a full Cursor clone. -- R0–R4.5 remain the only completed product evidence path. ADR-003's sandbox implementation gate has passed, but arbitrary participant or corpus execution remains blocked until the preregistered collection, eligibility, and freeze artifacts exist. +- R0–R4.5 remain the only completed product mechanism path. The R7 corpus infrastructure can execute the frozen supported patches, but that audit is technical evidence and not a usable general-project takeover product. +- ADR-007 proposes an Executable Operator Model so episodes update a durable map of demonstrated control instead of remaining isolated exercises. It is architecture, not implemented evidence. +- ADR-008 proposes an Intent Ledger that reconciles every changed line to supported, claimed, stale, contradicted, or explicitly unattributed semantic units. It is also architecture, not implemented evidence. +- ADR-009 proposes Decision Futures so a bounded pre-reveal human commitment can determine a live integrated branch while agents implement alternatives and continue unrelated work. It is architecture, not implemented evidence. ## Architecture shorthand ```text production agent run → observable Flight Recorder -→ high-value changed seam -→ disposable Takeover Twin -→ human prediction / diagnosis / intervention / recovery +→ bidirectional Intent Ledger / total change account +→ high-value changed seam or attribution gap +→ controllability surface + operator-model delta +→ Decision Future or disposable Takeover Twin / context-starved relay +→ human prediction / evidence choice / direction / recovery → executable judge -→ local delayed readiness evidence -→ future attention and delegation policy +→ reusable control artifact + local delayed readiness evidence +→ Takeover Envelope + future attention and delegation policy ``` The production plane continues by default while the readiness plane runs. If the human does no control work, PureFlow makes no claim that capability was preserved. diff --git a/docs/v0.3/RESEARCH.md b/docs/v0.3/RESEARCH.md index 236e78d..dba74dc 100644 --- a/docs/v0.3/RESEARCH.md +++ b/docs/v0.3/RESEARCH.md @@ -57,6 +57,14 @@ The [Onnasch et al. meta-analysis](https://doi.org/10.1177/0018720813501549) syn This supports the central trade-off but does not determine the correct IDE interaction. +### Negotiated initiative and interruption cost — 2026-08-01 + +Classic [mixed-initiative interaction](https://www.microsoft.com/en-us/research/publication/mixed-initiative-interaction/) rejects a binary choice between total automation and constant user control. It instead assigns initiative to the human or machine at the moment each contribution is valuable. NASA's [human-autonomy teaming](https://ntrs.nasa.gov/api/citations/20170010163/downloads/20170010163.pdf) work similarly identifies negotiated decisions and greater interaction as responses to brittle automation and out-of-the-loop situation-awareness loss. + +Interaction is not free. [Collaborative interruption research](https://www.microsoft.com/en-us/research/publication/modeling-information-exchange-opportunities-for-effective-human-computer-teamwork/) models when information from a person is valuable enough to justify interruption, while [task-interruption field work](https://www.microsoft.com/en-us/research/publication/conversations-amidst-computing-a-study-of-interruptions-and-recovery-of-task-activity/) documents the resumption cost after attention switches. Recent [interviews with science journalists](https://www.microsoft.com/en-us/research/publication/helping-me-versus-doing-it-for-me-designing-for-agency-in-llm-infused-writing-tools-for-science-journalism/) found that automating information gathering or feedback could preserve agency when core editorial judgment remained human, whereas automating core ideation and drafting was perceived as threatening autonomy and skill development. + +The design implication is narrower than “keep the human in the loop.” PureFlow should transfer initiative only at real, high-leverage decision forks; execute alternatives speculatively; offer the choice at a low-cost task boundary; and record whether the decision actually affected integration. ADR-009 calls this a Decision Future. The cited work motivates the mechanism but does not validate it for programming. + ## 2. Learning: active operation is different from reading ### Generation and self-explanation @@ -186,6 +194,17 @@ This comparison leads to ADR-006: spike Codex App Server first over local stdio, The research refresh also found a naming collision. OpenAI now uses [Chronicle](https://learn.chatgpt.com/docs/customization/chronicle) for screen-derived Codex memory. ADR-005 therefore renames PureFlow's observable run ledger to **Flight Recorder** before R1 data or a public wire format ships. +### Representation and traceability prior art — 2026-08-01 + +The artifact-accountability problem has credible neighbors, so PureFlow cannot claim novelty from adding an intent graph alone. + +- Microsoft Research's [Programming with Representations](https://www.microsoft.com/en-us/research/project/pwr/) puts a domain-specific representation between natural-language intent and generated code. Its published goal includes reducing the coding expertise needed to validate the result. This supports the value of an intermediate representation, but its direction is different from PureFlow's goal of preserving professional control in arbitrary existing repositories. +- [ReqToCode](https://arxiv.org/abs/2603.13999) embeds bidirectional requirement links in code and validates them during the build. It shows why traceability should fail visibly as artifacts evolve rather than live in detached documentation. It does not supply live agent-event provenance or behavioral evidence that a human can take over. +- Necula's [Proof-Carrying Code](https://doi.org/10.1145/263699.263712) lets a consumer validate that untrusted code satisfies a defined safety policy. PureFlow must not borrow the word “proof” for ordinary tests, traces, or agent claims; its proposed evidence-carrying generation is an explicitly weaker accountability mechanism. +- Code summaries and code-to-text representations may improve navigation, but model-generated prose can share the generator's original error. A reverse account needs structural and executable authority plus an honest unsupported state. + +ADR-008 therefore narrows the contribution to a combined mechanism: total changed-line reconciliation, untrusted agent claim references, executable evidence joins, visible attribution debt, and a separate Operator Model that only human actions can refresh. That combination remains a hypothesis until held-out expert attribution and delayed-transfer studies pass. + ## 5. Adjacent precedent: operational drills Reliability engineering already treats human response as something to exercise: @@ -198,19 +217,21 @@ These are important precedents for takeover muscle. They are periodic team event ## 6. Differentiation statement -In the reviewed public landscape, no system was found that documents all five properties together: +In the reviewed public landscape, no system was found that documents all seven properties together: 1. the production agent swarm continues without waiting for training; -2. a semantic compiler selects a causal seam from the current real change; -3. the human receives a concurrent executable fault or counterfactual twin; -4. tests and runtime evidence judge the action rather than an LLM response alone; -5. delayed adjacent transfer updates a module-scoped readiness model that influences future delegation. +2. every changed line reconciles to a bounded semantic unit and an honest attribution state; +3. a semantic compiler selects a causal seam from the current real change; +4. the human receives a concurrent executable fault or counterfactual twin; +5. tests and runtime evidence judge the action rather than an LLM response alone; +6. an on-time pre-reveal human decision can determine the live integrated branch while agents retain implementation; +7. delayed adjacent transfer updates a module-scoped readiness model that influences future delegation. This is meaningful differentiation within the reviewed landscape, not proof of absolute global or patent novelty. The defensible product core is therefore the combination: -> semantic scenario compiler + attention scheduler + executable evidence judge + longitudinal takeover model +> bidirectional Intent Ledger + Decision Futures + semantic scenario compiler + interruption-aware scheduler + executable evidence judge + longitudinal Takeover Envelope If PureFlow collapses back to explanations, questions, code tours, manual `TODO`s, or isolated bug games, it enters an already occupied category. diff --git a/docs/v0.3/THESIS.md b/docs/v0.3/THESIS.md index 50b68e6..38f37e0 100644 --- a/docs/v0.3/THESIS.md +++ b/docs/v0.3/THESIS.md @@ -83,6 +83,12 @@ Retrieval questions can improve retention. They usually exercise recall or recog Approval is cheap to rubber-stamp, especially while several agents are running. A real design fork becomes useful only when the developer predicts trade-offs and later sees evidence from the chosen or counterfactual path. +### Decision Futures + +PureFlow converts a small number of genuine live forks into bounded Decision Futures. Agents implement alternatives speculatively and continue unrelated work. Before decisive evidence is revealed, the developer predicts a consequence, chooses a discriminating observation, and may determine the branch that is actually integrated. A late or skipped choice never receives retroactive live-influence credit. + +This is the missing link between practice and ownership: the human makes the causal or architectural decision while the AI retains implementation work. Shadow control remains necessary for recovery practice, but it cannot substitute for all live authority. + ### Manual coding quotas A fixed quota puts human work on the critical path and can allocate it to trivial glue. One five-minute diagnosis of a high-blast-radius invariant may create more useful ownership than manually typing hundreds of predictable lines. @@ -141,6 +147,8 @@ Parallelism preserves throughput. It does not cause learning by itself. The epis Every changed line is either covered by a supported semantic unit with available provenance or visibly unattributed. PureFlow must not infer a convincing invariant for a line when the build plane and executable evidence do not provide one. +ADR-008 makes this concrete as a bidirectional Intent Ledger. Agent-emitted intent links are untrusted claims; the ledger independently reconciles every changed line to a bounded semantic unit and labels the link `supported`, `claimed`, `unattributed`, `contradicted`, or `stale`. Derived artifacts may bind to a deterministic generation receipt. This creates navigable accountability without claiming that the developer read or memorized every token. + ### Readiness of the operator For a critical seam, the developer can: @@ -152,6 +160,8 @@ For a critical seam, the developer can: - recover a failure; - transfer the model to an adjacent task later. +The Takeover Envelope lists where those claims remain current and contrasts them with the Autonomy Envelope where agents can execute. Divergence is visible by critical seam; it is not hidden behind one readiness percentage. + The second property requires behavior, not self-report. ## What would falsify the thesis diff --git a/prd.md b/prd.md index 5ba8331..68c0c03 100644 --- a/prd.md +++ b/prd.md @@ -12,12 +12,13 @@ PureFlow lets an agent swarm build at full speed while compiling the same work i AI code generation is not the feature to remove. It is the production engine. -The missing product is a second engine that maintains the human operator while the first engine maintains the software. Every meaningful agent run should be capable of producing two outputs: +The missing product is a second engine that maintains the human operator while the first engine maintains the software. Every meaningful agent run should be capable of producing three outputs: 1. a tested software change; -2. evidence that its human owner can take over a critical part of that change. +2. an executable `observe → actuate → recover` surface for a critical seam; +3. evidence that its human owner can take over that seam. -The second output is not a summary, a diff, a quiz score, or a quota of manually typed lines. It is an executable experience drawn from the real change: predict behavior, choose a control point, diagnose from evidence, intervene, observe the result, and recover. +The control and readiness outputs are not summaries, diffs, quiz scores, or quotas of manually typed lines. They are drawn from the real change: predict behavior, choose a control point, diagnose from evidence, intervene, observe the result, and recover. ## The honest constraint @@ -94,6 +95,8 @@ The Experience Compiler builds a semantic change graph and selects one high-info - expected future takeover value; - the user's current attention budget. +The Controllability Compiler attempts to bind that seam to an executable observation, a bounded actuator, and a recovery judge. The Operator Model Compiler then compares the new checkpoint with causal claims the developer previously demonstrated. Prompts are selected from model divergence, contradiction, or uncovered recovery routes—not from arbitrary changed functions. + ### 4. Drive While the production swarm continues, PureFlow opens a disposable **Takeover Twin** of that seam. The developer gets one short control episode without seeing the finished answer: @@ -112,6 +115,14 @@ Control episodes have two classes: This distinction prevents the product from becoming a simulated game. Some human decisions must genuinely steer the software, while lower-frequency recovery practice can stay entirely off the critical path. +### Decision Futures + +Live steering uses a versioned **Decision Future**, not plan approval. When agents encounter a real high-leverage fork, they speculatively implement viable alternatives and continue unrelated work. PureFlow freezes the autonomous default, evidence catalog, and bounded integration deadline before asking the developer to predict a consequence, select a discriminating observation, or choose a path. + +An on-time commitment can determine the integrated branch without requiring the developer to write its implementation. A skipped or late commitment leaves the autonomous policy in control; a late choice may still run as a counterfactual, but never receives retroactive live-influence credit. False forks, cosmetic alternatives, and post-reveal rubber stamps produce no decision evidence. + +Decision Futures maintain a **Takeover Envelope** beside the agent Autonomy Envelope. The former lists critical seams with current human control and transfer evidence; the latter lists seams the agents can operate. PureFlow highlights their divergence rather than collapsing it into a developer score or restricting full agentic coding. ADR-009 defines the protocol, authority, timing, and falsifiers. + ### 5. Judge The episode runs against executable evidence. PureFlow records what the developer inspected, predicted, changed, and recovered. An LLM may coach after the attempt, but it is not the sole grader. @@ -178,7 +189,7 @@ States decay with time and meaningful code changes. The UI shows evidence and ag ## Line-level accountability -The Flight Recorder groups generated lines into supported semantic units. Each unit may link to: +The Flight Recorder and Intent Ledger group generated lines into supported semantic units. Each unit may link to: - the user intent or requirement it serves; - the invariant or public behavior it changes; @@ -186,7 +197,21 @@ The Flight Recorder groups generated lines into supported semantic units. Each u - the tests, traces, benchmarks, or receipts that cover it; - unresolved assumptions or evidence gaps. -Agent adapters cannot reliably infer an invariant or intent for every line from Git history alone. A build plane may emit structured claim IDs to improve attribution; otherwise PureFlow shows unattributed lines as a gap rather than inventing a causal story. This is **evidence-carrying generation**, not proof of correctness. The graph makes a large patch navigable and supplies raw material for the Experience Compiler; it does not pretend that a graph teaches the human by itself. +Agent adapters cannot reliably infer an invariant or intent for every line from Git history alone. A build plane may emit structured claim IDs to improve attribution, but its own references begin as `claimed`. Only a deterministic structural join with allowed evidence can mark an attribution `supported`; contradictions, stale links, and unattributed units remain visible. Mechanical output may be covered by a reproducible generator receipt instead of thousands of fake decisions. + +This is **evidence-carrying generation**, not proof of correctness. The Intent Ledger reconciles every changed line, makes a large patch navigable, and supplies raw material for the Experience Compiler; it does not pretend that a graph teaches the human by itself. The bidirectional flow is `committed intent → agent obligations → actual patch/evidence → intent delta`, so the system can expose when generated behavior drifted from the account rather than merely produce another summary. + +The full decision, integrity rules, and falsifiable evaluation are specified in `docs/v0.3/ADR-008-EVIDENCE-CARRYING-GENERATION.md`. Like the Operator Model, it remains a post-R7 hypothesis. + +## Executable Operator Model + +PureFlow maintains a second, local versioned view beside the program: the causal claims the developer has actually demonstrated for this project. + +It is not an AI-generated repository summary. A claim can become fresh only through a pre-reveal prediction, evidence choice, intervention, recovery, or delayed transfer bound to executable evidence. Claims become stale when their semantic unit or control surface changes and contradicted when observed behavior disproves the committed prediction. Unknown and unattributed areas remain visible gaps. + +This gives the IDE a concrete answer to “what should the developer think about while agents keep coding?” It chooses the smallest action that reduces divergence between the running software and the operator's demonstrated model. A successful action should create a control dividend such as a regression probe, observation recipe, rollback path, or reusable recovery handle. + +The Operator Model is specified in `docs/v0.3/ADR-007-EXECUTABLE-OPERATOR-MODEL.md`. Decision Futures and the Takeover Envelope are specified in `docs/v0.3/ADR-009-DECISION-FUTURES.md`. Both remain post-R7 hypotheses until the complete expert gate and human experiments pass. ## Autonomy Router @@ -275,7 +300,7 @@ The R&D architecture is: - VSCodium plus a bundled extension for the product surface; - adapters over existing coding-agent runtimes for the build plane; -- PureFlow-owned Flight Recorder, Experience Compiler, Takeover Twin, Evidence Judge, Readiness Map, and Autonomy Router. +- PureFlow-owned Flight Recorder, Intent Ledger, Controllability and Experience Compilers, Decision Futures, Executable Operator Model, Takeover Envelope, Takeover Twin, Evidence Judge, Readiness Map, and Autonomy Router. This keeps the differentiating layer portable and avoids tying the thesis to a commercial editor fork. A deeper editor fork is justified only if a validated control experience requires unsupported workbench primitives.