From ea655f53cf6b52fccd2ef8f692426b603c896acb Mon Sep 17 00:00:00 2001 From: Error Lover Date: Sat, 1 Aug 2026 13:02:17 +0300 Subject: [PATCH] docs: define combined human-control pilot --- docs/BUILD_LOG.md | 10 + docs/PROJECT_STATE.md | 5 +- docs/v0.3/AGENT_EXECUTION.md | 13 ++ docs/v0.3/EXPERIMENTS.md | 53 +++++- docs/v0.3/GOAL_COMPLETION_AUDIT.md | 4 +- docs/v0.3/R8_COMBINED_PILOT_PROTOCOL.md | 236 ++++++++++++++++++++++++ docs/v0.3/README.md | 8 +- 7 files changed, 316 insertions(+), 13 deletions(-) create mode 100644 docs/v0.3/R8_COMBINED_PILOT_PROTOCOL.md diff --git a/docs/BUILD_LOG.md b/docs/BUILD_LOG.md index 69d22f8..687dfee 100644 --- a/docs/BUILD_LOG.md +++ b/docs/BUILD_LOG.md @@ -383,3 +383,13 @@ Evidence: `docs/v0.3/ADR-008-EVIDENCE-CARRYING-GENERATION.md`; updated PRD, thes - Added a three-arm ablation against autonomous explanation and shadow-only prediction. No product-effectiveness, live-influence, or skill-retention result is claimed. Evidence: `docs/v0.3/ADR-009-DECISION-FUTURES.md`; updated PRD, thesis, research, project state, completion audit, and v0.3 index. Implementation remains gated by the complete two-rater R7 expert audit. + +## 2026-08-01 — Combined R8 human-control protocol + +- Replaced the outdated three-condition delayed-transfer design with a draft four-condition protocol: Autonomous, fixed Active comparator, Shadow control, and Full PureFlow with genuine Decision Futures. +- Mapped the original product goal to behavioral evidence for autonomy throughput, total change accountability, executable control, real integration authority, voluntary attention, and AI-off takeover. +- Defined an intention-to-treat primary contrast of Full PureFlow versus Active comparator, followed by frozen mechanism tests of Full versus Shadow and Takeover Envelope calibration against confidence/exposure baselines. +- Added delayed AI-off transfer, longer-delay and field boundaries, privacy/minimization rules, ordered analysis, event contracts, and kill criteria that remove mechanisms when they do not add behavioral value. +- Grounded the design in current coding-skill and human-agency research without treating immediate quizzes, surveys, interviews, or target thresholds as product evidence. + +Evidence: `docs/v0.3/R8_COMBINED_PILOT_PROTOCOL.md` and updated experiment/execution handoff documents. No participants were enrolled, no target was measured, and implementation remains gated by the complete R7 expert audit plus R5/R6 runtime. diff --git a/docs/PROJECT_STATE.md b/docs/PROJECT_STATE.md index 411d6d9..79d9ed3 100644 --- a/docs/PROJECT_STATE.md +++ b/docs/PROJECT_STATE.md @@ -36,6 +36,7 @@ Branch `codex/shadow-cockpit-rnd` resets the product R&D thesis around **Dual-Co - The preregistered R7 collector froze 30 eligible patches from six repositories after evaluating 457 bounded eligibility records. Manifest `a4ef6cbfa48c66cb9d384bcc2834ecbfae8ff08810abfd1863b395b8aa47d149` contains 12 development and 18 held-out patches; `docs/v0.3/results/R7_CORPUS_COLLECTION.md` reports repository and first-match exclusion counts. No compiler or human outcome influenced selection. - The frozen R7 automatic audit passed its preregistered automatic threshold: 17/18 held-out identities compiled and 16/18 were valid end-to-end. The frozen blind expert packet set and deterministic rating join exist, but two independent ratings and adjudication remain pending. Full R7 has not passed; R5/R6 stay gated. - ADR-007 proposes an Executable Operator Model and shadow-control protocol. ADR-008 adds a bidirectional Intent Ledger for artifact accountability. ADR-009 adds Decision Futures and a Takeover Envelope so an on-time pre-reveal human commitment can determine a live integrated branch while agents retain implementation. Together they cover artifact accountability, demonstrated control, and real decision authority; none is implementation evidence. +- `docs/v0.3/R8_COMBINED_PILOT_PROTOCOL.md` now defines the draft four-condition human study needed to test the combined architecture against ordinary autonomous use and a fixed active-review comparator. It has no participants or measured outcomes and cannot be frozen until full R7 and the R5/R6 runtime exist. - `GOAL_COMPLETION_AUDIT.md` maps the original product goal to current evidence. It explicitly records that the live agent plane, v0.3 cockpit, cold relay, delayed-transfer study, and skill-preservation claim remain incomplete. - The readiness ledger and v0.3 cockpit do not exist yet. R0–R4.5 remain a closed reviewed-fixture mechanism and do not execute arbitrary participant or workspace code. - No skill-retention or speed metric has been measured. Values in the PRD are predeclared R&D targets. @@ -155,9 +156,9 @@ No external input blocks the repository-owned fixture R0–R4.5 mechanism in `do ## Next ordered actions 1. Obtain and adjudicate two independent blind expert ratings for packet index `54d78382b3ddbe15cba1f8153275e8149d32ddaa5192163f99ca5f43d903e8fe`. -2. If full R7 passes, review ADR-007, ADR-008, and ADR-009 together, then freeze the Operator Model, Intent Ledger, Decision Future, and Takeover Envelope schemas, hash domains, projections, evidence authorities, timing policy, and selector ablations before implementation. +2. If full R7 passes, review ADR-007, ADR-008, ADR-009, and the R8 combined protocol together, then freeze the Operator Model, Intent Ledger, Decision Future, and Takeover Envelope schemas, hash domains, projections, evidence authorities, timing policy, selector ablations, and four-condition study contrasts before implementation. 3. Add the local readiness ledger and minimal cockpit only after the complete R7 expert gate passes. -4. Run the preregistered delayed-transfer pilot before making any skill-retention claim. +4. Freeze and run the four-condition delayed-transfer pilot only after the technical runtime can instantiate every condition; run the longitudinal field pilot before making a sustained skill-retention claim. ## Recent milestone commits diff --git a/docs/v0.3/AGENT_EXECUTION.md b/docs/v0.3/AGENT_EXECUTION.md index ad81d73..de4555e 100644 --- a/docs/v0.3/AGENT_EXECUTION.md +++ b/docs/v0.3/AGENT_EXECUTION.md @@ -373,6 +373,16 @@ Use the thresholds in `EXPERIMENTS.md`. If the compiler misses the gate, narrow Agents may prepare fixtures, instrumentation, recruitment copy, randomization code, and analysis notebooks. A real human study, consent, outcome labeling, and claims cannot be automated away. +Use [`R8_COMBINED_PILOT_PROTOCOL.md`](R8_COMBINED_PILOT_PROTOCOL.md) as the preregistration template. It compares Autonomous, Active comparator, Shadow control, and Full PureFlow so executable practice, live decision authority, accountability, and readiness prediction can be falsified separately. + +### Entry gate + +- complete the frozen R7 expert ratings and adjudication; +- implement one versioned R5/R6 runtime capable of all four conditions; +- freeze ADR-007–009 schemas, hash domains, evidence authorities, invalidation, and timing rules; +- pass technical fixtures for total changed-line reconciliation, operator-model invalidation, Decision Future integration integrity, and delayed-task isolation; +- freeze the protocol, analysis code, task pairs, exclusions, and artifact hashes before enrollment. + ### Required artifacts - preregistered hypotheses and exclusions; @@ -382,6 +392,8 @@ Agents may prepare fixtures, instrumentation, recruitment copy, randomization co - delayed adjacent-task oracle; - blinded scoring rubric; - raw-data minimization plan; +- frozen four-condition protocol and ordered mechanism contrasts; +- Intent Ledger, Operator Model, Decision Future, and Takeover Envelope event schemas; - result report separating pilot targets from observed values. ### Gate @@ -404,6 +416,7 @@ flowchart LR R7 --> R6["R6 Cockpit"] R5 --> R6 R7 --> R8["R8 Human pilot"] + R5 --> R8 R6 --> R8 ``` diff --git a/docs/v0.3/EXPERIMENTS.md b/docs/v0.3/EXPERIMENTS.md index 966ed8a..e654aa8 100644 --- a/docs/v0.3/EXPERIMENTS.md +++ b/docs/v0.3/EXPERIMENTS.md @@ -28,6 +28,18 @@ A narrow compiler can generate coherent, deterministic episodes from normal test An event-triggered Explain-to-Break pulse selected from a causally important seam will produce better delayed adjacent-task performance per minute of attention than asking the developer to explain a randomly selected function. +### H6 — Live decision authority adds value beyond shadow control + +A pre-reveal Decision Future that can determine a real integrated path will improve delayed takeover and calibrated engineering agency beyond an otherwise identical shadow-control episode, without breaching the production-speed or attention margins. + +### H7 — The Takeover Envelope predicts blackout performance + +Project-scoped, time-stamped control evidence will predict delayed AI-off takeover better than self-confidence, diff exposure, episode count, or immediate explanation quality. + +### H8 — Evidence-carrying generation makes large changes accountable + +The Intent Ledger will let developers locate the relevant intent, invariant, and evidence in generated changes faster than raw diffs or AI summaries, without false `supported` provenance. Navigation success is an accountability outcome, not proof of skill retention. + ## What does not count as success - more questions answered correctly immediately after generation; @@ -157,6 +169,8 @@ Each participant completes three short project-derived episodes: 2. diagnose and repair a change-derived fault; 3. direct an intervention through a cold agent that receives only requested evidence. +Across the set, include one Intent Ledger evidence-navigation task and one genuine Decision Future whose on-time pre-reveal commitment can alter the integrated path. These additions test interaction integrity only; this small pilot cannot establish H6–H8. + The production agent runs a separate real task concurrently to test interruption and attention switching. ### Measures @@ -206,6 +220,8 @@ H5 requires a separately powered confirmatory comparison of policy 1 versus poli ## Experiment 3 — Controlled delayed-transfer study +The complete preregistration template is [`R8_COMBINED_PILOT_PROTOCOL.md`](R8_COMBINED_PILOT_PROTOCOL.md). Freeze it before enrollment; this section is the decision summary. + ### Research design Run a randomized controlled study on unfamiliar but realistic modules. Use a power analysis after the pilot to choose sample size; do not present a small convenience sample as definitive. @@ -214,9 +230,10 @@ Run a randomized controlled study on unfamiliar but realistic modules. Use a pow 1. **Autonomous:** full agent execution plus normal result view. 2. **Active comparator:** autonomous agent plus a fixed 10-minute protocol containing the same diff, a standardized explanation, and three preregistered post-hoc questions. Do not substitute a manual seam after seeing results. -3. **Dual control:** autonomous agent plus a compiled prediction–diagnosis–recovery episode. +3. **Shadow control:** autonomous agent plus Intent Ledger navigation, an Operator Model snapshot, and a compiled prediction–diagnosis–recovery episode that cannot affect production. +4. **Full PureFlow:** the same shadow-control mechanisms plus one genuine Decision Future whose on-time pre-reveal commitment can determine the integrated production path. -All groups get the same agent model, task, time budget, repository state, and production tests. +All groups get the same agent model, task, time and token budgets, repository state, documentation access, and production tests. Active comparator, Shadow control, and Full PureFlow receive the same maximum active-attention budget. ### Phase A — Production task @@ -236,7 +253,7 @@ The task must share the underlying invariant or data path but not the exact prac ### Primary outcomes -- **Primary estimand:** intention-to-treat risk difference between Dual control and Active comparator in the proportion completing a regression-free adjacent change within 45 minutes; +- **Primary estimand:** intention-to-treat risk difference between Full PureFlow and Active comparator in the proportion completing a regression-free adjacent change within 45 minutes; - successful regression-free adjacent change in the Autonomous condition is exploratory; - time to first valid causal hypothesis; - fault localization; @@ -247,6 +264,9 @@ The task must share the underlying invariant or data path but not the exact prac - confidence calibration; - architecture explanation scored blind by experts; +- Intent Ledger navigation accuracy and false-provenance rate; +- live-influence, false-fork, skip, late, default, and speculative-compute rates; +- Takeover Envelope calibration against confidence and exposure baselines; - retention after an additional delay; - subjective workload and product preference. @@ -260,9 +280,19 @@ The task must share the underlying invariant or data path but not the exact prac - distinguish immediate performance from delayed transfer. - randomize before the production task and analyze participants in their assigned condition; - count missing primary outcomes as unsuccessful in the conservative primary analysis and report a preregistered missing-data sensitivity analysis; -- cap both Active comparator and Dual control at 10 active minutes during Phase A so attention, not just elapsed time, is comparable; +- cap Active comparator, Shadow control, and Full PureFlow at 10 active minutes during Phase A so attention, not just elapsed time, is comparable; - standardize the delay and task timeout above rather than selecting them post hoc. +### Ordered mechanism tests + +Run comparisons in this frozen order: + +1. Full PureFlow versus Active comparator tests the combined product claim; +2. Full PureFlow versus Shadow control tests whether real live authority adds value beyond matched executable practice; +3. the Takeover Envelope is compared with self-confidence, diff exposure, and episode-count baselines. + +The pilot estimates variance. Freeze the confirmatory sample size, minimum worthwhile Full-versus-Shadow effect, and predictive-calibration margin before enrollment. Do not select the best arm post hoc. + ### Mechanism ablation Before scale-up, run a separately powered or explicitly exploratory ablation across equivalent tasks: @@ -275,14 +305,14 @@ This tests whether prediction and evidence selection add transfer beyond executi ### Product gate -Proceed to a longitudinal field pilot only if dual control: +Proceed to a longitudinal field pilot only if Full PureFlow: - improves delayed adjacent-task success over the fixed Active comparator by at least 20 percentage points and the 95% confidence interval for the primary risk difference excludes zero; - does not increase median production critical-path time by more than 5% and p90 by more than 10%; - stays within a median 10-minute human attention budget; - outperforms the active comparator on behavior, not just confidence. -If only immediate recall improves, reclassify the feature as a tutor and reject the revolutionary-IDE claim. +If Full PureFlow and Shadow control are equivalent within the frozen worthwhile-effect margin, remove Decision Futures from the product core or retain them only as an optional agency feature. If the Takeover Envelope does not outperform simple confidence/exposure baselines, remove readiness and routing claims. If only immediate recall improves, reclassify the feature as a tutor and reject the revolutionary-IDE claim. ## Experiment 4 — Longitudinal field pilot @@ -339,6 +369,17 @@ Every generated episode must pass these checks before presentation: The pilot should capture event types, not raw private content by default: ```text +intent_ledger.opened +intent_ledger.unit_selected +intent_ledger.evidence_opened +operator_model.snapshot +takeover_envelope.snapshot +decision_future.offered +decision_future.committed +decision_future.integrated +decision_future.auto_defaulted +decision_future.skipped +decision_future.late experience.offered experience.started prediction.recorded diff --git a/docs/v0.3/GOAL_COMPLETION_AUDIT.md b/docs/v0.3/GOAL_COMPLETION_AUDIT.md index 69bca1e..82e5b06 100644 --- a/docs/v0.3/GOAL_COMPLETION_AUDIT.md +++ b/docs/v0.3/GOAL_COMPLETION_AUDIT.md @@ -34,7 +34,7 @@ The original goal is broader than a compiler benchmark or an explanatory extensi - live Codex/Claude/OpenCode adapter accessible from this checkout; - context-starved relay on arbitrary supported project code; - Decision Futures with real bounded integration authority and an honest Takeover Envelope; -- delayed-transfer human experiment; +- frozen four-condition delayed-transfer human experiment and participant evidence; - longitudinal evidence of skill preservation; - evidence that developers voluntarily keep a non-zero attention budget; - any honest basis for saying PureFlow already preserves skills. @@ -50,7 +50,7 @@ The active goal may be marked complete only when all of these are true: 5. local Operator Model state invalidates correctly as the code changes; 6. every changed line resolves through the local Intent Ledger to a bounded unit and an honest attribution state; 7. a delayed adjacent task demonstrates takeover without answer-generating AI; -8. a controlled study beats the active diff/explanation/question comparator while preserving production speed; +8. the Full PureFlow condition beats the active diff/explanation/question comparator while preserving production speed, and its planned comparison with matched Shadow control isolates whether live decision authority adds value; 9. the runtime, deletion/export, packaging, and protected CI pass on supported platforms. Green unit tests for any individual component are necessary evidence, not completion of this goal. diff --git a/docs/v0.3/R8_COMBINED_PILOT_PROTOCOL.md b/docs/v0.3/R8_COMBINED_PILOT_PROTOCOL.md new file mode 100644 index 0000000..4f27c8b --- /dev/null +++ b/docs/v0.3/R8_COMBINED_PILOT_PROTOCOL.md @@ -0,0 +1,236 @@ +# R8 combined human-control pilot protocol + +**Status:** Draft. Freeze before enrollment. No participant run is authorized until the complete R7 expert gate passes and the R5/R6 runtime implements the tested conditions. + +**Purpose:** test whether PureFlow can preserve autonomous coding speed while producing measurable project takeover ability, accountable understanding, and real engineering agency. + +This protocol is a preregistration template, not a result report. Target values are decision rules, not measured outcomes. + +## 1. Why a combined pilot is necessary + +The product claim is not that questions, explanations, or code reading feel useful. The claim is that a developer may delegate most implementation and still retain enough causal control to modify and recover the project when answer-generating AI is unavailable. + +Recent evidence motivates the problem but does not prove this solution: + +- Anthropic's randomized study of 52 mostly junior developers reported lower immediate mastery with AI assistance, with the largest gap in debugging. It used an immediate quiz and explicitly leaves long-term skill development unresolved. +- A 319-person Microsoft survey associates higher confidence in generative AI with less self-reported critical-thinking effort. It is observational self-report, not a behavioral coding outcome. +- Interviews with 20 science journalists suggest that automating execution while retaining judgment may preserve agency better than automating core decisions. That is a useful design hypothesis from another domain, not evidence for PureFlow. + +Therefore R8 measures delayed executable behavior. It does not use confidence, explanation quality, time in a diff, or an LLM grade as proof of understanding. + +## 2. Product requirement to evidence map + +| Product requirement | Phase-A evidence | Delayed evidence | Falsifier | +| --- | --- | --- | --- | +| Full agentic coding remains first-class | agent completes the production task; production critical-path time and correctness are recorded | none | PureFlow requires a manual-code quota or breaches the speed margin | +| Generated changes remain accountable | every changed line reconciles to an Intent Ledger unit and honest attribution state | participant can locate the relevant intent, invariant, and evidence without reading the whole diff | false `supported` provenance, hidden answer, or irreducible attention cost | +| The developer exercises engineering judgment | pre-reveal Decision Future commitment can determine the integrated path | adjacent decision and recovery are scored independently | commitment is decorative, post-reveal, or does not alter integration | +| Challenges exercise control rather than recall | prediction, discriminating observation, actuation, and recovery are executable | a different fault in the same invariant/data path must be resolved | only the practiced mutation or remembered wording improves | +| The developer can take over if AI disappears | Operator Model and Takeover Envelope are frozen before transfer | regression-free adjacent task with answer-generating AI disabled | better explanation without better behavior | +| Participation remains voluntary | skip, late, abandonment, and attention-budget events are first-class and carry no penalty | voluntary return is measured | coercive gating, negative score, or misleading readiness claim | +| Almost all code may still be written by AI | no manual-code quota in Phase A; cold relay may implement a human directive | Phase B deliberately requires direct human editing because it tests AI disappearance | the product claim depends on typing volume rather than control | + +“Every changed line is accountable” means deterministic reconciliation to a bounded semantic unit and an honest attribution state. It does not mean the human memorized every token. + +## 3. Study boundary and entry gate + +Enrollment may begin only after all of the following are true: + +1. the two independent R7 ratings and adjudication clear the frozen expert gate; +2. R5/R6 implement all four conditions from one versioned runtime; +3. ADR-007, ADR-008, and ADR-009 schemas, hash domains, evidence authorities, timing rules, and invalidation rules are frozen; +4. held-out fixtures pass hidden-answer, production-isolation, line-reconciliation, and deterministic-oracle checks; +5. the protocol, analysis code, exclusions, minimum worthwhile effects, task pairs, and raw artifact hashes are frozen before condition outcomes are opened. + +The pilot estimates feasibility and variance. A power analysis based on blinded pilot aggregates must freeze the confirmatory enrollment size. A small convenience sample cannot clear the product claim. + +## 4. Participants and assignment + +Recruit developers who already use an agentic coding tool for multi-file work. Record experience, TypeScript familiarity, agent-use frequency, and baseline task skill without treating seniority as a proxy for takeover ability. + +Randomize before Phase A, stratified by: + +- experience band; +- baseline AI reliance; +- baseline task-skill score. + +Randomize fixture/task-pair order with a preregistered balanced schedule. Participants remain in their assigned condition for intention-to-treat analysis. Outcome reviewers receive de-identified artifacts without condition labels. + +## 5. Four conditions + +All conditions use the same agent model and version, task prompt, repository state, tests, documentation access, total production window, and token budget. + +### A. Autonomous / Cursor-like + +Full agent execution with normal progress, result, and diff access. No forced review and no PureFlow control episode. This arm measures natural autonomy throughput and exploratory delayed takeover. + +### B. Active comparator + +The same autonomous agent plus a fixed ten-minute post-hoc protocol: the same diff, a standardized explanation, and three preregistered questions. This represents the strongest ordinary “AI writes, human reviews” alternative without a control surface. + +### C. Shadow control + +The same autonomous agent plus: + +- Intent Ledger navigation; +- an Executable Operator Model snapshot; +- one matched prediction, discriminating-observation, diagnosis, and recovery episode in a disposable twin. + +The participant's choice cannot alter the production branch. This isolates executable control practice from live authority. + +### D. Full PureFlow / Decision Futures + +Everything in Shadow control plus one genuine Decision Future. Before any outcome or autonomous preference is revealed, the system freezes comparable alternatives, an autonomous default, evidence catalog, deadline, and integration rule. An on-time participant commitment can determine the integrated production path. + +Late, skipped, defaulted, or counterfactual choices are recorded honestly and never upgraded to live-influence evidence. + +Conditions B–D receive the same maximum ten minutes of active attention. The system may continue unrelated agent work while C or D is active. Speculative compute and any merge delay are reported separately from human attention. + +## 6. Procedure + +### Phase 0 — Baseline + +Run a short unfamiliar-module pretest with no target answer. Record causal localization, task skill, AI reliance, and self-confidence. Do not teach the target invariant. + +### Phase A — Agentic production + +The participant directs the agent to implement a realistic multi-file feature. The technical oracle must accept the base, reject registered faults, and score the final patch without condition-dependent tests. + +Record: + +- regression-free completion and production critical-path time; +- total wall time, agent time, human active time, and interruption recovery; +- changed-line reconciliation and attribution states; +- Decision Future offer, commitment, deadline, default, and actual merge effect; +- speculative compute, discarded alternatives, and branch confusion; +- skip, late, hint, abandonment, and assistance behavior. + +### Phase B — Delayed AI-off takeover + +After 48 hours, within a preregistered ±4-hour window: + +- hide the generation transcript and Phase-A explanations; +- disable answer-generating and code-generating AI; +- allow documentation, syntax reference, tests, debugger, and ordinary IDE navigation; +- give 45 minutes for a regression-free adjacent task; +- require direct human diagnosis and editing because this phase models AI disappearance. + +The transfer task shares an invariant, dependency seam, or data path with Phase A but uses a different failure family and cannot be solved by replaying the practiced mutation. + +### Phase C — Longer-delay subset + +Repeat an equivalent adjacent task after 7–14 days for a preregistered subset. The 48-hour result is delayed transfer, not long-term retention. A four- to eight-week field pilot remains necessary for a sustained skill-preservation claim. + +## 7. Outcomes and estimands + +### Primary outcome + +The primary estimand is the intention-to-treat risk difference between Full PureFlow and Active comparator in regression-free Phase-B completion within 45 minutes. + +The confirmatory product gate requires: + +- at least a 20 percentage-point improvement; +- a 95% confidence interval excluding zero; +- median Phase-A production critical-path slowdown no greater than 5%; +- p90 slowdown no greater than 10%; +- median active attention no greater than 10 minutes. + +### Ordered mechanism tests + +Run tests in the frozen order; do not promote a post-hoc winning contrast: + +1. Full PureFlow versus Active comparator: does the combined product improve takeover? +2. Full PureFlow versus Shadow control: does real live authority add value beyond identical shadow practice? +3. Takeover Envelope versus self-confidence, diff exposure, and episode count: does the readiness model predict Phase-B behavior better? + +The pilot must freeze the minimum worthwhile Full-versus-Shadow effect and the predictive calibration margin before confirmatory enrollment. + +### Secondary outcomes + +- time to first valid causal hypothesis; +- fault-localization accuracy; +- discriminating-evidence selection; +- recovery after an incorrect first move; +- blinded architecture explanation, reported separately from behavior; +- Intent Ledger navigation accuracy and time; +- `humanSelectionAffectedIntegration` rate; +- false-fork, late/default, skip, and abandonment rates; +- control-dividend reuse and stale-model detection; +- Brier score, calibration error, and discrimination for the Takeover Envelope; +- workload, preference, and voluntary next-episode budget. + +Intent Ledger navigation and explanation quality are not skill outcomes. Confidence is a comparator, not evidence. + +## 8. Analysis integrity + +- Freeze task pairs, oracles, assignments, analysis code, exclusions, and artifact hashes before unblinding. +- Count missing primary outcomes as unsuccessful in the conservative primary analysis and publish a frozen sensitivity analysis. +- Report all conditions, tested outcomes, effect sizes, confidence intervals, assistance use, and deviations. +- Keep correctness and behavioral evidence separate from prose ratings. +- Use independent blinded reviewers for qualitative artifacts and publish agreement plus adjudication. +- Treat executable oracles as outcome authority; an LLM may annotate prose but cannot certify success. +- Report exact mutation/fault-family transfer separately from adjacent-family transfer. + +## 9. Kill and narrowing rules + +| Result | Required decision | +| --- | --- | +| Full PureFlow misses the primary transfer or speed gate | reject the current revolutionary-IDE claim; do not compensate with confidence or explanation scores | +| Full and Shadow are equivalent within the frozen worthwhile-effect margin | remove Decision Futures from the product core or retain them only as an optional agency feature | +| Shadow does not beat Active on behavior | narrow or remove the Operator Model/control-episode mechanism | +| Takeover Envelope does not predict better than confidence/exposure baselines | remove readiness and routing claims; expose only raw evidence | +| Any held-out line is falsely labeled `supported` by invented or circular provenance | stop enrollment and repair Intent Ledger evidence authority | +| Benefits occur only for the practiced mutation or fault family | reject transfer and narrow to rehearsal tooling | +| Branch confusion, false forks, or live-choice integrity failures exceed the frozen tolerance | disable live integration authority until corrected | +| Voluntary use collapses after novelty | remove gamification/frequency pressure and reassess the mechanism before field expansion | +| Only immediate recall or explanation improves | classify the feature as a tutor, not a skill-preserving agentic IDE | + +Pilot observations set variance and tolerance values; tolerances cannot be chosen after confirmatory condition labels are opened. + +## 10. Instrumentation contract + +Export de-identified event envelopes, not raw repository content, by default: + +```text +intent_ledger.opened +intent_ledger.unit_selected +intent_ledger.evidence_opened +operator_model.snapshot +takeover_envelope.snapshot +decision_future.offered +decision_future.committed +decision_future.integrated +decision_future.auto_defaulted +decision_future.skipped +decision_future.late +experience.started +prediction.recorded +evidence.requested +intervention.applied +judge.completed +experience.abandoned +transfer.started +transfer.completed +``` + +Each event carries a local participant pseudonym, condition, timestamp, scoped capability, revision and schema digests, and bounded outcome fields. Raw code, prompts, prose, paths, and transcripts stay local unless the participant explicitly consents to a separately enumerated research export. + +## 11. Consent and non-coercion + +- Explain the AI-off phase, live-decision authority, data categories, withdrawal, and deletion before consent. +- Never attach the Takeover Envelope, streaks, or episode history to employer evaluation or a public leaderboard. +- Skipping an episode does not block agent execution, lower a public score, or imply incompetence. +- Keep project-scoped evidence time-stamped, deletable, and exportable by the participant. +- Report adverse events, accidental production edits, and integrity failures even when the final task passes. + +## 12. Claim boundary + +Passing this study would support a bounded claim that the tested PureFlow mechanism improved 48-hour adjacent-project takeover under the tested tasks and population while staying inside the specified production cost. It would not prove general intelligence, permanent skill retention, security correctness, or transfer across languages and domains. + +Only a later four- to eight-week field study on participants' own repositories can test whether the effect survives novelty and predicts real project takeover. + +## Research anchors + +- Anthropic, “How AI assistance impacts the formation of coding skills,” 2026: +- Lee et al., “The Impact of Generative AI on Critical Thinking,” CHI 2025: +- Nishal et al., “Helping Me Versus Doing It for Me,” CHI 2026: diff --git a/docs/v0.3/README.md b/docs/v0.3/README.md index fd1cc1e..f1fb3d0 100644 --- a/docs/v0.3/README.md +++ b/docs/v0.3/README.md @@ -23,9 +23,10 @@ PureFlow v0.3 asks whether an AI IDE can keep autonomous coding fast while behav 17. [`OPERATOR_MODEL_SPEC.md`](OPERATOR_MODEL_SPEC.md) — draft post-R7 vertical-slice requirements, contracts, and ablations. 18. [`ADR-008-EVIDENCE-CARRYING-GENERATION.md`](ADR-008-EVIDENCE-CARRYING-GENERATION.md) — proposed bidirectional Intent Ledger for total change accountability without passive full-diff review. 19. [`ADR-009-DECISION-FUTURES.md`](ADR-009-DECISION-FUTURES.md) — proposed speculative live-steering protocol and project-scoped Takeover Envelope. -20. [`GOAL_COMPLETION_AUDIT.md`](GOAL_COMPLETION_AUDIT.md) — requirement-by-requirement evidence separating infrastructure from the requested final product. -21. [`AGENT_EXECUTION.md`](AGENT_EXECUTION.md) — ordered implementation workstreams and acceptance gates. -22. [`JULES_LOOP.md`](JULES_LOOP.md) — guarded server-side execution queue for the audited R0–R4 slice. +20. [`R8_COMBINED_PILOT_PROTOCOL.md`](R8_COMBINED_PILOT_PROTOCOL.md) — draft four-condition human protocol for delayed takeover, live authority, accountability, and readiness calibration. +21. [`GOAL_COMPLETION_AUDIT.md`](GOAL_COMPLETION_AUDIT.md) — requirement-by-requirement evidence separating infrastructure from the requested final product. +22. [`AGENT_EXECUTION.md`](AGENT_EXECUTION.md) — ordered implementation workstreams and acceptance gates. +23. [`JULES_LOOP.md`](JULES_LOOP.md) — guarded server-side execution queue for the audited R0–R4 slice. ## Current truth @@ -38,6 +39,7 @@ PureFlow v0.3 asks whether an AI IDE can keep autonomous coding fast while behav - ADR-007 proposes an Executable Operator Model so episodes update a durable map of demonstrated control instead of remaining isolated exercises. It is architecture, not implemented evidence. - ADR-008 proposes an Intent Ledger that reconciles every changed line to supported, claimed, stale, contradicted, or explicitly unattributed semantic units. It is also architecture, not implemented evidence. - ADR-009 proposes Decision Futures so a bounded pre-reveal human commitment can determine a live integrated branch while agents implement alternatives and continue unrelated work. It is architecture, not implemented evidence. +- The draft R8 protocol now tests those mechanisms in four conditions. It is a preregistration template, not participant evidence, and remains blocked by full R7 plus the R5/R6 runtime. ## Architecture shorthand