diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index 2044038..a76993b 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -1,73 +1,37 @@ # Cross-project architecture -The authoritative conceptual topology is -[`program/THEORY_MAP.md`](program/THEORY_MAP.md). The machine-readable mapping is -[`program/theory-registry.json`](program/theory-registry.json). - -```text - PRIMARY DEVELOPER - | - v - AGENT GEARBOX - chooses WHERE work executes - / | \ - deterministic tool bounded helper optional isolated worker - \ | / - v - raw result/evidence - | - v - CONTEXT FIREWALL - chooses WHAT evidence returns - | - v - compact decision evidence - | - Decision Evidence conformance - | - Trajectory / Visible Value telemetry - | - v - PRIMARY DEVELOPER -``` - -Gearbox is a public narrow prototype with deterministic execution and an -injected bounded-helper transport contract. Context Firewall, Decision Evidence -Protocol, Agent Trajectory Profiler, Verifiable Agent Handoff, and Ephemeral -Agent Workers remain external, independently reusable mechanisms. - -Durable orchestration is a separate family: - -```text -Durable Supervisor (objective owner) - | - +-- Agent State Ledger (history/projection) - +-- Agent Scheduler Runtime (readiness/time/pause) - +-- Event-Driven Agent Wakeup (durable activation) - +-- Agent Discovery Control (new-work admission) - +-- Agent Recovery Policy (bounded recovery permission) -``` - -It can own autonomous progress across multiple activations without a -continuously active primary developer. Gearbox performs one bounded delegation -for a primary developer and returns one terminal result. - -Supporting boundaries: - -- Affected Verification independently decides which verification checks are - defensible for a change; humans, CI, or Gearbox may execute the plan, and - Context Firewall may subsequently reduce the resulting evidence. -- Gearbox admits deterministic versus cognitive work; Routing Policy selects an - eligible cognitive route after admission. -- Resource Claims establishes current concurrent ownership; Execution - Authorization validates permission for one exact action. -- Scheduler decides when durable work is ready; Routing decides where it may - run; Recovery decides whether another attempt is justified. -- Verifiable Handoff preserves exact result artifacts across source destruction; - Context Firewall determines which verified facts enter model context. -- Semantic Edit Protocol owns mutation semantics; Agent Trajectory Profiler - measures the resulting trajectory. - -No core primitive depends on a Taslos Tasks database, worker, scheduler, -package, path, or service. Future host-specific adapters remain optional and -versioned. +Opsle Tasks (`opsle/tasks`) owns current workload, task lifecycle and remote +execution. Its control plane and Incus execution targets are independently +deployed. This research repository records evidence and program state; it has no +runtime task-management authority. + +Tasks calls the versioned capability contract implemented for Task 15. Trusted +installed manifests declare hooks; operator grants bound repository selection at +the immutable task base. Capabilities supply independent public interfaces: + +- Gearbox chooses deterministic or cognitive routes; model-backed capability work + requires an execution-scoped, single-use gateway authorization. +- Context Firewall reduces command evidence, retaining canonical audit packets + separately from semantic model evidence and actual delivery measurements. +- Affected Verification supplies verification plans and exact change capture. + Incomplete evidence broadens verification or stops; bounded manifest-backed + integration does not promote research beyond OBSERVE/SHADOW. +- Visible Value validates receipts and produces operator summaries without + inventing missing values or causal savings. + +Decision Evidence Protocol remains standalone because independent consumers pin +its contract. Agent Trajectory Profiler and Semantic Edit Protocol retain their +independent evidence and research boundaries. Graphify is an optional standalone +CLI, not a Tasks capability. + +The final consolidation places ten source concepts under Tasks contracts and +Routing Policy under Gearbox. Source repositories are historical and cannot be +active dependencies. Durable Supervisor, Taslos Tasks, Paperclip and historical +agent-run are retired. Findings, source attribution, licenses and experiment +artifacts remain evidence; none supplies a current execution prerequisite. + +The [generated theory map](program/THEORY_MAP.md) maps each concept and records +its current home. The [historical architecture](program/history/pre-consolidation/ARCHITECTURE.md) +preserves the earlier hypotheses without reviving their repository boundaries. +`.github` housekeeping remains separate from workload eligibility. Deployment and +release authorization remain separate from implementation and research maturity. diff --git a/CONCEPTS.md b/CONCEPTS.md index a47ddeb..2ab161d 100644 --- a/CONCEPTS.md +++ b/CONCEPTS.md @@ -1,79 +1,27 @@ # Concept overview -This is the bootstrap narrative and subordinate-concept map. The authoritative -current repository inventory and lifecycle state are in -[`program/registry.json`](program/registry.json); the human-readable generated -view is [`PROGRAM_STATUS.md`](PROGRAM_STATUS.md). - -The bootstrap extraction list is not the canonical product/repository map. The -2026-08-29 source reconciliation is in -[`program/THEORY_MAP.md`](program/THEORY_MAP.md), with machine-readable concept -state in [`program/theory-registry.json`](program/theory-registry.json). It -records Agent Gearbox in its restored public home, formally separates it from -Durable Supervisor, and recommends future consolidation for several -durable-orchestration hypotheses without executing those dispositions. - -All current repositories are public, independently versioned, and part of Opsle -Research. That current versioning records provenance; it does not establish that -every hypothesis should remain a standalone implementation repository. Initial -SHA means the first public commit created during the 2026-08-25 bootstrap. - -| Repository | Maturity | Initial SHA | Implementation | Benchmark | -|---|---:|---|---|---| -| [gearbox](https://github.com/opsle/gearbox) | PROTOTYPE | `4d2a7cf902b52f099638b02a4fdec34fd5705a75` | provider-free bounded execution core | one deterministic dogfood fixture and plan; no comparative result | -| [agent-trajectory-profiler](https://github.com/opsle/agent-trajectory-profiler) | PROTOTYPE | `ce2fd532731d9bf5a0b7a271289bfdcc404f57c1` | sanitized dependency-free prototype | plan and fixtures only; no measured result | -| [semantic-edit-protocol](https://github.com/opsle/semantic-edit-protocol) | THEORY | `29caad5c03827cde17aabd71c38bc25899413a33` | theory/specification only | plan and fixtures only; no measured result | -| [durable-supervisor](https://github.com/opsle/durable-supervisor) | THEORY | `555ebedb992ac74236bb7da8230b4d6b0489830b` | theory/specification only | plan and fixtures only; no measured result | -| [event-driven-agent-wakeup](https://github.com/opsle/event-driven-agent-wakeup) | PROTOTYPE | `a6209860c2151450cc28ed648bc8c2631c8db7ef` | sanitized dependency-free prototype | plan and fixtures only; no measured result | -| [context-firewall](https://github.com/opsle/context-firewall) | THEORY | `e865644e86a3f820a120548e91284c048e515671` | theory/specification only | plan and fixtures only; no measured result | -| [decision-evidence-protocol](https://github.com/opsle/decision-evidence-protocol) | PROTOTYPE | `7050f83406a709da84e4b4770556319f767bbeaf` | sanitized dependency-free prototype | plan and fixtures only; no measured result | -| [agent-state-ledger](https://github.com/opsle/agent-state-ledger) | THEORY | `acab03b1ff7168222552050e21e7553b07d00e7c` | theory/specification only | plan and fixtures only; no measured result | -| [agent-scheduler-runtime](https://github.com/opsle/agent-scheduler-runtime) | THEORY | `d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0` | theory/specification only | plan and fixtures only; no measured result | -| [verifiable-agent-handoff](https://github.com/opsle/verifiable-agent-handoff) | PROTOTYPE | `399e5cfae94345affa3f087f0f6eb9e77669d33c` | sanitized dependency-free prototype | plan and fixtures only; no measured result | -| [agent-routing-policy](https://github.com/opsle/agent-routing-policy) | THEORY | `43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8` | theory/specification only | plan and fixtures only; no measured result | -| [agent-resource-claims](https://github.com/opsle/agent-resource-claims) | THEORY | `dfe0fbc90c67ce5ef4256354bb62f1d511b1304c` | theory/specification only | plan and fixtures only; no measured result | -| [agent-discovery-control](https://github.com/opsle/agent-discovery-control) | THEORY | `926547dd9fd713990b1d6f1f2e650aa6c0883564` | theory/specification only | plan and fixtures only; no measured result | -| [agent-execution-authorization](https://github.com/opsle/agent-execution-authorization) | THEORY | `8fb02e943c83fae121c63185d2f0d0dde8c4260a` | theory/specification only | plan and fixtures only; no measured result | -| [controlled-agent-acceptance](https://github.com/opsle/controlled-agent-acceptance) | THEORY | `2d652adf56e53953327d09b1ba9c4a9c3445f052` | theory/specification only | plan and fixtures only; no measured result | -| [agent-recovery-policy](https://github.com/opsle/agent-recovery-policy) | THEORY | `1b733a111e26e0a409fee3b96f627048531daefe` | theory/specification only | plan and fixtures only; no measured result | -| [ephemeral-agent-workers](https://github.com/opsle/ephemeral-agent-workers) | THEORY | `ad96fcfdfac06d340b5e96d369634980cee78ef4` | theory/specification only | plan and fixtures only; no measured result | - -## Ecosystem repositories - -| Repository | Purpose | Initial SHA | -|---|---|---| -| [opsle/site](https://github.com/opsle/site) | Source foundation for future opsle.com; not deployed | `85c3bc7c04f6a2774d589a65df008ad1f8837794` | -| [opsle/.github](https://github.com/opsle/.github) | Organization profile | Recorded in that repository | - -## Subordinate concepts - -These mechanisms remain inside parent projects until they show independent falsifiability, a reusable interface, and meaningful independent install/removal value. - -| Concept | Parent / current role | -|---|---| -| Mutation Amplification | agent-trajectory-profiler metric | -| Edit Payload Amplification | agent-trajectory-profiler metric | -| Semantic Region Revisit Rate | agent-trajectory-profiler metric | -| Change Intent Graph | semantic-edit-protocol mechanism | -| failure-only test reporting | context-firewall reducer | -| semantic Git adapter | decision-evidence/context-firewall adapter | -| semantic shell adapter | decision-evidence/context-firewall adapter | -| provider result envelope | decision-evidence-protocol adapter | -| evidence-gated completion | decision-evidence + state-ledger policy | -| already-satisfied proof | agent-discovery-control mechanism | -| Global Pause | agent-scheduler-runtime safety gate | -| provider cooldown registry | routing/scheduler evidence source | -| execution-target snapshots | authorization/routing binding | -| project configuration snapshots | state-ledger/authorization binding | -| independent reviewer routing | routing/controlled-acceptance mechanism | -| fresh-worker remediation | ephemeral-workers/recovery mechanism | -| objective graph planner | durable-supervisor/discovery mechanism | -| multi-agent write-region coordination | semantic-edit/resource-claims bridge | -| PLAN/DESIGN/BUILD/REVIEW/TEST pipeline | historical scheduler experiment, not final architecture | -| universal redaction | subordinate defense; never a substitute for authorization | -| provider availability countdown | routing UI projection | -| durable route-reason UI | routing evidence presentation | - -## Promotion rule - -Promote only with independent falsifiability, an independent reusable interface, and meaningful independent install/removal value. Rejected and superseded concepts remain recorded. +The [theory registry](program/theory-registry.json) and generated +[theory map](program/THEORY_MAP.md) define current concept state. The +[program ledger](program/registry.json) distinguishes current repositories, +consolidated sources, retired concepts and workload eligibility. + +There are 19 recorded concepts: seven active standalone concepts, eleven +consolidated concepts, and one retired concept. Eighteen live concepts share eight +current homes. Tasks is program infrastructure and the workload authority; +Visible Value is a standalone concept and receipt-validation CLI. + +Ten source concepts now live in Tasks contracts: discovery, execution +authorization, recovery, resource claims, scheduler, state ledger, controlled +acceptance, ephemeral workers, event wakeup and verifiable handoff. Routing Policy +is consolidated into Gearbox. Durable Supervisor is retired, with findings +preserved in Tasks and no current conceptual home or activation gate. + +Graphify's final role is an optional standalone CLI, not an Opsle Tasks capability. +Task 15's versioned capability contract integrates Gearbox, Context Firewall, +Affected Verification and Visible Value through independent public interfaces. + +The [historical bootstrap overview](program/history/pre-consolidation/CONCEPTS.md) +preserves initial SHAs, subordinate concepts and extraction provenance. Its +repository inventory and architectural recommendations are historical, not current +work. Consolidation preserves licenses, experiment identities and negative results; +it does not combine maturity stages or prove comparative benefit. diff --git a/MATURITY.md b/MATURITY.md index f9df7a2..dea6a13 100644 --- a/MATURITY.md +++ b/MATURITY.md @@ -18,4 +18,4 @@ Only these exact states are valid: | REJECTED | Evidence does not support the hypothesis or the cost/risk defeats it. | | SUPERSEDED | A later concept/version replaces it while preserving history. | -Existence inside Taslos Tasks never qualifies a concept as PROVEN. Maturity can move backward when evidence or scope changes. +Historical existence inside the retired Taslos Tasks never qualifies a concept as PROVEN. Maturity can move backward when evidence or scope changes. diff --git a/PROGRAM_STATUS.md b/PROGRAM_STATUS.md index 6713754..1cb08d6 100644 --- a/PROGRAM_STATUS.md +++ b/PROGRAM_STATUS.md @@ -2,15 +2,15 @@ -**Coverage: 21/21 expected repositories; duplicates: 0.** +**Coverage: 23 recorded repositories; 11 current, 12 historical; 10 workload eligible.** -Last verified: `2026-09-05T15:28:58Z`. HEADs are the verified default-branch revisions, not an assumption about later changes. +Last verified: `2026-09-09T00:00:00Z`. HEADs are the verified default-branch revisions, not an assumption about later changes. ## Program direction -> What prevents Durable Supervisor from successfully finishing Opsle Tasks with less intelligence, less context, less human involvement, and defensible evidence? +> How can current Opsle Tasks workloads use less intelligence and context while preserving correctness and defensible evidence? -Current lane: **NOW**. Durable Supervisor v0.1: **IN_PROGRESS** (0/10 stopping criteria satisfied). +Current lane: **NOW**. Workload and task-management authority: **Opsle Tasks**. Full generated priority view: `program/PRIORITY.md`. @@ -18,23 +18,23 @@ Full generated priority view: `program/PRIORITY.md`. | Lane | Objective | Repositories | Entry gate | |---|---|---|---| -| **NOW** | Finish Durable Supervisor v0.1 as a bounded measured system, then freeze feature work. | `durable-supervisor`, `research` | Current program lane. | -| **NEXT** | Use Opsle Tasks as the primary real-world workload and collect integrated measurements from the current implementation. | `gearbox`, `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `affected-verification` | Durable Supervisor v0.1 is declared and feature-frozen. | -| **THEN** | Advance an individual concept only when Durable Supervisor or Opsle Tasks evidence demonstrates a concrete need. | `semantic-edit-protocol`, `event-driven-agent-wakeup`, `agent-state-ledger`, `agent-scheduler-runtime`, `verifiable-agent-handoff`, `agent-routing-policy`, `agent-resource-claims`, `agent-discovery-control`, `agent-execution-authorization`, `controlled-agent-acceptance`, `agent-recovery-policy`, `ephemeral-agent-workers` | A qualifying work-item reason and real workload evidence identify the smallest relevant concept boundary. | -| **LATER** | Run controlled experiments, freeze a real-workload benchmark corpus, seek independent replication, and only then consider public product and research-site releases. | `site` | NOW, NEXT, and evidence-triggered THEN work establish a defensible need and release prerequisites are separately authorized. | -| **PARKED** | Retain useful non-priority ideas without turning them into active work. | `.github` | The idea is useful but lacks a qualifying reason to compete with the current objective. | +| **NOW** | Maintain the current Tasks workload and evidence-backed program state. | `tasks`, `research` | An authorized Tasks work item and a qualifying evidence-backed reason are required. | +| **NEXT** | Address demonstrated integration or measurement deficiencies from Tasks workloads. | `gearbox`, `context-firewall`, `visible-value`, `affected-verification`, `decision-evidence-protocol`, `agent-trajectory-profiler` | An authorized Tasks work item and a qualifying evidence-backed reason are required. | +| **THEN** | Investigate semantic edits only when workload evidence establishes a need. | `semantic-edit-protocol` | An authorized Tasks work item and a qualifying evidence-backed reason are required. | +| **LATER** | Prepare controlled research and public material when evidence and separate release authority support it. | `site` | An authorized Tasks work item and a qualifying evidence-backed reason are required. | +| **PARKED** | Retain organization housekeeping separately from workload eligibility. | `.github` | An authorized Tasks work item and a qualifying evidence-backed reason are required. | ### Exact next execution -In opsle/durable-supervisor, make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening Durable Supervisor v0.1. +Select the next authorized task from Opsle Tasks; use its project scope and workload evidence to identify the smallest justified change. This ledger does not create a parallel queue. -## Portfolio totals +## Highest evidenced research maturity (including historical sources) | Lifecycle stage | Count | |---|---:| | `THEORY` | 11 | | `SPECIFIED` | 0 | -| `PROTOTYPED` | 6 | +| `PROTOTYPED` | 8 | | `VERIFIED` | 4 | | `BENCHMARK_READY` | 0 | | `EXPERIMENTED` | 0 | @@ -42,37 +42,51 @@ In opsle/durable-supervisor, make the selected Gearbox route, routing rationale, | `DOCUMENTED` | 0 | | `COMPLETE` | 0 | -Program state totals: active 2; waiting 19; complete 0. +Program state totals: active 2; waiting 9; complete 0; historical 12. ## Repository inventory -| # | Repository | Type | HEAD | Stage | Evidence | Blocker | Next task | Dependencies | State | +| # | Repository | Type | HEAD | Stage | Evidence | Blocker | Next task | Dependencies | State / disposition | |---:|---|---|---|---|---|---|---|---|---| -| 1 | [agent-trajectory-profiler](https://github.com/opsle/agent-trajectory-profiler) | concept | `0a8966164072` | `VERIFIED` | [dependency-free JavaScript trajectory profiler with value-receipt ingestion, observational run records, class/unit/trust-safe per-run and cumulative summaries, Context Firewall adapter, canonical CLI/API, and content-addressed conformance corpus](https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/value-summary.js); 74 of 74 automated tests and 14 of 14 content-addressed measurement fixtures passed locally at the verified HEAD; 5 of 5 pinned exact-revision interoperability cases and 5 determinism tests passed; PR #2 CI passed | EXP-001 exact budgeted authorization, catalogue/pricing preflight, measured experiment, and independent replication remain missing; the one-block coordinator is provider-free qualified and zero model/provider subjects have run. | Wait for a Durable Supervisor or Opsle Tasks workload to require trajectory measurement; do not advance the profiler merely to polish it. | — | waiting | -| 2 | [semantic-edit-protocol](https://github.com/opsle/semantic-edit-protocol) | concept | `29caad5c0382` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/THEORY.md); placeholder only; no automated tests | No executable semantic operation, validator, tests, or benchmark fixtures. | Wait until a real Durable Supervisor or Opsle Tasks workload demonstrates a semantic-edit deficiency that simpler bounded edits cannot satisfy. | `agent-resource-claims`, `agent-trajectory-profiler` | waiting | -| 3 | [durable-supervisor](https://github.com/opsle/durable-supervisor) | concept | `1b5ab7631ba6` | `VERIFIED` | [dependency-free Node.js Durable Supervisor v0.1 with persistent supervisor authority, task handoff, discovery, exact Gearbox routing, claims and fencing, detached Runner ownership, Context Firewall packets, acceptance and immutable evaluation, opsled wake delivery, bounded reconstruction, policy controls, and runtime-release fencing](https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/README.md); PR #2 recorded 141 full-suite tests and canonical validation passing; PR #3 recorded 176 passing with 1 intentional skip; D5 PR #4 passed 6 focused atomicity, operator, and wake checks; projection-reconciliation PR #5 passed 76 focused evaluation, invariant, operator, recovery, and wake checks plus syntax, release-manifest, and canonical validation at the exact head whose tree is main | Durable Supervisor v0.1 still lacks the operator-visible measurement surface, avoidable-intelligence classification, smallest evidence-driven preflight, and another measured foreign-repository workload required by the program stopping criteria. | Make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening v0.1. | `agent-state-ledger`, `agent-scheduler-runtime`, `event-driven-agent-wakeup`, `decision-evidence-protocol` | active | -| 4 | [event-driven-agent-wakeup](https://github.com/opsle/event-driven-agent-wakeup) | concept | `a6209860c215` | `PROTOTYPED` | [dependency-free JavaScript state-machine prototype](https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/src/runtime.js); 2 of 2 automated tests passed locally at the verified HEAD | Current prototype is in-memory and has no restart or durable-store harness. | Wait for a real Durable Supervisor or Opsle Tasks wake deficiency; advance the standalone concept only if the integrated runtime exposes one. | `agent-state-ledger` | waiting | -| 5 | [context-firewall](https://github.com/opsle/context-firewall) | concept | `953c48f1cfd1` | `PROTOTYPED` | [dependency-free deterministic JavaScript TAP-subset reducer with compact evidence packets, separate opsle.value-receipt.v1 sidecars, named operator indicators, payload ceilings, raw escalation, and synthetic conformance corpus](https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/src/value-receipt.js); 42 of 42 automated tests and 30 of 30 synthetic conformance fixtures passed locally at the verified HEAD; deterministic stdout/sidecar separation and operator stderr tests passed; PR #2 CI passed | EXP-001 exact budgeted authorization, catalogue/pricing preflight, measured correctness experiment, and replication remain missing; the one-block coordinator is provider-free qualified and zero model/provider subjects have run. | Wait for Durable Supervisor or Opsle Tasks evidence of context overload, unsafe omission, or unsupported input before advancing the standalone reducer. | `decision-evidence-protocol`, `agent-trajectory-profiler` | waiting | -| 6 | [decision-evidence-protocol](https://github.com/opsle/decision-evidence-protocol) | concept | `b17ae3b41cea` | `VERIFIED` | [dependency-free generic value-receipt validator plus independent Context Firewall packet/source/value-receipt validator, Decision Evidence validation receipts, canonical CLI with named operator indicators, and self-contained conformance suite](https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/context-firewall-value.js); 60 of 60 automated tests and 24 of 24 public-safe conformance vectors passed locally at the verified HEAD; 5 of 5 exact-revision interoperability cases, 8 determinism tests, operator separation, source-unverified, and tamper paths passed; PR #2 CI passed | Additional real tool classes, measured subject decision adequacy, EXP-001 exact budgeted authorization and catalogue/pricing preflight, a controlled comparative result, and independent replication are missing. | Wait for an integrated Durable Supervisor or Opsle Tasks receipt to expose a decision-evidence gap that blocks a defensible decision. | — | waiting | -| 7 | [agent-state-ledger](https://github.com/opsle/agent-state-ledger) | concept | `acab03b1ff71` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/THEORY.md); placeholder only; no automated tests | No portable schema, implementation, projection oracle, or replay fixtures. | Wait for a real Durable Supervisor or Opsle Tasks workload to expose a portable durable-state deficiency before advancing this repository. | — | waiting | -| 8 | [agent-scheduler-runtime](https://github.com/opsle/agent-scheduler-runtime) | concept | `d97cd3c218b2` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/THEORY.md); placeholder only; no automated tests | Ledger and claim interfaces are not portable or executable in these repositories. | Wait for a real workload to expose a scheduler-runtime deficiency not already satisfied by Durable Supervisor's local implementation. | `agent-resource-claims`, `agent-state-ledger` | waiting | -| 9 | [verifiable-agent-handoff](https://github.com/opsle/verifiable-agent-handoff) | concept | `399e5cfae943` | `PROTOTYPED` | [dependency-free JavaScript HMAC manifest prototype](https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/src/seal.js); 3 of 3 automated tests passed locally at the verified HEAD | Prototype authenticates a manifest but does not build, transport, or reconstruct artifacts. | Wait for a real workload to require evidence survival across an isolation or source-destruction boundary before advancing this protocol. | `decision-evidence-protocol` | waiting | -| 10 | [agent-routing-policy](https://github.com/opsle/agent-routing-policy) | concept | `43fc2a72d2c8` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/THEORY.md); placeholder only; no automated tests | No normative route schema, evaluator, fixtures, or provider-independent quality evidence. | Wait for measured Durable Supervisor routing deficiencies before advancing a standalone routing policy. | — | waiting | -| 11 | [agent-resource-claims](https://github.com/opsle/agent-resource-claims) | concept | `dfe0fbc90c67` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/THEORY.md); placeholder only; no automated tests | No portable resource catalog, claim-set state machine, concurrency tests, or fairness fixtures. | Wait for a workload-exposed resource-claim or fencing deficiency not already covered by Durable Supervisor's local claims. | — | waiting | -| 12 | [agent-discovery-control](https://github.com/opsle/agent-discovery-control) | concept | `926547dd9fd7` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/THEORY.md); placeholder only; no automated tests | No durable proposal dataset, similarity oracle, policy evaluator, or convergence harness. | Wait for real duplicate-discovery or already-satisfied work evidence before advancing this repository. | `agent-state-ledger` | waiting | -| 13 | [agent-execution-authorization](https://github.com/opsle/agent-execution-authorization) | concept | `8fb02e943c83` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/THEORY.md); placeholder only; no automated tests | No grant schema, current-authority oracle, revocation model, or adversarial fixtures. | Wait for a real execution-authorization gap that blocks the current Durable Supervisor or Opsle Tasks objective. | `agent-resource-claims`, `agent-state-ledger` | waiting | -| 14 | [controlled-agent-acceptance](https://github.com/opsle/controlled-agent-acceptance) | concept | `2d652adf56e5` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/THEORY.md); placeholder only; no automated tests | Authorization, routing, scheduling, and handoff contracts are not yet executable together. | Wait for real acceptance evidence to show a missing portable capability before advancing the standalone concept. | `agent-execution-authorization`, `agent-routing-policy`, `agent-scheduler-runtime`, `verifiable-agent-handoff` | waiting | -| 15 | [agent-recovery-policy](https://github.com/opsle/agent-recovery-policy) | concept | `1b733a111e26` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md); placeholder only; no automated tests | No shared failure schema, attempt ledger, route evaluator, or comparative fixture set. | Wait for repeated real recovery failures to demonstrate a policy deficiency; do not create speculative recovery work. | `agent-routing-policy`, `agent-state-ledger`, `decision-evidence-protocol` | waiting | -| 16 | [ephemeral-agent-workers](https://github.com/opsle/ephemeral-agent-workers) | concept | `ad96fcfdfac0` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md); placeholder only; no automated tests | Portable authority, claim, and handoff contracts are not ready; no safe synthetic containment harness exists. | Wait for a real workload to require an ephemeral-worker boundary before advancing this repository. | `agent-execution-authorization`, `agent-resource-claims`, `verifiable-agent-handoff` | waiting | -| 17 | [gearbox](https://github.com/opsle/gearbox) | concept | `f3fab9f292cf` | `PROTOTYPED` | [provider-free Python reference core with strict authority-policy admission, exact deterministic argv execution, content-addressed staged helper context, injected one-shot helper transport, passive process waiting, compact results, raw-artifact accounting, fail-closed budgets, and Visible Value receipts](https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py); 19 of 19 provider-free automated tests passed locally, in PR #1 CI, and in final-main CI; ruff, shellcheck, actionlint, gitleaks, wheel build, receipt validation, and public raw-locator/hash checks passed | A production-quality bounded helper transport, independently verified isolation and termination, full Context Firewall integration, and a frozen comparative benchmark remain missing. | Support Durable Supervisor receipt visibility and Opsle Tasks workload measurement; advance the standalone Gearbox only if integrated evidence exposes a routing or preflight deficiency. | `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `agent-routing-policy`, `agent-execution-authorization` | waiting | -| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `97f490a67337` | `VERIFIED` | [dependency-free Node.js deterministic plan-v2 planner with check-level dependency-completeness states, mechanism and boundary evidence, check-local fail-closed forced selection, explainable skips, bounded deterministic Python boundary inspection, identity-bound SHADOW validation, frozen-oracle repair/replay harnesses, and opsle.value-receipt.v1 safety-cost telemetry](https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/REPORT.md); 107 of 107 automated tests, 15 of 15 conformance scenarios, and 10 of 10 determinism checks passed locally, in PR #4 CI, in exact final-main CI, and from a fresh detached worktree at 97f490a67337552fee25757266f3dc034660dca0; the AV-EXP-003 verifier and full deterministic repair reproduction also passed with identical result and regression-matrix identities | The bounded repair does not close opaque boundaries or constitute dynamic analysis; historical real-change replay, a production-quality evidence adapter, and independent qualifying replication remain missing. AV remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized. | Remain OBSERVE/SHADOW and wait for Durable Supervisor-driven Opsle Tasks work to expose a concrete verification-selection deficiency before adding another experiment. | — | waiting | -| 19 | [research](https://github.com/opsle/research) | program infrastructure | `e8a2c36678ac` | `PROTOTYPED` | [authoritative 21-repository portfolio and priority ledger, machine-readable 18-concept theory registry, five-experiment registry, lifecycle and anti-nitpick controls, normative Visible Value target, deterministic generated status and priority views, provider-free EXP-001 preparation, and recorded AV-EXP-001/002/003 shadow evidence with integrity CI](program/registry.json); 90 of 90 repository tests passed locally at the verified default-branch HEAD; the 21-repository and 18-concept registries, five experiment records, generated dashboard, and deterministic experiment evidence checks passed | Durable Supervisor v0.1 measurement and foreign-workload stopping criteria remain open; Opsle Tasks cannot become the primary workload until v0.1 is declared and frozen. | Keep the authoritative priority and portfolio views current while Durable Supervisor completes its ten bounded v0.1 stopping criteria. | — | active | -| 20 | [site](https://github.com/opsle/site) | program infrastructure | `28ad65be4750` | `PROTOTYPED` | [React/Vinext source implementation with content routes](https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md); automated build/render tests present; not rerun because this reconciliation kept other repositories read-only | Wait for validated registry data and measured research; deployment requires separate authorization. | Remain later until measured research and separate site-release authorization justify registry-derived public content. | `research` | waiting | -| 21 | [.github](https://github.com/opsle/.github) | program infrastructure | `01c38e726db7` | `THEORY` | [documentation-only organization profile](https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md); not applicable to current single Markdown profile; consistency is unverified | No mechanical registry consistency check exists in this repository. | Remain parked until a broken organization-profile link or external release condition creates a concrete need. | `research` | waiting | +| 1 | [agent-trajectory-profiler](https://github.com/opsle/agent-trajectory-profiler) | concept | `0a8966164072` | `VERIFIED` | [dependency-free JavaScript trajectory profiler with value-receipt ingestion, observational run records, class/unit/trust-safe per-run and cumulative summaries, Context Firewall adapter, canonical CLI/API, and content-addressed conformance corpus](https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/README.md); Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation. | EXP-001 is prepared but unconsumed; measured subject decision adequacy, broader workload coverage and independent replication remain absent. | Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository. | — | waiting / ACTIVE | +| 2 | [semantic-edit-protocol](https://github.com/opsle/semantic-edit-protocol) | concept | `29caad5c0382` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/README.md); Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation. | No executable semantic operation, validator, tests, or benchmark fixtures. | Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository. | — | waiting / ACTIVE | +| 3 | [durable-supervisor](https://github.com/opsle/durable-supervisor) | concept | `a69f3b1ce523` | `VERIFIED` | [Historical source preserved; runtime and repository retired; findings preserved in Tasks.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json); Historical observation before retirement: PR #2 recorded 141 full-suite tests and canonical validation passing; PR #3 recorded 176 passing with 1 intentional skip; D5 PR #4 passed 6 focused atomicity, operator, and wake checks; projection-reconciliation PR #5 passed 76 focused evaluation, invariant, operator, recovery, and wake checks plus syntax, release-manifest, and canonical validation at the exact head whose tree is main | None | None; historical source is retired from workload selection. | — | historical / RETIRED | +| 4 | [event-driven-agent-wakeup](https://github.com/opsle/event-driven-agent-wakeup) | concept | `8278bfb064eb` | `PROTOTYPED` | [Historical source preserved; concept contracts consolidated into tasks.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json); Historical observation before retirement: 2 of 2 automated tests passed locally at the verified HEAD | None | None; historical source is retired from workload selection. | — | historical / CONSOLIDATED → tasks | +| 5 | [context-firewall](https://github.com/opsle/context-firewall) | concept | `6dd6e5fdf21f` | `PROTOTYPED` | [v0.5.0 reducer for flat TAP and Node spec/dot reporters, semantic-only model projection, canonical audit packet and separate value receipt.](https://github.com/opsle/context-firewall/blob/6dd6e5fdf21f28dc5ebfa07954aaa9bed2dbcc32/README.md); Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation. | EXP-001 remains PLANNED with zero consumed authorizations and no subject runs; comparative correctness and replication remain absent. | Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository. | — | waiting / ACTIVE | +| 6 | [decision-evidence-protocol](https://github.com/opsle/decision-evidence-protocol) | concept | `b17ae3b41cea` | `VERIFIED` | [dependency-free generic value-receipt validator plus independent Context Firewall packet/source/value-receipt validator, Decision Evidence validation receipts, canonical CLI with named operator indicators, and self-contained conformance suite](https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/README.md); Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation. | EXP-001 is prepared but unconsumed; measured subject decision adequacy, broader workload coverage and independent replication remain absent. | Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository. | — | waiting / ACTIVE | +| 7 | [agent-state-ledger](https://github.com/opsle/agent-state-ledger) | concept | `509a2a55066f` | `THEORY` | [Historical source preserved; concept contracts consolidated into tasks.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json); Historical observation before retirement: placeholder only; no automated tests | None | None; historical source is retired from workload selection. | — | historical / CONSOLIDATED → tasks | +| 8 | [agent-scheduler-runtime](https://github.com/opsle/agent-scheduler-runtime) | concept | `261adff7792b` | `THEORY` | [Historical source preserved; concept contracts consolidated into tasks.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json); Historical observation before retirement: placeholder only; no automated tests | None | None; historical source is retired from workload selection. | — | historical / CONSOLIDATED → tasks | +| 9 | [verifiable-agent-handoff](https://github.com/opsle/verifiable-agent-handoff) | concept | `b832c770d996` | `PROTOTYPED` | [Historical source preserved; concept contracts consolidated into tasks.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json); Historical observation before retirement: 3 of 3 automated tests passed locally at the verified HEAD | None | None; historical source is retired from workload selection. | — | historical / CONSOLIDATED → tasks | +| 10 | [agent-routing-policy](https://github.com/opsle/agent-routing-policy) | concept | `179fb15add41` | `THEORY` | [Historical source preserved; concept contracts consolidated into gearbox.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json); Historical observation before retirement: placeholder only; no automated tests | None | None; historical source is retired from workload selection. | — | historical / CONSOLIDATED → gearbox | +| 11 | [agent-resource-claims](https://github.com/opsle/agent-resource-claims) | concept | `f80f82e79ed0` | `THEORY` | [Historical source preserved; concept contracts consolidated into tasks.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json); Historical observation before retirement: placeholder only; no automated tests | None | None; historical source is retired from workload selection. | — | historical / CONSOLIDATED → tasks | +| 12 | [agent-discovery-control](https://github.com/opsle/agent-discovery-control) | concept | `09c11f8a8ffd` | `THEORY` | [Historical source preserved; concept contracts consolidated into tasks.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json); Historical observation before retirement: placeholder only; no automated tests | None | None; historical source is retired from workload selection. | — | historical / CONSOLIDATED → tasks | +| 13 | [agent-execution-authorization](https://github.com/opsle/agent-execution-authorization) | concept | `6babce3929e0` | `THEORY` | [Historical source preserved; concept contracts consolidated into tasks.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json); Historical observation before retirement: placeholder only; no automated tests | None | None; historical source is retired from workload selection. | — | historical / CONSOLIDATED → tasks | +| 14 | [controlled-agent-acceptance](https://github.com/opsle/controlled-agent-acceptance) | concept | `258325416c85` | `THEORY` | [Historical source preserved; concept contracts consolidated into tasks.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json); Historical observation before retirement: placeholder only; no automated tests | None | None; historical source is retired from workload selection. | — | historical / CONSOLIDATED → tasks | +| 15 | [agent-recovery-policy](https://github.com/opsle/agent-recovery-policy) | concept | `a529954c784a` | `THEORY` | [Historical source preserved; concept contracts consolidated into tasks.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json); Historical observation before retirement: placeholder only; no automated tests | None | None; historical source is retired from workload selection. | — | historical / CONSOLIDATED → tasks | +| 16 | [ephemeral-agent-workers](https://github.com/opsle/ephemeral-agent-workers) | concept | `e758697528e7` | `THEORY` | [Historical source preserved; concept contracts consolidated into tasks.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json); Historical observation before retirement: placeholder only; no automated tests | None | None; historical source is retired from workload selection. | — | historical / CONSOLIDATED → tasks | +| 17 | [gearbox](https://github.com/opsle/gearbox) | concept | `12112c3a04d7` | `PROTOTYPED` | [Provider-free bounded execution core and deterministic Tasks phase routing; Agent Routing Policy contracts consolidated with source provenance.](https://github.com/opsle/gearbox/blob/12112c3a04d72b768f3252a6d0f6ea455816c9a1/README.md); Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation. | General helper transport, independent isolation evidence and comparative benchmark remain separate research work. | Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository. | — | waiting / ACTIVE | +| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `792c4bb7881f` | `VERIFIED` | [dependency-free Node.js deterministic plan-v2 planner with check-level dependency-completeness states, mechanism and boundary evidence, check-local fail-closed forced selection, explainable skips, bounded deterministic Python boundary inspection, identity-bound SHADOW validation, frozen-oracle repair/replay harnesses, and opsle.value-receipt.v1 safety-cost telemetry; Tasks verification planning and capture contract integrated.](https://github.com/opsle/affected-verification/blob/792c4bb7881f6f430b2b1ba238ba05be43f50d38/README.md); Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation. | Research remains OBSERVE/SHADOW; AV-EXP-002 failure is permanent. Tasks bounded manifest-backed selection requires complete impact, catalog and check-boundary evidence; incomplete evidence broadens or stops. No general TRUSTED_BOUNDED promotion. | Keep OBSERVE/SHADOW research limits; admit a new calibration or repair only from a demonstrated Tasks verification deficiency. | — | waiting / ACTIVE | +| 19 | [research](https://github.com/opsle/research) | program infrastructure | `840c79c04199` | `PROTOTYPED` | [Authoritative current/historical research ledger, concept mappings, experiment records and deterministic generated views; Tasks owns workload management.](https://github.com/opsle/research/blob/840c79c04199762acdd59e89e807ca9086f3e45e/README.md); Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation. | Research completion requires controlled evidence and replication; integration completion is insufficient. | Select the next authorized task from Opsle Tasks; use its project scope and workload evidence to identify the smallest justified change. This ledger does not create a parallel queue. | — | active / ACTIVE | +| 20 | [site](https://github.com/opsle/site) | program infrastructure | `29c94e6440ec` | `PROTOTYPED` | [React/Vinext source implementation with content routes](https://github.com/opsle/site/blob/29c94e6440ec91f5e09aded959fe7144187af6d7/README.md); Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation. | Wait for validated registry data and measured research; deployment requires separate authorization. | Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository. | — | waiting / ACTIVE | +| 21 | [.github](https://github.com/opsle/.github) | program infrastructure | `01c38e726db7` | `THEORY` | [documentation-only organization profile](https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/README.md); Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation. | No mechanical registry consistency check exists in this repository. | Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository. | — | waiting / ACTIVE | +| 22 | [tasks](https://github.com/opsle/tasks) | program infrastructure | `e1207c5264c5` | `PROTOTYPED` | [Current workload lifecycle, task management and remote execution substrate with plug-and-play capability contract.](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/README.md); Implementation and boundary-test sources inspected at recorded HEAD; upstream proof receipts are observational evidence. Final research verification belongs to the pipeline. | Controlled comparative evidence and independent replication remain absent. | Use Tasks workload evidence to admit only the smallest justified deficiency repair. | `gearbox`, `context-firewall`, `affected-verification`, `visible-value` | active / ACTIVE | +| 23 | [visible-value](https://github.com/opsle/visible-value) | concept | `f011395cae86` | `PROTOTYPED` | [Independent receipt validation and operator summaries preserving evidence classes and missing measurements.](https://github.com/opsle/visible-value/blob/f011395cae86a7a736b424d658bbe3afe65ba870/README.md); Implementation and boundary-test sources inspected at recorded HEAD; upstream proof receipts are observational evidence. Final research verification belongs to the pipeline. | Controlled comparative evidence and independent replication remain absent. | Use Tasks workload evidence to admit only the smallest justified deficiency repair. | — | waiting / ACTIVE | + +## Completed integration and retirement work + +- **taslos-tasks-retirement — COMPLETED**: Retired execution target; historical extraction provenance retained. [evidence 1](https://github.com/opsle/tasks/commit/9915dd4a7f0a4bf73c62b2310523b1429f46d9b1) +- **paperclip-retirement — COMPLETED**: Retired orchestration references removed; no current dependency. [evidence 1](https://github.com/opsle/tasks/commit/2d06527f583e49dbe75adb84051806812712e920) +- **repository-consolidation — COMPLETED**: Final 22:42 UTC receipt supersedes partial observations, including Durable Supervisor host move. Source licenses, bundles, history and project records preserved. [evidence 1](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json), [evidence 2](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.md) +- **remote-incus-execution — COMPLETED**: Recorded normal Tasks Codex and Claude SSH executions to Incus targets completed; unreachable target did not fall back locally. Integration observations, not controlled savings evidence. [evidence 1](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/REMOTE_EXECUTION_PROOF_2026-09-07.md), [evidence 2](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/remote-execution/no-local-fallback.json) +- **task-15-capability-contract — COMPLETED**: Versioned manifest discovery, immutable selection, operator grants, lifecycle hooks, receipts and scoped model gateway implemented. Historical agent-run remains retired. Implementation completion does not imply deployment or research completion. [evidence 1](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/TASK_15.md), [evidence 2](https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/CAPABILITIES.md) + +Graphify is an optional standalone CLI, not an Opsle Tasks capability. Historical capability acceptance is superseded. + +Historical source evidence and superseded controls: `program/history/pre-consolidation/`. Retirement is not research completion. ## Theory reconciliation -Concept coverage: 18 canonical concepts; 18 current concept repositories mapped exactly once; Agent Gearbox maps to opsle/gearbox. +Concept coverage: 19 canonical concepts; 18 live concepts sharing 8 current homes; Agent Gearbox maps to opsle/gearbox. Canonical map: `program/THEORY_MAP.md`. Machine registry: `program/theory-registry.json`. @@ -84,7 +98,8 @@ Blockers: The validated live authorization set remains unconsumed; separate exec ## Mechanical source of truth -- Portfolio and priority authority: `program/registry.json` +- Research portfolio and priority evidence: `program/registry.json` +- Workload and task-management authority: Opsle Tasks (`opsle/tasks`) - Generated priority view: `program/PRIORITY.md` - Experiments: `program/experiments.json` - Theory registry: `program/theory-registry.json` @@ -92,3 +107,14 @@ Blockers: The validated live authorization set remains unconsumed; separate exec - Lifecycle: `program/LIFECYCLE.md` - Operating rules: `program/OPERATING_RULES.md` - Integrity check: `python3 tools/validate_program.py` + + +### All recorded experiments + +| Experiment | Status | Verdict | +|---|---|---| +| `EXP-001` | PLANNED | PENDING | +| `AV-EXP-001` | RECORDED | PASS for completing the preregistered shadow calibration: no AV selection miss was observed in the frozen corpus. This does not establish general safety, correctness equivalence, causal savings, or bounded trust. | +| `AV-EXP-002` | RECORDED | FAIL for the safety hypothesis: AV_CORE and AV_WITH_SELECTOR_EVIDENCE each selected 86/87 oracle-relevant checks and omitted the same runtime/subprocess import test in AV2-006. The controlled shadow calibration itself completed and FULL exposed the miss. | +| `AV-EXP-003` | RECORDED | PASS only for the preregistered defect-repair claim: the known AV2-006 check is selected, ten generalized cases have zero misses, no frozen prior AV-relevant check becomes newly missed, and exact precision cost is measured. This does not solve dynamic dependencies, prove general safety, or authorize trusted execution. | +| `LEGACY-001` | RECORDED | An integration path was observed; completeness and comparative benefit were not established. | diff --git a/README.md b/README.md index 10dcd28..2a0b896 100644 --- a/README.md +++ b/README.md @@ -30,8 +30,8 @@ Important mechanisms should remain understandable, falsifiable, benchmarkable, r ## Start here -- [PROGRAM_STATUS.md](PROGRAM_STATUS.md) — generated 21-repository dashboard. -- [program/registry.json](program/registry.json) — authoritative machine-readable portfolio and priority ledger. +- [PROGRAM_STATUS.md](PROGRAM_STATUS.md) — generated current and historical repository dashboard. +- [program/registry.json](program/registry.json) — authoritative machine-readable research ledger (Tasks owns workload management). - [program/PRIORITY.md](program/PRIORITY.md) — generated NOW / NEXT / THEN / LATER / PARKED view. - [program/THEORY_MAP.md](program/THEORY_MAP.md) — canonical conceptual topology and Gearbox boundary. - [program/theory-registry.json](program/theory-registry.json) — machine-readable concept classifications and dispositions. @@ -53,7 +53,7 @@ dashboard and priority view with `python3 tools/render_program_status.py`. ## Product relationship -Opsle Tasks is the current integrated reference implementation in +Opsle Tasks is the authoritative current workload and task-management system and integrated reference implementation in `opsle/tasks`. Its control plane and project execution targets are independently deployed and do not depend on this research repository at runtime. diff --git a/program/LIFECYCLE.md b/program/LIFECYCLE.md index d822204..59377a3 100644 --- a/program/LIFECYCLE.md +++ b/program/LIFECYCLE.md @@ -71,3 +71,29 @@ or tooling plus automated checks. Later gates require correctness evidence, reproducible operations, and accurate public documentation appropriate to that repository. Infrastructure is not `COMPLETE` merely because it renders or because one control document exists. + +## Repository disposition and concept homes + +`repository_disposition` is independent of `lifecycle_stage`: ACTIVE repositories +belong to current membership; CONSOLIDATED and RETIRED sources belong to historical +membership and cannot own active work. Retirement records SUPERSEDED completion +status and preserves the highest evidenced stage; it does not mean COMPLETE. +`consolidated_into` must name a current repository. Tasks project visibility and +`workload_eligible` do not imply research maturity; `.github` remains current but +is ineligible for workload execution. + +Concept identity remains stable in the theory registry. `source_repository` +preserves origin; `current_repository` can be shared by several consolidated +concepts or null for a retired concept. `highest_evidenced_stage` preserves source +research evidence without inheriting the destination's maturity. Tasks is program +infrastructure, and Visible Value is an independently recorded concept. + +Integration milestones (including remote execution and Task 15) use completed-work +records, not research COMPLETE promotions. Current and historical repository counts, +concept dispositions, homes and workload authority validate independently. + +Frozen experiment participants and roles remain historical source identities. +`historical_experiment_ids` preserves reciprocal participation for retired sources +without activation. A future authorized experiment must resolve support through +the concept's current home; historical potential support is not a dependency to +revive. Experiment next-task proposals do not bypass current Tasks admission. diff --git a/program/OPERATING_RULES.md b/program/OPERATING_RULES.md index 0716c17..b72dcf5 100644 --- a/program/OPERATING_RULES.md +++ b/program/OPERATING_RULES.md @@ -1,21 +1,18 @@ # Program operating rules -`program/registry.json` is authoritative for both portfolio state and the -NOW / NEXT / THEN / LATER / PARKED priority state. Generated Markdown is never -an independent planning authority. +`program/registry.json` is authoritative for research portfolio state and evidence-driven priority lanes. Opsle Tasks (`opsle/tasks`) is authoritative for current workload, task management and execution. Lanes describe research priorities; they do not create a parallel queue. Generated Markdown is never an independent planning authority. Every Opsle execution must: -1. Read `program/registry.json` before selecting or performing work. +1. Use the authorized Opsle Tasks record to select work; consult `program/registry.json` for research evidence and scope. 2. Verify the relevant repository default branch and HEAD before relying on recorded state. -3. Start from the current program lane and operating question before selecting - an explicit repository or experiment objective. +3. Reconcile the authorized task with the operating question and current repository membership; historical sources are not activation candidates. 4. Preserve immutable or content-addressed evidence for every material claim. 5. Promote lifecycle state only after satisfying the canonical gate in `program/LIFECYCLE.md`. 6. Update the registry when verified state changes. 7. Update the experiment registry when an experiment is planned, run, failed, replicated, or judged. 8. Record blockers and unknowns instead of silently bypassing them. -9. Identify one exact next meaningful task for every touched repository. +9. Record a justified next action for current repositories; historical sources have no active next task. 10. Return a bounded outcome summary rather than raw execution transcripts. 11. When any Opsle mechanism runs, preserve its machine-readable Visible Value receipt and keep its named operator indicator outside canonical model @@ -35,10 +32,7 @@ Every Opsle execution must: 15. Park cosmetic cleanup, architectural taste, hypothetical robustness, and speculative future requirements unless qualifying evidence appears. -Run `python3 tools/validate_program.py` and -`python3 tools/render_program_status.py --check` before committing a registry -change. The renderer checks both `PROGRAM_STATUS.md` and -`program/PRIORITY.md`. +Registry changes require the CI integrity checks: `python3 tools/validate_program.py`, `python3 tools/render_program_status.py --check`, the research unittest suite, and existing offline-freeze and receipt checks with pinned dependencies. The renderer updates `PROGRAM_STATUS.md`, `program/PRIORITY.md`, and `program/THEORY_MAP.md` together. Task-specific execution instructions may leave final verification to the pipeline. ## Portfolio discipline @@ -59,3 +53,9 @@ meaningful transition or completion point. A full machine value receipt may use a caller-requested deterministic sidecar when embedding it would inflate compact model context. Display timestamps, ambient repository state, and other nondeterministic fields must not contaminate deterministic semantic output. + +## Retirement and current ownership + +Durable Supervisor, Taslos Tasks, Paperclip and historical agent-run are retired. Their archived controls, migration rollback notes and experiments authorize no current work. Completed consolidation receipts supersede dated partial observations. Graphify is an optional standalone CLI, not a Tasks capability. Do not recreate retired repositories or execution systems from stale references. + +Repository activity, Tasks project visibility and research maturity are independent. `.github` is retained for housekeeping but excluded from workload eligibility. Consolidated concepts share active homes; historical source records retain evidence without becoming active dependencies. Release and data-safety requirements remain separate from task authorization. diff --git a/program/PRIORITY.md b/program/PRIORITY.md index 6a938fc..b0be822 100644 --- a/program/PRIORITY.md +++ b/program/PRIORITY.md @@ -2,11 +2,11 @@ -`program/registry.json` is the sole priority authority. Edit the registry, not this file. +`program/registry.json` records research priorities. Opsle Tasks is the current workload and task-management authority. Edit the registry, not this generated view. ## Operating question -> What prevents Durable Supervisor from successfully finishing Opsle Tasks with less intelligence, less context, less human involvement, and defensible evidence? +> How can current Opsle Tasks workloads use less intelligence and context while preserving correctness and defensible evidence? ## Anti-nitpick guardrail @@ -33,64 +33,47 @@ Park by default: | Lane | Objective | Repositories | Entry gate | |---|---|---|---| -| **NOW** | Finish Durable Supervisor v0.1 as a bounded measured system, then freeze feature work. | `durable-supervisor`, `research` | Current program lane. | -| **NEXT** | Use Opsle Tasks as the primary real-world workload and collect integrated measurements from the current implementation. | `gearbox`, `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `affected-verification` | Durable Supervisor v0.1 is declared and feature-frozen. | -| **THEN** | Advance an individual concept only when Durable Supervisor or Opsle Tasks evidence demonstrates a concrete need. | `semantic-edit-protocol`, `event-driven-agent-wakeup`, `agent-state-ledger`, `agent-scheduler-runtime`, `verifiable-agent-handoff`, `agent-routing-policy`, `agent-resource-claims`, `agent-discovery-control`, `agent-execution-authorization`, `controlled-agent-acceptance`, `agent-recovery-policy`, `ephemeral-agent-workers` | A qualifying work-item reason and real workload evidence identify the smallest relevant concept boundary. | -| **LATER** | Run controlled experiments, freeze a real-workload benchmark corpus, seek independent replication, and only then consider public product and research-site releases. | `site` | NOW, NEXT, and evidence-triggered THEN work establish a defensible need and release prerequisites are separately authorized. | -| **PARKED** | Retain useful non-priority ideas without turning them into active work. | `.github` | The idea is useful but lacks a qualifying reason to compete with the current objective. | +| **NOW** | Maintain the current Tasks workload and evidence-backed program state. | `tasks`, `research` | An authorized Tasks work item and a qualifying evidence-backed reason are required. | +| **NEXT** | Address demonstrated integration or measurement deficiencies from Tasks workloads. | `gearbox`, `context-firewall`, `visible-value`, `affected-verification`, `decision-evidence-protocol`, `agent-trajectory-profiler` | An authorized Tasks work item and a qualifying evidence-backed reason are required. | +| **THEN** | Investigate semantic edits only when workload evidence establishes a need. | `semantic-edit-protocol` | An authorized Tasks work item and a qualifying evidence-backed reason are required. | +| **LATER** | Prepare controlled research and public material when evidence and separate release authority support it. | `site` | An authorized Tasks work item and a qualifying evidence-backed reason are required. | +| **PARKED** | Retain organization housekeeping separately from workload eligibility. | `.github` | An authorized Tasks work item and a qualifying evidence-backed reason are required. | -### NOW — Finish Durable Supervisor v0.1 as a bounded measured system, then freeze feature work. +### NOW — Maintain the current Tasks workload and evidence-backed program state. -Entry: Current program lane. +Entry: An authorized Tasks work item and a qualifying evidence-backed reason are required. -Exit: Every Durable Supervisor v0.1 stopping criterion is satisfied and the release is explicitly declared and frozen. +Exit: The scoped defect or question is resolved with bounded evidence; release requirements remain separate. -### NEXT — Use Opsle Tasks as the primary real-world workload and collect integrated measurements from the current implementation. +### NEXT — Address demonstrated integration or measurement deficiencies from Tasks workloads. -Entry: Durable Supervisor v0.1 is declared and feature-frozen. +Entry: An authorized Tasks work item and a qualifying evidence-backed reason are required. -Exit: Durable Supervisor has driven the remaining authorized Opsle Tasks readiness work and produced defensible integrated measurements. +Exit: The scoped defect or question is resolved with bounded evidence; release requirements remain separate. -### THEN — Advance an individual concept only when Durable Supervisor or Opsle Tasks evidence demonstrates a concrete need. +### THEN — Investigate semantic edits only when workload evidence establishes a need. -Entry: A qualifying work-item reason and real workload evidence identify the smallest relevant concept boundary. +Entry: An authorized Tasks work item and a qualifying evidence-backed reason are required. -Exit: The demonstrated deficiency is resolved or falsified at its narrowest justified boundary; do not march mechanically through the portfolio. +Exit: The scoped defect or question is resolved with bounded evidence; release requirements remain separate. -### LATER — Run controlled experiments, freeze a real-workload benchmark corpus, seek independent replication, and only then consider public product and research-site releases. +### LATER — Prepare controlled research and public material when evidence and separate release authority support it. -Entry: NOW, NEXT, and evidence-triggered THEN work establish a defensible need and release prerequisites are separately authorized. +Entry: An authorized Tasks work item and a qualifying evidence-backed reason are required. -Exit: Applicable evidence and separate release authorization exist. +Exit: The scoped defect or question is resolved with bounded evidence; release requirements remain separate. -### PARKED — Retain useful non-priority ideas without turning them into active work. +### PARKED — Retain organization housekeeping separately from workload eligibility. -Entry: The idea is useful but lacks a qualifying reason to compete with the current objective. +Entry: An authorized Tasks work item and a qualifying evidence-backed reason are required. -Exit: New evidence supplies a qualifying work-item reason. - -## Durable Supervisor v0.1 stopping criteria - -Status: **IN_PROGRESS**. Foundation: `VERIFIED_ENOUGH_TO_STOP_NITPICKING_UNLESS_REAL_WORKLOAD_EXPOSES_A_DEFECT`. - -Verified main: `1b5ab7631ba651a32592bbbdab8001865a3baf3d`. Runtime: `PAUSED_NO_ACTIVE_TASK_OR_ATTEMPT` as durably recorded at `2026-09-05T12:20:01.551Z`. - -1. **OPEN** — Expose the selected Gearbox route, routing rationale, and Context Firewall accounting in an operator-visible per-child receipt. -2. **OPEN** — Expose each child's model and reasoning effort. -3. **OPEN** — Record actual tokens and cost when provider evidence supplies them, and clearly labeled estimates otherwise. -4. **OPEN** — Measure raw evidence or context versus Context Firewall retained context and report the reduction. -5. **OPEN** — Report first-pass success rate. -6. **OPEN** — Report repair-child or retry rate and the token cost of retries. -7. **OPEN** — Estimate avoidable intelligence consumption by identifying failures whose needed facts were discoverable through deterministic preflight before model execution. -8. **OPEN** — Implement only the smallest useful deterministic preflight or reconnaissance mechanism shown necessary by the avoidable-failure evidence. -9. **OPEN** — Complete another real foreign-repository workload that exercises and preserves the new measurements. -10. **OPEN** — Declare Durable Supervisor v0.1 and freeze feature work except for defects exposed by real workloads. +Exit: The scoped defect or question is resolved with bounded evidence; release requirements remain separate. ## Opsle Tasks boundary -Opsle Tasks is the NEXT primary real-world workload after Durable Supervisor v0.1. Its current repository remains `opsle/tasks`. +Opsle Tasks is the current workload and task-management authority. Its repository is `opsle/tasks`. -Measure: Gearbox, Context Firewall, Decision Evidence Protocol, Agent Trajectory Profiler, Affected Verification. +Measure: Gearbox, Context Firewall, Visible Value, Affected Verification. Without separate authorization, do not: @@ -100,16 +83,16 @@ Without separate authorization, do not: ## Evidence-triggered concept activation -- routing → `agent-routing-policy` -- durable state → `agent-state-ledger` +- routing → `gearbox` +- durable state → `tasks` - retry or deterministic preflight → `gearbox` - context overload → `context-firewall` - verification selection → `affected-verification` -- recovery → `agent-recovery-policy` +- recovery → `tasks` ## Visible Value target -Baseline: The configured primary supervisor model and reasoning effort performs all child work itself and receives raw unfiltered evidence. +Baseline: The configured primary model and reasoning effort performs the same task work and receives raw unfiltered evidence. Savings require an inspectable same-work baseline or a clearly labeled estimate derived from that baseline; marketing counterfactuals are prohibited. @@ -120,7 +103,7 @@ Savings require an inspectable same-work baseline or a clearly labeled estimate Per-child receipt: child/task identity; model; reasoning effort; Gearbox route; routing rationale; input tokens; output tokens; raw evidence/context size; Context Firewall retained size; reduction percentage; estimated tokens avoided; estimated cost avoided; duration; attempt number; success/failure; escalation/retry reason; retry potentially avoidable through deterministic preflight. -Supervisor/run summary: total children; model/effort distribution; total model tokens consumed; work completed deterministically without model use; Context Firewall reduction; estimated token/cost savings; first-pass success rate; repair-child rate; tokens spent on retries; avoidable-intelligence estimate. +Task/run summary: total children; model/effort distribution; total model tokens consumed; work completed deterministically without model use; Context Firewall reduction; estimated token/cost savings; first-pass success rate; repair-child rate; tokens spent on retries; avoidable-intelligence estimate. ## Later @@ -133,11 +116,8 @@ Supervisor/run summary: total children; model/effort distribution; total model t ## Parked -- Durable Supervisor src/cli.js is 1,570 lines at the verified main SHA and may eventually warrant decomposition; this is a maintenance smell, not a current objective, unless measured work shows material reliability or efficiency harm. -- Background projection repair or reconciliation is not justified merely because explicit retry exists. -- Historical pre-fix cleanup or migration requires evidence of need. -- General architectural polishing remains parked unless a real workload exposes a concrete defect. +- Cosmetic cleanup, architectural polishing, hypothetical robustness and speculative requirements without qualifying evidence. ## Exact next execution -In opsle/durable-supervisor, make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening Durable Supervisor v0.1. +Select the next authorized task from Opsle Tasks; use its project scope and workload evidence to identify the smallest justified change. This ledger does not create a parallel queue. diff --git a/program/THEORY_MAP.md b/program/THEORY_MAP.md index eb393ab..e5d88d8 100644 --- a/program/THEORY_MAP.md +++ b/program/THEORY_MAP.md @@ -1,511 +1,54 @@ -# Canonical Opsle theory map +# Current theory map -Status: authoritative conceptual reconciliation plus the 2026-08-29 public -Gearbox extraction and 2026-08-31 Affected Verification project creation. -Consolidation and disposition operations remain recommendations only. + -Machine source: [`theory-registry.json`](theory-registry.json). +Opsle Tasks owns current workload and task management. The program ledger records research state; it does not create a parallel task queue. -Theory registry canonical SHA-256: -`2538ad9b59ebb0cbd34735df8ae3e49ae2f24a1cc98d8f55ac7ee426e6c69e6d`. - -This map corrects an extraction-boundary error. The original 16 concept -repositories were useful hypotheses isolated from one production system, but -hypothesis granularity was treated as repository and product granularity. The -resulting cross-repository diagram then placed almost every concept in one autonomous -supervisor pipeline. Agent Gearbox initially had no canonical home, and shared words such -as *bounded child*, *delegation*, and *waiting without inference* allowed its -primary-developer transmission theory to be conflated with Durable Supervisor. - -The separately authorized Gearbox extraction created one public repository and -assigned its new, evidence-backed implementation `PROTOTYPED`. No existing -repository was promoted, demoted, consolidated, renamed, archived, or deleted. -Existing evidence remains attached to the exact implementation or profile it -actually exercises. - -## Canonical topology - -```text -OPSLE -├── Agent Gearbox public narrow prototype -│ ├── deterministic-versus-cognitive admission -│ ├── gear and model/effort selection -│ ├── content-addressed bounded context -│ ├── deterministic executor / bounded helper executor -│ ├── OS/transport passive wait -│ ├── compact result, literal budget, cleanup, and value telemetry -│ └── consumes external policies/protocols where required -├── Independent Opsle tools -│ ├── Context Firewall -│ ├── Agent Trajectory Profiler -│ ├── Affected Verification -│ └── Semantic Edit Protocol -├── Cross-cutting protocols -│ ├── Decision Evidence Protocol -│ └── Verifiable Agent Handoff -├── Gearbox-facing supporting policy research -│ ├── Agent Routing Policy -│ ├── Agent Execution Authorization -│ └── Agent Resource Claims -├── Durable orchestration research family -│ ├── Durable Supervisor family hypothesis / objective owner -│ ├── Agent State Ledger history and projection module -│ ├── Agent Scheduler Runtime readiness and time module -│ ├── Event-Driven Agent Wakeup durable activation module -│ ├── Agent Discovery Control discovered-work admission module -│ └── Agent Recovery Policy bounded recovery permission module -├── Execution and isolation infrastructure -│ └── Ephemeral Agent Workers -└── Research-only system acceptance - └── Controlled Agent Acceptance -``` - -The tree is conceptual, not a repository-creation plan. In particular, the -Durable Supervisor family has multiple recorded repositories today but likely -fewer future implementation boundaries. Conversely, Context Firewall, Decision -Evidence, Profiler, Semantic Edit, Handoff, and isolation have meaningful reuse -outside either Gearbox or durable orchestration. - -## Agent Gearbox - -> Agent Gearbox lets a powerful primary developer delegate routine operations -> and bounded work to deterministic software or less expensive models, then -> receive only the compact result needed to continue. - -Guiding idea: **Stop using intelligence for work that does not require -intelligence.** - -The primary developer retains end-to-end project understanding, architecture, -safety decisions, integration, ambiguity resolution, and the final definition -of completion. Gearbox is subordinate to that developer and one explicit -bounded call resembling: - -```text -gearbox_run( - task, - task_type, - requested_gear, - allowed_context, - output_contract, - authority, - budget -) -``` - -### Irreducible mechanism - -1. Validate the task, authority, context manifest, requested gear, model and - effort limits, output contract, and literal budget. -2. Admit deterministic software whenever cognition is unnecessary. -3. Admit one fresh bounded helper only when cognition is necessary. -4. Resolve only explicitly allowed, content-addressed context. -5. Execute through a deterministic or bounded cognitive gear. -6. Block at the OS or transport layer until the bounded execution terminates; - the primary model does not poll. -7. Keep raw logs outside primary model context and addressable. -8. Return one compact structured success, failure, indeterminate, or escalation - result. -9. Terminate temporary helpers, reconcile residue, and fail closed if bounded - cleanup cannot be established. -10. Emit exact or observed Visible Value telemetry without inventing token, - cost, latency, or causal savings. - -### Current minimal package boundary - -```text -src/ -├── contract/ request, result, gear, budget, and error schemas -├── admission/ deterministic-vs-cognitive and requested-gear validation -├── context/ content-addressed allow-manifest resolution -├── gears/ -│ ├── deterministic/ -│ └── bounded-helper/ -├── transport/ passive blocking and terminal-result capture -├── budget/ literal provider/session/tool/time accounting -├── cleanup/ helper termination and residue reconciliation -├── result/ compact result assembly and fail-closed escalation -└── telemetry/ Visible Value and trajectory emission hooks -``` - -The public reference core currently implements these responsibilities in a -small Python package rather than one module per diagram node. The diagram is a -responsibility map, not a requirement to split packages or repositories. - -### External dependencies and profiles - -- **Context Firewall** adapts raw tool or helper evidence for primary context. - Gearbox does not own reducer policies. -- **Decision Evidence Protocol** defines or validates compact result facts. - Gearbox owns execution, not general evidence-protocol governance. -- **Agent Trajectory Profiler** consumes execution and Visible Value telemetry. - It never selects a gear. -- **Agent Routing Policy** may supply one Gearbox-specific model, effort, and - provider route after Gearbox admits cognitive execution. -- **Agent Execution Authorization** may validate a capability-style authority - envelope. Reduction and redaction never become authorization. -- **Agent Resource Claims** is optional when shared mutable resources require - leases and fences; it is not needed for every bounded call. -- **Ephemeral Agent Workers** is optional execution/isolation infrastructure for - risky helpers; deterministic local gears need not use it. -- **Verifiable Agent Handoff** is optional when a helper source environment will - be destroyed before independent verification. - -### Explicit non-responsibilities - -Gearbox must not absorb: - -- durable objective ownership, cross-activation reconstruction, queues, - schedules, Global Pause, or autonomous discovery; -- autonomous retry, consultation, alternate-provider fallback, or recovery; -- a persistent hierarchy of agents or a second permanent supervisor; -- exact-session resume or replacement of ChatGPT Remote; -- generic worker isolation, handoff storage, or every Opsle protocol; -- final integration or completion authority from the primary developer. - -### Visible Value contract +Gearbox determines where bounded work executes; Context Firewall determines what evidence returns. Visible Value validates receipts and presents operator measurements. Affected Verification plans checks under explicit evidence boundaries. -A Gearbox run should expose, when directly observed: +Task 15 completed the versioned capability contract: trusted manifests, operator grants, immutable repository selection, lifecycle hooks, receipts and an execution-scoped model gateway. Graphify is an optional standalone CLI, not a Tasks capability. -- requested, admitted, and executed gear; -- deterministic operation identity or helper model/effort identity; -- content manifest identity and permitted bytes/artifacts; -- provider sessions actually started, with missing counts left missing; -- passive-wait duration only when directly measured outside deterministic - semantic output; -- helper termination and residue state; -- raw and returned evidence bytes from compatible Context Firewall receipts; -- exact budget use and typed rejection/escalation reasons. +The final consolidation receipt supersedes earlier partial observations. Durable Supervisor is retired, with no current home or activation gate. Other consolidated concepts share Tasks or Gearbox; their source maturity does not increase through migration. `.github` remains a repository but is ineligible for workload execution. -It may report a deterministic task completed with zero provider sessions when -that fact is directly recorded. It may not call that value *tokens saved* or -*cost avoided* without a controlled baseline and provider-recorded usage. +Historical definitions, extraction analysis, initial SHAs, licenses, negative evidence and superseded recommendations are preserved in [the historical map](history/pre-consolidation/THEORY_MAP.md) and [historical registry](history/pre-consolidation/theory-registry.json). They authorize no work. -## Gearbox versus Durable Supervisor +Coverage: 19 concepts; active 7; consolidated 11; retired 1; 8 distinct current homes. -| Dimension | Agent Gearbox | Durable Supervisor / autonomous orchestration | -|---|---|---| -| Primary owner | One continuously responsible powerful primary developer | A durable objective-owning supervisor or orchestrator | -| Persistence | Bounded call state; durable receipts/artifacts as needed | Durable objective, work, event, attempt, and reconstruction state | -| Session lifecycle | One primary-developer activation delegates and receives one terminal result | May stop and resume reasoning across multiple activations | -| Child purpose | Perform one routine operation or bounded cognitive assignment cheaply | Advance autonomous objective state between supervisor decisions | -| Waiting | Synchronous OS/transport blocking without primary-model polling | Persisted wait registration and event-driven reactivation | -| State | Content manifest, request, result, budget, and cleanup state for one call | Ledger projection spanning objectives, tasks, attempts, schedules, and events | -| Retry/recovery | None by default; return failure or indeterminate and fail closed | May use a separately authorized bounded recovery policy across attempts | -| Scheduling | None beyond executing the current bounded call | Queues, dependencies, timeouts, schedules, pause, and readiness transitions | -| Completion | Returns a compact result; the primary developer decides project completion | Records durable task/objective transitions under evidence gates | -| Context ownership | Primary developer owns project context; Gearbox sees only an allowlisted slice | Supervisor reconstructs the model-readable state required for the next activation | -| Intended use | Save the primary developer's intelligence and context on routine bounded work | Operate long-horizon autonomous work despite inactive or interrupted reasoning sessions | -| Failure model | One typed terminal result, escalation, literal budget, helper termination, no hidden fallback | Crash/restart, duplicate/lost events, stale projections, retries, recovery, and terminal orchestration states | - -Canonical invariant: - -> Gearbox enhances a primary developer. Durable orchestration owns progress -> across activations without requiring that developer to remain continuously -> active. - -The earlier wording “can operate across multiple activations” is necessary but -not sufficient. The decisive distinction is **objective ownership**: Gearbox -never takes the primary developer's end-to-end ownership, whereas a Durable -Supervisor owns durable autonomous progress between decisions. - -The prior drift occurred because Durable Supervisor and Event Wakeup also use a -fresh child, bounded assignment, and no model inference while waiting. Their -durable objective, reconstruction, scheduler, and reactivation semantics are the -parts that make them a different theory. - -## Context Firewall - -> Context Firewall is a deterministic boundary that keeps operational noise out -> of an AI agent's context while preserving the compact evidence, provenance, -> and escalation path the agent needs to make correct decisions. - -```text -tool or process - ↓ -raw addressable evidence - ↓ -deterministic versioned adapter and policy - ↓ -compact evidence packet - ↓ -model context -``` - -The packet carries the functional equivalent of schema/policy version, -producer/tool class, source identity and SHA-256, exit state, compact facts, -retained/suppressed accounting, completeness, raw locator, escalation state and -reason, and a deterministic packet identity. - -Context Firewall is not an AI summarizer, sandbox, authorization system, -redaction system, evidence deletion mechanism, supervisor, or proof that model -correctness was preserved. It fails closed or escalates for unsupported formats, -incomplete parsing, uncertain identity, unexpected truncation, unsafe evidence -ceilings, unclassifiable evidence, missing or hash-invalid raw artifacts, and -source/policy incompatibility. - -The current repository implements one strict flat TAP-compatible adapter. That -adapter supports `PROTOTYPED` for its scoped implementation; it does not imply -the adapter families below exist. - -### Intended adapter families - -| Family | Likely compact facts | Unsafe to hide | Escalation conditions | Raw artifact handling | -|---|---|---|---|---| -| Tests | runner/profile, exit/interruption, pass/fail/skip totals, failed identities and complete failure regions, duration when supplied | every failure, crash, timeout, contradiction, flaky/retry marker that changes verdict | unsupported dialect, nested structure not understood, truncated failure, contradictory counts, unexplained exit, ceiling cannot retain failures | retain exact stdout/stderr bytes, framed hash, runner identity, and immutable locator | -| Lint | tool/config/version, file/rule/severity counts, affected paths, fixability, exit state | errors, parser/config failures, suppressed-rule changes, unsafe autofix warnings | unknown formatter, omitted diagnostics, path ambiguity, config mismatch, truncated output | retain complete diagnostics and config identity; packet links line/rule facts to raw bytes | -| Typecheck | compiler/project version, diagnostic codes, files/locations, error counts, build mode | all type errors, project-reference failures, emit-on-error behavior, config resolution defects | unrecognized diagnostic grammar, incremental cache uncertainty, truncated chains, project/config drift | retain compiler streams, project graph/config hashes, and artifact locator | -| Git | repository/worktree identity, HEAD/base/target, status counts, changed paths, diff/stat/ref outcomes | conflicts, unmerged entries, rejected refs, dirty-overlap evidence, signature or object failures, destructive target ambiguity | ambiguous repository, stale refs, truncated diff, unsafe path encoding, missing objects, hash mismatch | retain exact command output and required diff/bundle/object evidence with repository identity | -| Build/compiler | toolchain/config/target, exit state, produced artifact identities, error/warning counts, cache state | compile/link errors, missing outputs, nondeterminism warnings, unsafe fallback, signing or packaging defects | unsupported formatter, missing artifact, cache provenance uncertain, truncated diagnostic, artifact hash mismatch | retain logs plus content-addressed outputs, manifests, and toolchain/config identity | -| Processes/services | command/unit identity, exit/signal, health state, revision, bounded recent error facts | crash loops, unhealthy identity/revision, permission failures, timeouts, resource exhaustion, cleanup residue | uncertain process identity, log gap, unexpected truncation, health/workload disagreement, missing start/end evidence | retain bounded journal/process artifacts with cursor/time window, command/revision identity, and hashes | -| Helper-agent results | task/context/authority/budget/model identity, terminal status, changed artifacts, verification, limitations, cleanup | any failure, uncertainty, unauthorized access, budget breach, incomplete verification, residue, hidden retry/fallback | malformed result contract, context mismatch, missing raw transcript/artifact, provider identity uncertainty, truncated result, helper not terminated | retain raw transcript/logs outside primary context, content-addressed result artifacts, request/response hashes, and cleanup receipt | - -## Seventeen repository classifications - -The classification is conceptual. The disposition is a future recommendation, -not an executed repository action. +Theory registry canonical SHA-256: +`f079b141834b738d9df6dbd2f39c2ae063e4c4186e39477130beb3dc9db7cf7b`. -| Repository | Primary classification | Recommended disposition | Confidence | One-sentence rationale | +| Concept identity | Classification | Executed / retained disposition | Confidence | Current home and maturity | |---|---|---|---|---| -| `gearbox` | `GEARBOX_CORE` | `KEEP_STANDALONE` | `HIGH` | One bounded primary-developer transmission operation now has a narrow public home without absorbing durable orchestration or external policies and protocols. | -| `affected-verification` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | Minimum-defensible verification planning composes native impact evidence, catalogs, and risk policy without owning check execution or model-context reduction. | -| `agent-trajectory-profiler` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | Reusable correctness-gated telemetry is external to both Gearbox and durable orchestration, although generic Visible Value scope needs ownership cleanup. | -| `semantic-edit-protocol` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | Bounded structural editing is independently useful and can be selected as a deterministic Gearbox tool without becoming Gearbox core. | -| `durable-supervisor` | `DURABLE_ORCHESTRATION` | `KEEP_AS_RESEARCH` | `HIGH` | Durable objective ownership and reconstruction form a distinct autonomous-orchestration hypothesis whose package boundary remains unproven. | -| `event-driven-agent-wakeup` | `DURABLE_ORCHESTRATION` | `CONSOLIDATE_WITH_OTHER` | `MEDIUM_HIGH` | Durable wait registration and event reactivation share the supervisor/scheduler state boundary and are not Gearbox's synchronous passive wait. | -| `context-firewall` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | Deterministic evidence adaptation has standalone value; the current TAP reducer is one adapter, not the whole concept. | -| `decision-evidence-protocol` | `CROSS_CUTTING_PROTOCOL` | `KEEP_AS_PROTOCOL` | `HIGH` | Vendor-neutral evidence representation and independent conformance are reused by many producers and consumers. | -| `agent-state-ledger` | `DURABLE_ORCHESTRATION` | `CONSOLIDATE_WITH_OTHER` | `MEDIUM_HIGH` | The ledger is the Durable Supervisor family's authoritative history/projection module rather than a second product boundary. | -| `agent-scheduler-runtime` | `DURABLE_ORCHESTRATION` | `CONSOLIDATE_WITH_OTHER` | `HIGH` | Readiness, time, pause, and queue transitions belong to the durable-orchestration runtime and currently overlap claims, routing, and recovery. | -| `verifiable-agent-handoff` | `CROSS_CUTTING_PROTOCOL` | `KEEP_AS_PROTOCOL` | `HIGH` | Durable result transfer across source destruction is reusable, while the current HMAC code proves only a seal subcomponent. | -| `agent-routing-policy` | `GEARBOX_SUPPORTING_POLICY` | `FUTURE_GEARBOX_POLICY` | `MEDIUM_HIGH` | Gearbox needs a scoped model/effort/provider route policy after gear admission, but reviewer and fallback routing remain external concerns. | -| `agent-resource-claims` | `GEARBOX_SUPPORTING_POLICY` | `KEEP_AS_RESEARCH` | `MEDIUM` | Leases and fencing can support Gearbox and other systems, but a portable policy boundary and independent packaging value are not yet proven. | -| `agent-discovery-control` | `DURABLE_ORCHESTRATION` | `CONSOLIDATE_WITH_OTHER` | `HIGH` | Autonomous discovered-work admission depends on the supervisor's durable objective graph, ledger, and budgets and is outside explicit-task Gearbox. | -| `agent-execution-authorization` | `GEARBOX_SUPPORTING_POLICY` | `FUTURE_GEARBOX_POLICY` | `HIGH` | Exact capability-style authority is required before Gearbox execution but remains reusable by orchestrators and isolation infrastructure. | -| `controlled-agent-acceptance` | `RESEARCH_ONLY_HYPOTHESIS` | `KEEP_AS_RESEARCH` | `MEDIUM_HIGH` | One-shot autonomy acceptance is an experiment-control method, not a production runtime or Gearbox component. | -| `agent-recovery-policy` | `DURABLE_ORCHESTRATION` | `CONSOLIDATE_WITH_OTHER` | `HIGH` | Recovery permission relies on durable attempts and shared budgets; canonical Gearbox explicitly performs no autonomous retry or fallback. | -| `ephemeral-agent-workers` | `EXECUTION_ISOLATION_INFRASTRUCTURE` | `KEEP_STANDALONE` | `HIGH` | Brokered disposable execution and destruction proof are reusable infrastructure independent of reasoning and orchestration. | - -Detailed original problems, implementation fidelity, drift status, gains, risks, -provenance concerns, relationships, dependencies, consumers, evidence, and -unresolved questions are machine-readable in `theory-registry.json`. - -## Accidental duplication and exact boundaries - -### Gearbox selection versus routing - -Gearbox decides whether the explicit task should use deterministic software or -bounded cognition and validates the requested gear. Routing selects one eligible -model/provider/profile after cognitive admission. Routing does not admit -cognition and neither decision authorizes autonomous fallback. - -### Discovery versus routing - -Discovery decides whether newly proposed autonomous work should exist. Routing -decides where an already admitted cognitive attempt may run. Gearbox accepts an -explicit primary-developer task and creates no recursive work. - -### Authorization versus resource claims - -Resource Claims establishes current ownership of a conflicting resource set -through leases and fences. Execution Authorization decides whether a subject may -perform one exact action and binds current claim identity plus immutable source, -target, purpose, and route authority. Authorization references claims; it does -not reimplement claim lifecycle. - -### Scheduler versus resource claims and routing - -Scheduler owns **when** durable work is ready. Claims owns **who currently -controls** a resource. Routing owns **where/by which eligible executor** a -cognitive attempt may run. Provider availability is route evidence; retry timing -is scheduler/recovery state. - -### Recovery versus routing and acceptance - -Recovery decides whether another attempt can add information or change -conditions and which recovery class is allowed. Routing selects the exact route -only for that separately authorized attempt. Controlled Acceptance preserves -first-failure behavior and cannot silently enable recovery unless its immutable -manifest says so. - -### State Ledger versus Durable Supervisor - -The ledger records immutable facts and projects current state. Durable -Supervisor owns objective decisions based on that projection. No supervisor, -scheduler, or discovery module may maintain a competing authoritative history. - -### Event wakeup versus scheduler - -The scheduler owns wait readiness and timeout policy; the wakeup module persists -registrations and delivers idempotent decision-relevant activation events. One -shared event identity and store contract is required. - -### Handoff versus Decision Evidence and workers - -Decision Evidence owns generic fact/receipt conformance. Handoff adds exact -artifact identity, durable-before-destroy ordering, and fresh reconstruction. -Worker infrastructure enforces containment and lifecycle, then consumes Handoff -for result transfer. A worker must not mint its own authorization or duplicate -the handoff protocol. - -### Context Firewall versus Decision Evidence and Profiler - -Context Firewall reduces and emits. Decision Evidence validates the packet and -its claims independently. Profiler records exposure and value observations. The -Decision Evidence validator may reclassify source bytes for a pinned conformance -profile, but it must not evolve into a competing reduction policy. Profiler must -not become the normative owner of every receipt protocol merely because it -aggregates them. - -### Affected Verification versus selectors, Gearbox, and Context Firewall - -Native affected/related systems remain authoritative evidence providers for -their own graphs and test frameworks. Affected Verification composes that -evidence with a complete verification catalog and explicit risk policy to -produce selected checks, explained skips, and a sufficiency or uncertainty -state. It does not reimplement those selectors, execute CI, or claim global -mathematical minimality. - -Gearbox may consume a plan and choose an execution gear, but does not own the -verification theory or planner. After checks execute, Context Firewall decides -which results enter model context; Affected Verification decides only what -verification should execute. Decision Evidence may validate plan provenance, -and Agent Trajectory Profiler may measure planned versus actual or shadow work, -without initial package coupling. - -## Restored home and remaining gaps - -`opsle/gearbox` now coherently owns deterministic task admission, requested-gear -validation before model/provider routing, content-addressed staged helper -context, exact model and reasoning-effort profiles, one injected bounded helper -transport, OS-level blocking wait, compact result and failure contracts, -literal command/provider/context/output/time budgets, termination checks, raw -artifact accounting, and directly observed provider-session counts. - -The portable core deliberately leaves production helper isolation, provider -routing, general execution authorization, Context Firewall adapters, Decision -Evidence conformance, and trajectory aggregation outside its ownership. A -production helper transport and controlled provider-session/correctness/context -comparison remain implementation and evidence gaps. No concept now lacks a -machine-recorded home, and these responsibilities do not justify additional -repositories. - -## Implementation fidelity and lifecycle scope - -- Agent Gearbox is prototyped for its provider-free public core: deterministic - execution and injected-helper contract paths are runnable and tested, while - no production helper transport, Context Firewall integration, comparative - benchmark, or model/provider subject exists. -- Agent Trajectory Profiler is verified for its implemented metrics, Context - Firewall profile, run records, and Visible Value summaries; predictive or - causal benefit remains unmeasured. -- Context Firewall is prototyped for the TAP-subset adapter. No other adapter - family exists, and no safe correctness frontier is known. -- Decision Evidence is verified for the Context Firewall and Visible Value - profiles. Its generic multi-tool envelope remains narrow. -- Event-Driven Agent Wakeup is prototyped only as an in-memory transition - function; durability, restart, event delivery, and timeout behavior are not - implemented. -- Verifiable Agent Handoff is prototyped only for HMAC manifest binding and - caller-supplied destruction state; publication, transport, destruction proof, - reconstruction, and independent verification are not implemented. -- Affected Verification is verified for its narrow deterministic planner and - the revision-bound AV-EXP-001 Zustand/Vitest shadow calibration. Its frozen - corpus observed 8/8 relevant checks selected by both AV arms, including - conservative full broadening under incomplete impact evidence. The result is - still one repository, one ecosystem, and mostly synthetic change/fault - shapes; no production adapter, independent qualifying replication, or trusted - selective-verification class exists. -- The remaining theory repositories contain coherent falsifiable narratives, - but their `SPEC.md` files are generic templates and their source/tests are - placeholders. Their current lifecycle remains `THEORY`. - -Conceptual reclassification neither promotes nor demotes any repository. A -future consolidation does not combine lifecycle stages arithmetically: evidence -continues to support only its exact historical concept, revision, and scope. - -## EXP-001 reconciliation - -`EXP-001` remains the correct first controlled Context Firewall experiment: - -> How much context can an AI coding agent safely not see? - -Three layers must remain separate: - -1. **Concept definition:** deterministic evidence boundary with provenance, - completeness, raw addressability, and fail-closed escalation. -2. **Current prototype:** one strict TAP-subset adapter and its packet/value - contracts. -3. **Experimental hypothesis:** reduced context preserves task correctness under - equal fixtures, models, prompts, and correctness gates. - -The conformance corpora and Prompt 005 dogfood evidence validate implementation -behavior; they are not an EXP-001 result. The content-addressed task corpus, -correctness oracle, baseline and arms, offline harness, exact model/provider -configuration, blinded allocation, coordinator, and provider-free live -preflight are frozen or qualified at their recorded revisions. `EXP-001` stays -`PLANNED` with zero run identities and zero result artifacts. - -EXP-001 has no technical dependency on Gearbox, and no provider/model subject is -authorized by the provider-free preparation. Controlled experiments are now -LATER program work rather than the immediate execution; the authoritative -current priority is the machine state in `program/registry.json`. - -## Public Opsle and predecessor boundary - -The 2026-08-25 extraction snapshot is historical provenance, not a dependency -direction. The retired predecessor demonstrated that several mechanisms were -possible, but does not establish their general primitives, repository -boundaries, or comparative benefit. - -The long-term direction is: - -```text -public Opsle mechanisms and protocols - ↓ -Opsle Tasks and other products consume pinned versions/adapters -``` - -It is not: - -```text -the retired predecessor remains the canonical implementation - ↓ -public concepts are repeatedly rediscovered after production coupling -``` - -The separately authorized 2026-08-29 extraction adapted the canonical portable -Gearbox core into a public AGPL repository from Taslos Tasks revision -`7734caf208366a0515cf4d78efc17a86363f2238`. The public provenance file records -the exact source and introduction commits. It copied no credentials, private -evidence, host paths, provider configuration, services, databases, product -state, or Durable Supervisor machinery, and the source repository remained -unchanged. Future product adoption still requires public contracts, independent -evidence, explicit versioned adapters, and separate authorization. - -## Future consolidation and provenance policy - -A later consolidation may occur only after a reviewed plan identifies the exact -source repository, target module, preserved commit/history path, lifecycle and -benchmark evidence mapping, and public-link redirect strategy. - -Minimum requirements: - -1. Never delete or silently replace the source repository. -2. Preserve commit attribution and cite the original extraction revision. -3. Preserve theory, specification, benchmark, experiment, negative-result, and - lifecycle history under stable public references. -4. Map each old artifact to a named target module; do not claim the target's - broader lifecycle from narrow source evidence. -5. Publish an archive/redirect note only after the target is available and the - mapping is independently checked. -6. Keep citations and public links working or provide explicit redirects. -7. Record rejected consolidation proposals as research history. -8. Execute repository transfer, archive, rename, or deletion only under new - explicit authority. - -## Repository anti-forgetting invariant - -The authoritative program registry tracks exactly 21 repositories: 18 concept -repositories, `research`, `site`, and `.github`. Agent Gearbox and Affected -Verification are each mapped exactly once to their public repositories. Concept -tracking does not replace repository tracking. +| `gearbox` | `GEARBOX_CORE` | `KEEP_STANDALONE` | `HIGH` | gearbox; PROTOTYPED | +| `agent-trajectory-profiler` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | agent-trajectory-profiler; VERIFIED | +| `semantic-edit-protocol` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | semantic-edit-protocol; THEORY | +| `durable-supervisor` | `DURABLE_ORCHESTRATION` | `RETIRED` | `HIGH` | none (retired); VERIFIED | +| `event-driven-agent-wakeup` | `DURABLE_ORCHESTRATION` | `CONSOLIDATED` | `MEDIUM_HIGH` | tasks; PROTOTYPED | +| `context-firewall` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | context-firewall; PROTOTYPED | +| `decision-evidence-protocol` | `CROSS_CUTTING_PROTOCOL` | `KEEP_AS_PROTOCOL` | `HIGH` | decision-evidence-protocol; VERIFIED | +| `agent-state-ledger` | `DURABLE_ORCHESTRATION` | `CONSOLIDATED` | `MEDIUM_HIGH` | tasks; THEORY | +| `agent-scheduler-runtime` | `DURABLE_ORCHESTRATION` | `CONSOLIDATED` | `HIGH` | tasks; THEORY | +| `verifiable-agent-handoff` | `CROSS_CUTTING_PROTOCOL` | `CONSOLIDATED` | `HIGH` | tasks; PROTOTYPED | +| `agent-routing-policy` | `GEARBOX_SUPPORTING_POLICY` | `CONSOLIDATED` | `MEDIUM_HIGH` | gearbox; THEORY | +| `agent-resource-claims` | `GEARBOX_SUPPORTING_POLICY` | `CONSOLIDATED` | `MEDIUM` | tasks; THEORY | +| `agent-discovery-control` | `DURABLE_ORCHESTRATION` | `CONSOLIDATED` | `HIGH` | tasks; THEORY | +| `agent-execution-authorization` | `GEARBOX_SUPPORTING_POLICY` | `CONSOLIDATED` | `HIGH` | tasks; THEORY | +| `controlled-agent-acceptance` | `RESEARCH_ONLY_HYPOTHESIS` | `CONSOLIDATED` | `MEDIUM_HIGH` | tasks; THEORY | +| `agent-recovery-policy` | `DURABLE_ORCHESTRATION` | `CONSOLIDATED` | `HIGH` | tasks; THEORY | +| `ephemeral-agent-workers` | `EXECUTION_ISOLATION_INFRASTRUCTURE` | `CONSOLIDATED` | `HIGH` | tasks; THEORY | +| `affected-verification` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | affected-verification; VERIFIED | +| `visible-value` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | visible-value; PROTOTYPED | + +## Evidence limits + +Context Firewall remains PROTOTYPED: flat TAP plus Node spec/dot reporters and a semantic-only projection. Canonical packet metrics are audit measurements; downstream delivery is measured separately. This is not an arbitrary-log parser or production security boundary. + +Gearbox remains PROTOTYPED with provider-free bounded execution and deterministic phase routing. Routing Policy is consolidated into its documented contract. Capability integration does not prove comparative savings. + +Visible Value is PROTOTYPED: an independent validator and reporting CLI. It does not authenticate external producer evidence. Missing measurements remain absent and incompatible totals are excluded. + +Affected Verification remains VERIFIED, with research limited to OBSERVE/SHADOW. AV-EXP-002 remains FAIL; AV-EXP-003 repaired its known skip without establishing general selector completeness. Tasks may use bounded manifest-backed selection only with complete impact, catalog and check-boundary evidence; unknown evidence broadens or stops. This is not a general TRUSTED_BOUNDED promotion. + +EXP-001 remains PLANNED and unconsumed. Remote Incus execution proofs are operational observations; Task 15 implementation and repository consolidation do not complete research or establish token, cost, correctness or causal savings. + +Revision-linked completed work and verified inventory: [PROGRAM_STATUS](../PROGRAM_STATUS.md), [evidence](evidence/post-consolidation/README.md). diff --git a/program/evidence/post-consolidation/README.md b/program/evidence/post-consolidation/README.md new file mode 100644 index 0000000..d2f25dd --- /dev/null +++ b/program/evidence/post-consolidation/README.md @@ -0,0 +1,43 @@ +# Post-consolidation evidence inventory + +`inventory.json` records default-branch revisions inspected through Git on +2026-09-09. The timestamp denotes the UTC observation date, not a runtime health +probe. `consolidation.json` is the exact parsed final committed Tasks migration +receipt at the revision linked by the inventory. It includes dated partial +observations for provenance; `overall: complete`, `completedAt`, and the final +Durable Supervisor host-move receipt supersede them. Archive paths are provenance +locators, not instructions to restore or reactivate sources. + +The inventory includes 11 current repositories and 12 historical sources. Eleven +sources consolidated into Tasks (ten) or Gearbox (one); Durable Supervisor retired +separately. `.github` is retained but its execution project is retired. Repository +presence does not assert that a project is visible, enabled, or currently running. +The manifest's temporary proof projects are not new research repositories. + +Current source inspection covered Tasks `docs/TASK_15.md`, `docs/CAPABILITIES.md`, +remote execution proof and no-local-fallback receipt, and the current README and +limitations of Context Firewall, Gearbox, Visible Value and Affected Verification. +The exact revision URLs and retirement commits are in `registry.completed_work`. +Later Taslos Tasks and Paperclip retirement commits supersede the remote proof's +historical statements that predecessor services still existed. + +Tasks' historical commit `00d2bac` accepted an external Graphify capability. +The approved Task 21 scope explicitly supplies the final superseding policy: +Graphify is an optional standalone CLI, not a Tasks capability. The current +bundled capability tree contains Gearbox, Context Firewall, Affected Verification +and Visible Value only. This ledger does not claim to inspect or change an external +operator installation, and absence from the bundled tree alone is not evidence +about such installations. + +Research maturity is conservative: Context Firewall, Gearbox and Visible Value +remain PROTOTYPED; Affected Verification remains VERIFIED with OBSERVE/SHADOW +research limits. Tasks integration and remote proofs are operational evidence, +not controlled comparative savings or research completion. EXP-001 remains +unconsumed; AV-EXP-002 permanently remains FAIL. + +Historical ledger and conceptual analysis are preserved under +`program/history/pre-consolidation/` at research revision +`840c79c04199762acdd59e89e807ca9086f3e45e`. Initial Gearbox publication and all pinned +experiment identities remain unchanged. No deployment, provider experiment, +external task, or PR #18 operation was performed for this reconciliation. +Final verification is delegated to the existing research pipeline. diff --git a/program/evidence/post-consolidation/consolidation.json b/program/evidence/post-consolidation/consolidation.json new file mode 100644 index 0000000..ea8e59d --- /dev/null +++ b/program/evidence/post-consolidation/consolidation.json @@ -0,0 +1,449 @@ +{ + "observedAt": "2026-09-07T20:32:15.587232+00:00", + "overall": "complete", + "sources": [ + { + "source": "agent-discovery-control", + "destination": "tasks", + "preserved": "926547dd9fd713990b1d6f1f2e650aa6c0883564", + "branch": "retire-consolidation-20260907", + "noticeHead": "b6db7271dbbdc525ba58c29478d35d27ea8628ed", + "pr": "https://github.com/opsle/agent-discovery-control/pull/1", + "worktree": "/home/deploy/worktrees/retire-agent-discovery-control-20260907", + "backup": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/before-notices/agent-discovery-control.tgz", + "noticeMerge": "09c11f8a8ffd57521110a7696ab4f6a5181c0218", + "githubStatus": "archived", + "destinationPath": "docs/contracts/agent-discovery-control", + "directoryStatus": { + "host": "moved", + "opsle-dev": "moved", + "rollbackPath": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/repos/agent-discovery-control" + }, + "projectStatus": { + "id": 5, + "retiredAt": "2026-09-07T20:28:00.910Z", + "enabled": false, + "hidden": true, + "historicalRecordsPreserved": true + } + }, + { + "source": "agent-execution-authorization", + "destination": "tasks", + "preserved": "8fb02e943c83fae121c63185d2f0d0dde8c4260a", + "branch": "retire-consolidation-20260907", + "noticeHead": "b3fa7402ac844bb78a3915c4be47f9a1a6563f17", + "pr": "https://github.com/opsle/agent-execution-authorization/pull/1", + "worktree": "/home/deploy/worktrees/retire-agent-execution-authorization-20260907", + "backup": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/before-notices/agent-execution-authorization.tgz", + "noticeMerge": "6babce3929e0284cdbb530ed5ebfb29a36388bf7", + "githubStatus": "archived", + "destinationPath": "docs/contracts/agent-execution-authorization", + "directoryStatus": { + "host": "moved", + "opsle-dev": "moved", + "rollbackPath": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/repos/agent-execution-authorization" + }, + "projectStatus": { + "id": 6, + "retiredAt": "2026-09-07T20:28:00.928Z", + "enabled": false, + "hidden": true, + "historicalRecordsPreserved": true + } + }, + { + "source": "agent-recovery-policy", + "destination": "tasks", + "preserved": "1b733a111e26e0a409fee3b96f627048531daefe", + "branch": "retire-consolidation-20260907", + "noticeHead": "03ce5c63fc47aa86ebd7c40050d1bccdefda2b03", + "pr": "https://github.com/opsle/agent-recovery-policy/pull/1", + "worktree": "/home/deploy/worktrees/retire-agent-recovery-policy-20260907", + "backup": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/before-notices/agent-recovery-policy.tgz", + "noticeMerge": "a529954c784a7763fec452c23e143e96ccb97a72", + "githubStatus": "archived", + "destinationPath": "docs/contracts/agent-recovery-policy", + "directoryStatus": { + "host": "moved", + "opsle-dev": "moved", + "rollbackPath": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/repos/agent-recovery-policy" + }, + "projectStatus": { + "id": 7, + "retiredAt": "2026-09-07T20:28:00.960Z", + "enabled": false, + "hidden": true, + "historicalRecordsPreserved": true + } + }, + { + "source": "agent-resource-claims", + "destination": "tasks", + "preserved": "dfe0fbc90c67ce5ef4256354bb62f1d511b1304c", + "branch": "retire-consolidation-20260907", + "noticeHead": "ad64f79a7b1147963ec77cf38c6423bc6eeb7755", + "pr": "https://github.com/opsle/agent-resource-claims/pull/1", + "worktree": "/home/deploy/worktrees/retire-agent-resource-claims-20260907", + "backup": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/before-notices/agent-resource-claims.tgz", + "noticeMerge": "f80f82e79ed07915feaea6a3b873d9af4aed144c", + "githubStatus": "archived", + "destinationPath": "docs/contracts/agent-resource-claims", + "directoryStatus": { + "host": "moved", + "opsle-dev": "moved", + "rollbackPath": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/repos/agent-resource-claims" + }, + "projectStatus": { + "id": 8, + "retiredAt": "2026-09-07T20:28:00.979Z", + "enabled": false, + "hidden": true, + "historicalRecordsPreserved": true + } + }, + { + "source": "agent-routing-policy", + "destination": "gearbox", + "preserved": "43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8", + "branch": "retire-consolidation-20260907", + "noticeHead": "cef2246412ee0073c8015102751eaed9f0c7d78a", + "pr": "https://github.com/opsle/agent-routing-policy/pull/1", + "worktree": "/home/deploy/worktrees/retire-agent-routing-policy-20260907", + "backup": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/before-notices/agent-routing-policy.tgz", + "noticeMerge": "179fb15add41f338813bc04cc77c308563942a93", + "githubStatus": "archived", + "destinationPath": "docs/agent-routing-policy", + "directoryStatus": { + "host": "moved", + "opsle-dev": "moved", + "rollbackPath": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/repos/agent-routing-policy" + }, + "projectStatus": { + "id": 9, + "retiredAt": "2026-09-07T20:28:00.989Z", + "enabled": false, + "hidden": true, + "historicalRecordsPreserved": true + } + }, + { + "source": "agent-scheduler-runtime", + "destination": "tasks", + "preserved": "d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0", + "branch": "retire-consolidation-20260907", + "noticeHead": "4d3d9de7e6efa09cc9a4f24068cea65456a5e871", + "pr": "https://github.com/opsle/agent-scheduler-runtime/pull/1", + "worktree": "/home/deploy/worktrees/retire-agent-scheduler-runtime-20260907", + "backup": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/before-notices/agent-scheduler-runtime.tgz", + "noticeMerge": "261adff7792b12257e57b408a84699479d2b992d", + "githubStatus": "archived", + "destinationPath": "docs/contracts/agent-scheduler-runtime", + "directoryStatus": { + "host": "moved", + "opsle-dev": "moved", + "rollbackPath": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/repos/agent-scheduler-runtime" + }, + "projectStatus": { + "id": 10, + "retiredAt": "2026-09-07T20:28:01.005Z", + "enabled": false, + "hidden": true, + "historicalRecordsPreserved": true + } + }, + { + "source": "agent-state-ledger", + "destination": "tasks", + "preserved": "acab03b1ff7168222552050e21e7553b07d00e7c", + "branch": "retire-consolidation-20260907", + "noticeHead": "20916de029ba4cd53e0b05c57fc44be096b1ff86", + "pr": "https://github.com/opsle/agent-state-ledger/pull/1", + "worktree": "/home/deploy/worktrees/retire-agent-state-ledger-20260907", + "backup": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/before-notices/agent-state-ledger.tgz", + "noticeMerge": "509a2a55066fd30ca67ceca14eaf640772b84c31", + "githubStatus": "archived", + "destinationPath": "docs/contracts/agent-state-ledger", + "directoryStatus": { + "host": "moved", + "opsle-dev": "moved", + "rollbackPath": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/repos/agent-state-ledger" + }, + "projectStatus": { + "id": 11, + "retiredAt": "2026-09-07T20:28:01.023Z", + "enabled": false, + "hidden": true, + "historicalRecordsPreserved": true + } + }, + { + "source": "controlled-agent-acceptance", + "destination": "tasks", + "preserved": "2d652adf56e53953327d09b1ba9c4a9c3445f052", + "branch": "retire-consolidation-20260907", + "noticeHead": "8ad116afe8e441c41da70ed492cdb1680e690e1d", + "pr": "https://github.com/opsle/controlled-agent-acceptance/pull/1", + "worktree": "/home/deploy/worktrees/retire-controlled-agent-acceptance-20260907", + "backup": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/before-notices/controlled-agent-acceptance.tgz", + "noticeMerge": "258325416c852be719c545a622e8a9d175447512", + "githubStatus": "archived", + "destinationPath": "docs/contracts/controlled-agent-acceptance", + "directoryStatus": { + "host": "moved", + "opsle-dev": "moved", + "rollbackPath": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/repos/controlled-agent-acceptance" + }, + "projectStatus": { + "id": 14, + "retiredAt": "2026-09-07T20:28:01.062Z", + "enabled": false, + "hidden": true, + "historicalRecordsPreserved": true + } + }, + { + "source": "ephemeral-agent-workers", + "destination": "tasks", + "preserved": "ad96fcfdfac06d340b5e96d369634980cee78ef4", + "branch": "retire-consolidation-20260907", + "noticeHead": "c9431770f05f25903674033cb2a860aa2b415420", + "pr": "https://github.com/opsle/ephemeral-agent-workers/pull/1", + "worktree": "/home/deploy/worktrees/retire-ephemeral-agent-workers-20260907", + "backup": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/before-notices/ephemeral-agent-workers.tgz", + "noticeMerge": "e758697528e71290bec55391a6b4e6c5bae83422", + "githubStatus": "archived", + "destinationPath": "docs/contracts/ephemeral-agent-workers", + "directoryStatus": { + "host": "moved", + "opsle-dev": "moved", + "rollbackPath": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/repos/ephemeral-agent-workers" + }, + "projectStatus": { + "id": 17, + "retiredAt": "2026-09-07T20:28:01.091Z", + "enabled": false, + "hidden": true, + "historicalRecordsPreserved": true + } + }, + { + "source": "event-driven-agent-wakeup", + "destination": "tasks", + "preserved": "a6209860c2151450cc28ed648bc8c2631c8db7ef", + "branch": "retire-consolidation-20260907", + "noticeHead": "79868d8dfd91eabc9fd9c097025ad61b37ec7400", + "pr": "https://github.com/opsle/event-driven-agent-wakeup/pull/1", + "worktree": "/home/deploy/worktrees/retire-event-driven-agent-wakeup-20260907", + "backup": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/before-notices/event-driven-agent-wakeup.tgz", + "noticeMerge": "8278bfb064eba50ce05aeaf9c1b7cc54a748c210", + "githubStatus": "archived", + "destinationPath": "docs/contracts/event-driven-agent-wakeup", + "directoryStatus": { + "host": "moved", + "opsle-dev": "moved", + "rollbackPath": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/repos/event-driven-agent-wakeup" + }, + "projectStatus": { + "id": 18, + "retiredAt": "2026-09-07T20:28:01.097Z", + "enabled": false, + "hidden": true, + "historicalRecordsPreserved": true + } + }, + { + "source": "verifiable-agent-handoff", + "destination": "tasks", + "preserved": "399e5cfae94345affa3f087f0f6eb9e77669d33c", + "branch": "retire-consolidation-20260907", + "noticeHead": "2cfb6d72d132326bb9cc5c59f628ee6ec08adb7b", + "pr": "https://github.com/opsle/verifiable-agent-handoff/pull/1", + "worktree": "/home/deploy/worktrees/retire-verifiable-agent-handoff-20260907", + "backup": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/before-notices/verifiable-agent-handoff.tgz", + "noticeMerge": "b832c770d9965b6060fc21af4f8826812c80b0a9", + "githubStatus": "archived", + "destinationPath": "docs/contracts/verifiable-agent-handoff", + "directoryStatus": { + "host": "moved", + "opsle-dev": "moved", + "rollbackPath": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/repos/verifiable-agent-handoff" + }, + "projectStatus": { + "id": 24, + "retiredAt": "2026-09-07T20:28:01.107Z", + "enabled": false, + "hidden": true, + "historicalRecordsPreserved": true + } + } + ], + "tasks": { + "pr": "https://github.com/opsle/tasks/pull/15", + "deployedCommit": "f40cade3ad961f66feff794f70516caf6306e2e0", + "previousCommit": "63bd1e7187c3d79698fab69058119bec4daa7802", + "testsPassed": 85, + "container": "opsle-tasks", + "unit": "deploy user opsle-tasks.service", + "databaseBackup": "/home/deploy/.local/share/opsle-tasks/backups/consolidation-20260907/before.sqlite", + "integrity": "ok", + "foreignKeyViolations": 0, + "preexistingTasksUnchanged": 12, + "excludedAndUnlistedProjectsUnchanged": true + }, + "gearbox": { + "pr": "https://github.com/opsle/gearbox/pull/3", + "commit": "12112c3a04d72b768f3252a6d0f6ea455816c9a1", + "testsPassed": 25, + "ci": "SUCCESS" + }, + "normalTaskProof": { + "taskId": 13, + "projectId": 38, + "status": "DONE", + "host": "opsle-dev.incus", + "resultCommit": "41710de356a518845c1bc43f6cc9e39f7b7fb9e3", + "selected": [ + "docs.proof" + ], + "skipped": [ + "unrelated.proof" + ], + "deployed": true, + "worktreeRemoved": true, + "providers": [ + { + "phase": "PLAN", + "provider": "codex", + "model": "gpt-6-astra", + "effort": "none", + "modelSource": "provider reported", + "effortSource": "provider reported" + }, + { + "phase": "BUILD", + "provider": "codex", + "model": "gpt-6-astra", + "effort": "none", + "modelSource": "provider reported", + "effortSource": "provider reported" + } + ], + "visibleValueStatus": "validated", + "evidenceDirectory": "/home/deploy/.local/share/opsle-tasks/backups/consolidation-20260907/proof", + "visibleValueReport": "/home/deploy/.local/share/opsle-tasks/logs/9c9f01f3b576e855-task-13-attempt-280.visible-value.json" + }, + "retained": { + "decision-evidence-protocol": "Independent Agent Trajectory Profiler CI and interoperability tools pin b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09; preserve standalone contract packaging.", + ".github": "GitHub and directories retained; project #3 retired from execution choices only." + }, + "deferred": {}, + "archive": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation", + "followup": { + "observedAt": "2026-09-07T21:47:41.431180+00:00", + "pr": "https://github.com/opsle/tasks/pull/17", + "deployedRevision": "8c63aedec81f343d0ba2f15a41006aa36fbb4ebe", + "tests": 90, + "retiredFixtures": [ + { + "projectId": 32, + "name": "SMOKE · Context Firewall release 2026-09-07", + "path": "/home/deploy/.local/share/opsle-tasks/test-projects/context-firewall-release-smoke-20260907/project", + "host": "", + "retiredAt": "2026-09-07T21:41:33.380Z", + "historicalTaskIds": [ + 5, + 6 + ], + "historyPreserved": true + }, + { + "projectId": 34, + "name": "SSH proof 1788799827123", + "path": "/home/deploy/apps/ssh-proof-20260907", + "host": "opsle-dev.incus", + "retiredAt": "2026-09-07T21:41:33.395Z", + "historicalTaskIds": [ + 9, + 8 + ], + "historyPreserved": true + }, + { + "projectId": 35, + "name": "SSH proof 1788800868896", + "path": "/home/deploy/apps/ssh-proof-final-20260907", + "host": "opsle-dev.incus", + "retiredAt": "2026-09-07T21:41:33.410Z", + "historicalTaskIds": [ + 10 + ], + "historyPreserved": true + }, + { + "projectId": 36, + "name": "SSH proof 1788801236782", + "path": "/home/family/apps/ssh-proof-authority-20260907", + "host": "dearestellie.incus", + "retiredAt": "2026-09-07T21:41:33.421Z", + "historicalTaskIds": [ + 11 + ], + "historyPreserved": true + }, + { + "projectId": 37, + "name": "SSH proof 1788803044138", + "path": "/home/family/apps/ssh-proof-release-20260907", + "host": "dearestellie.incus", + "retiredAt": "2026-09-07T21:41:33.431Z", + "historicalTaskIds": [ + 12 + ], + "historyPreserved": true + } + ], + "proof": { + "taskId": 14, + "projectId": 39, + "status": "DONE", + "projectRetired": true, + "deployedCommit": "5c4f6970527dd15ee12f90bb707f83e1aec2e06e" + }, + "ui": { + "desktop": "1280x720", + "mobile": "390x844", + "hiddenTransitions": "passed", + "togglePersistence": "passed", + "retiredExcludedWithEitherToggle": "passed", + "historicalTask12": "timeline and receipts accessible" + }, + "durableSupervisor": "deferred: verified live Claude 3857676 started 2026-09-04T15:17:24Z still uses checkout; PR #6 reviewed and preserved", + "archive": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/followup-20260907" + }, + "completedAt": "2026-09-07T22:42:38.676251+00:00", + "completedDurableSupervisor": { + "reason": "All scoped retirement actions completed; no live checkout consumers remained.", + "githubStatus": "archived", + "directoryStatus": { + "host": "moved to dated consolidation archive", + "opsle-dev": "moved to dated consolidation archive" + }, + "projectStatus": "#16 retired; no history deleted", + "exclusiveDaemon": "Native opsled stop returned stopped=true; empty registry; PID 681198 no longer running.", + "runtimeBackup": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/opsled-before.tgz", + "noticePr": "https://github.com/opsle/durable-supervisor/pull/7", + "noticeMerge": "a69f3b1ce523912a1f840cea305db008bab8e03f", + "proposalPr6": "closed as superseded; branch and patch preserved", + "finalBundle": "followup-20260907/durable-supervisor-final.bundle", + "verifiedAt": "2026-09-07T22:03:03.404765+00:00", + "hostMoveReceipt": { + "completedAt": "2026-09-07T22:42:38.676251+00:00", + "source": "/home/deploy/apps/opsle-bootstrap/repos/durable-supervisor", + "archive": "/home/deploy/apps/opsle-bootstrap/archive/20260907-consolidation/repos/durable-supervisor", + "head": "b3974f7468c19f4dd6b666409badbdcade25f0c9", + "status": "moved; clean; history and external worktree repaired and preserved", + "applicationRecordsUnchanged": true + } + } +} diff --git a/program/evidence/post-consolidation/inventory.json b/program/evidence/post-consolidation/inventory.json new file mode 100644 index 0000000..d926b3f --- /dev/null +++ b/program/evidence/post-consolidation/inventory.json @@ -0,0 +1,67 @@ +{ + "schema_version": 1, + "observed_at": "2026-09-09T00:00:00Z", + "source_revision": "e1207c5264c59e14efe9838bba3a33ba504665d2", + "source": "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "active_repositories": [ + "agent-trajectory-profiler", + "semantic-edit-protocol", + "context-firewall", + "decision-evidence-protocol", + "gearbox", + "affected-verification", + "research", + "site", + ".github", + "tasks", + "visible-value" + ], + "historical_repositories": [ + "agent-discovery-control", + "agent-execution-authorization", + "agent-recovery-policy", + "agent-resource-claims", + "agent-routing-policy", + "agent-scheduler-runtime", + "agent-state-ledger", + "controlled-agent-acceptance", + "ephemeral-agent-workers", + "event-driven-agent-wakeup", + "verifiable-agent-handoff", + "durable-supervisor" + ], + "heads": { + "agent-trajectory-profiler": "0a89661640721d6a39f127514b993d29bd728d47", + "semantic-edit-protocol": "29caad5c03827cde17aabd71c38bc25899413a33", + "context-firewall": "6dd6e5fdf21f28dc5ebfa07954aaa9bed2dbcc32", + "decision-evidence-protocol": "b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09", + "gearbox": "12112c3a04d72b768f3252a6d0f6ea455816c9a1", + "affected-verification": "792c4bb7881f6f430b2b1ba238ba05be43f50d38", + "research": "840c79c04199762acdd59e89e807ca9086f3e45e", + "site": "29c94e6440ec91f5e09aded959fe7144187af6d7", + ".github": "01c38e726db7c3e45059d25fccce55e071e35938", + "tasks": "e1207c5264c59e14efe9838bba3a33ba504665d2", + "visible-value": "f011395cae86a7a736b424d658bbe3afe65ba870" + }, + "consolidation_destinations": { + "agent-discovery-control": "tasks", + "agent-execution-authorization": "tasks", + "agent-recovery-policy": "tasks", + "agent-resource-claims": "tasks", + "agent-routing-policy": "gearbox", + "agent-scheduler-runtime": "tasks", + "agent-state-ledger": "tasks", + "controlled-agent-acceptance": "tasks", + "ephemeral-agent-workers": "tasks", + "event-driven-agent-wakeup": "tasks", + "verifiable-agent-handoff": "tasks" + }, + "workload_ineligible": [ + ".github" + ], + "graphify": { + "role": "OPTIONAL_STANDALONE_CLI", + "tasks_capability": false, + "basis": "Task 21 approved scope fixes the final policy; the historical external-capability acceptance in Tasks commit 00d2bac is superseded. Current bundled manifests contain only Gearbox, Context Firewall, Affected Verification and Visible Value." + } +} diff --git a/program/history/pre-consolidation/ARCHITECTURE.md b/program/history/pre-consolidation/ARCHITECTURE.md new file mode 100644 index 0000000..2044038 --- /dev/null +++ b/program/history/pre-consolidation/ARCHITECTURE.md @@ -0,0 +1,73 @@ +# Cross-project architecture + +The authoritative conceptual topology is +[`program/THEORY_MAP.md`](program/THEORY_MAP.md). The machine-readable mapping is +[`program/theory-registry.json`](program/theory-registry.json). + +```text + PRIMARY DEVELOPER + | + v + AGENT GEARBOX + chooses WHERE work executes + / | \ + deterministic tool bounded helper optional isolated worker + \ | / + v + raw result/evidence + | + v + CONTEXT FIREWALL + chooses WHAT evidence returns + | + v + compact decision evidence + | + Decision Evidence conformance + | + Trajectory / Visible Value telemetry + | + v + PRIMARY DEVELOPER +``` + +Gearbox is a public narrow prototype with deterministic execution and an +injected bounded-helper transport contract. Context Firewall, Decision Evidence +Protocol, Agent Trajectory Profiler, Verifiable Agent Handoff, and Ephemeral +Agent Workers remain external, independently reusable mechanisms. + +Durable orchestration is a separate family: + +```text +Durable Supervisor (objective owner) + | + +-- Agent State Ledger (history/projection) + +-- Agent Scheduler Runtime (readiness/time/pause) + +-- Event-Driven Agent Wakeup (durable activation) + +-- Agent Discovery Control (new-work admission) + +-- Agent Recovery Policy (bounded recovery permission) +``` + +It can own autonomous progress across multiple activations without a +continuously active primary developer. Gearbox performs one bounded delegation +for a primary developer and returns one terminal result. + +Supporting boundaries: + +- Affected Verification independently decides which verification checks are + defensible for a change; humans, CI, or Gearbox may execute the plan, and + Context Firewall may subsequently reduce the resulting evidence. +- Gearbox admits deterministic versus cognitive work; Routing Policy selects an + eligible cognitive route after admission. +- Resource Claims establishes current concurrent ownership; Execution + Authorization validates permission for one exact action. +- Scheduler decides when durable work is ready; Routing decides where it may + run; Recovery decides whether another attempt is justified. +- Verifiable Handoff preserves exact result artifacts across source destruction; + Context Firewall determines which verified facts enter model context. +- Semantic Edit Protocol owns mutation semantics; Agent Trajectory Profiler + measures the resulting trajectory. + +No core primitive depends on a Taslos Tasks database, worker, scheduler, +package, path, or service. Future host-specific adapters remain optional and +versioned. diff --git a/program/history/pre-consolidation/CONCEPTS.md b/program/history/pre-consolidation/CONCEPTS.md new file mode 100644 index 0000000..a47ddeb --- /dev/null +++ b/program/history/pre-consolidation/CONCEPTS.md @@ -0,0 +1,79 @@ +# Concept overview + +This is the bootstrap narrative and subordinate-concept map. The authoritative +current repository inventory and lifecycle state are in +[`program/registry.json`](program/registry.json); the human-readable generated +view is [`PROGRAM_STATUS.md`](PROGRAM_STATUS.md). + +The bootstrap extraction list is not the canonical product/repository map. The +2026-08-29 source reconciliation is in +[`program/THEORY_MAP.md`](program/THEORY_MAP.md), with machine-readable concept +state in [`program/theory-registry.json`](program/theory-registry.json). It +records Agent Gearbox in its restored public home, formally separates it from +Durable Supervisor, and recommends future consolidation for several +durable-orchestration hypotheses without executing those dispositions. + +All current repositories are public, independently versioned, and part of Opsle +Research. That current versioning records provenance; it does not establish that +every hypothesis should remain a standalone implementation repository. Initial +SHA means the first public commit created during the 2026-08-25 bootstrap. + +| Repository | Maturity | Initial SHA | Implementation | Benchmark | +|---|---:|---|---|---| +| [gearbox](https://github.com/opsle/gearbox) | PROTOTYPE | `4d2a7cf902b52f099638b02a4fdec34fd5705a75` | provider-free bounded execution core | one deterministic dogfood fixture and plan; no comparative result | +| [agent-trajectory-profiler](https://github.com/opsle/agent-trajectory-profiler) | PROTOTYPE | `ce2fd532731d9bf5a0b7a271289bfdcc404f57c1` | sanitized dependency-free prototype | plan and fixtures only; no measured result | +| [semantic-edit-protocol](https://github.com/opsle/semantic-edit-protocol) | THEORY | `29caad5c03827cde17aabd71c38bc25899413a33` | theory/specification only | plan and fixtures only; no measured result | +| [durable-supervisor](https://github.com/opsle/durable-supervisor) | THEORY | `555ebedb992ac74236bb7da8230b4d6b0489830b` | theory/specification only | plan and fixtures only; no measured result | +| [event-driven-agent-wakeup](https://github.com/opsle/event-driven-agent-wakeup) | PROTOTYPE | `a6209860c2151450cc28ed648bc8c2631c8db7ef` | sanitized dependency-free prototype | plan and fixtures only; no measured result | +| [context-firewall](https://github.com/opsle/context-firewall) | THEORY | `e865644e86a3f820a120548e91284c048e515671` | theory/specification only | plan and fixtures only; no measured result | +| [decision-evidence-protocol](https://github.com/opsle/decision-evidence-protocol) | PROTOTYPE | `7050f83406a709da84e4b4770556319f767bbeaf` | sanitized dependency-free prototype | plan and fixtures only; no measured result | +| [agent-state-ledger](https://github.com/opsle/agent-state-ledger) | THEORY | `acab03b1ff7168222552050e21e7553b07d00e7c` | theory/specification only | plan and fixtures only; no measured result | +| [agent-scheduler-runtime](https://github.com/opsle/agent-scheduler-runtime) | THEORY | `d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0` | theory/specification only | plan and fixtures only; no measured result | +| [verifiable-agent-handoff](https://github.com/opsle/verifiable-agent-handoff) | PROTOTYPE | `399e5cfae94345affa3f087f0f6eb9e77669d33c` | sanitized dependency-free prototype | plan and fixtures only; no measured result | +| [agent-routing-policy](https://github.com/opsle/agent-routing-policy) | THEORY | `43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8` | theory/specification only | plan and fixtures only; no measured result | +| [agent-resource-claims](https://github.com/opsle/agent-resource-claims) | THEORY | `dfe0fbc90c67ce5ef4256354bb62f1d511b1304c` | theory/specification only | plan and fixtures only; no measured result | +| [agent-discovery-control](https://github.com/opsle/agent-discovery-control) | THEORY | `926547dd9fd713990b1d6f1f2e650aa6c0883564` | theory/specification only | plan and fixtures only; no measured result | +| [agent-execution-authorization](https://github.com/opsle/agent-execution-authorization) | THEORY | `8fb02e943c83fae121c63185d2f0d0dde8c4260a` | theory/specification only | plan and fixtures only; no measured result | +| [controlled-agent-acceptance](https://github.com/opsle/controlled-agent-acceptance) | THEORY | `2d652adf56e53953327d09b1ba9c4a9c3445f052` | theory/specification only | plan and fixtures only; no measured result | +| [agent-recovery-policy](https://github.com/opsle/agent-recovery-policy) | THEORY | `1b733a111e26e0a409fee3b96f627048531daefe` | theory/specification only | plan and fixtures only; no measured result | +| [ephemeral-agent-workers](https://github.com/opsle/ephemeral-agent-workers) | THEORY | `ad96fcfdfac06d340b5e96d369634980cee78ef4` | theory/specification only | plan and fixtures only; no measured result | + +## Ecosystem repositories + +| Repository | Purpose | Initial SHA | +|---|---|---| +| [opsle/site](https://github.com/opsle/site) | Source foundation for future opsle.com; not deployed | `85c3bc7c04f6a2774d589a65df008ad1f8837794` | +| [opsle/.github](https://github.com/opsle/.github) | Organization profile | Recorded in that repository | + +## Subordinate concepts + +These mechanisms remain inside parent projects until they show independent falsifiability, a reusable interface, and meaningful independent install/removal value. + +| Concept | Parent / current role | +|---|---| +| Mutation Amplification | agent-trajectory-profiler metric | +| Edit Payload Amplification | agent-trajectory-profiler metric | +| Semantic Region Revisit Rate | agent-trajectory-profiler metric | +| Change Intent Graph | semantic-edit-protocol mechanism | +| failure-only test reporting | context-firewall reducer | +| semantic Git adapter | decision-evidence/context-firewall adapter | +| semantic shell adapter | decision-evidence/context-firewall adapter | +| provider result envelope | decision-evidence-protocol adapter | +| evidence-gated completion | decision-evidence + state-ledger policy | +| already-satisfied proof | agent-discovery-control mechanism | +| Global Pause | agent-scheduler-runtime safety gate | +| provider cooldown registry | routing/scheduler evidence source | +| execution-target snapshots | authorization/routing binding | +| project configuration snapshots | state-ledger/authorization binding | +| independent reviewer routing | routing/controlled-acceptance mechanism | +| fresh-worker remediation | ephemeral-workers/recovery mechanism | +| objective graph planner | durable-supervisor/discovery mechanism | +| multi-agent write-region coordination | semantic-edit/resource-claims bridge | +| PLAN/DESIGN/BUILD/REVIEW/TEST pipeline | historical scheduler experiment, not final architecture | +| universal redaction | subordinate defense; never a substitute for authorization | +| provider availability countdown | routing UI projection | +| durable route-reason UI | routing evidence presentation | + +## Promotion rule + +Promote only with independent falsifiability, an independent reusable interface, and meaningful independent install/removal value. Rejected and superseded concepts remain recorded. diff --git a/program/history/pre-consolidation/README.md b/program/history/pre-consolidation/README.md new file mode 100644 index 0000000..89beaf7 --- /dev/null +++ b/program/history/pre-consolidation/README.md @@ -0,0 +1,7 @@ +# Historical pre-consolidation snapshot + +These files preserve the ledger and conceptual narrative at research revision +`840c79c04199762acdd59e89e807ca9086f3e45e`. All priorities, next actions, runtime +observations, topology recommendations and stopping criteria here are historical +and superseded by the current program registry. They authorize no work. +Experiment artifacts, source attribution and licenses retain their original identities. diff --git a/program/history/pre-consolidation/THEORY_MAP.md b/program/history/pre-consolidation/THEORY_MAP.md new file mode 100644 index 0000000..eb393ab --- /dev/null +++ b/program/history/pre-consolidation/THEORY_MAP.md @@ -0,0 +1,511 @@ +# Canonical Opsle theory map + +Status: authoritative conceptual reconciliation plus the 2026-08-29 public +Gearbox extraction and 2026-08-31 Affected Verification project creation. +Consolidation and disposition operations remain recommendations only. + +Machine source: [`theory-registry.json`](theory-registry.json). + +Theory registry canonical SHA-256: +`2538ad9b59ebb0cbd34735df8ae3e49ae2f24a1cc98d8f55ac7ee426e6c69e6d`. + +This map corrects an extraction-boundary error. The original 16 concept +repositories were useful hypotheses isolated from one production system, but +hypothesis granularity was treated as repository and product granularity. The +resulting cross-repository diagram then placed almost every concept in one autonomous +supervisor pipeline. Agent Gearbox initially had no canonical home, and shared words such +as *bounded child*, *delegation*, and *waiting without inference* allowed its +primary-developer transmission theory to be conflated with Durable Supervisor. + +The separately authorized Gearbox extraction created one public repository and +assigned its new, evidence-backed implementation `PROTOTYPED`. No existing +repository was promoted, demoted, consolidated, renamed, archived, or deleted. +Existing evidence remains attached to the exact implementation or profile it +actually exercises. + +## Canonical topology + +```text +OPSLE +├── Agent Gearbox public narrow prototype +│ ├── deterministic-versus-cognitive admission +│ ├── gear and model/effort selection +│ ├── content-addressed bounded context +│ ├── deterministic executor / bounded helper executor +│ ├── OS/transport passive wait +│ ├── compact result, literal budget, cleanup, and value telemetry +│ └── consumes external policies/protocols where required +├── Independent Opsle tools +│ ├── Context Firewall +│ ├── Agent Trajectory Profiler +│ ├── Affected Verification +│ └── Semantic Edit Protocol +├── Cross-cutting protocols +│ ├── Decision Evidence Protocol +│ └── Verifiable Agent Handoff +├── Gearbox-facing supporting policy research +│ ├── Agent Routing Policy +│ ├── Agent Execution Authorization +│ └── Agent Resource Claims +├── Durable orchestration research family +│ ├── Durable Supervisor family hypothesis / objective owner +│ ├── Agent State Ledger history and projection module +│ ├── Agent Scheduler Runtime readiness and time module +│ ├── Event-Driven Agent Wakeup durable activation module +│ ├── Agent Discovery Control discovered-work admission module +│ └── Agent Recovery Policy bounded recovery permission module +├── Execution and isolation infrastructure +│ └── Ephemeral Agent Workers +└── Research-only system acceptance + └── Controlled Agent Acceptance +``` + +The tree is conceptual, not a repository-creation plan. In particular, the +Durable Supervisor family has multiple recorded repositories today but likely +fewer future implementation boundaries. Conversely, Context Firewall, Decision +Evidence, Profiler, Semantic Edit, Handoff, and isolation have meaningful reuse +outside either Gearbox or durable orchestration. + +## Agent Gearbox + +> Agent Gearbox lets a powerful primary developer delegate routine operations +> and bounded work to deterministic software or less expensive models, then +> receive only the compact result needed to continue. + +Guiding idea: **Stop using intelligence for work that does not require +intelligence.** + +The primary developer retains end-to-end project understanding, architecture, +safety decisions, integration, ambiguity resolution, and the final definition +of completion. Gearbox is subordinate to that developer and one explicit +bounded call resembling: + +```text +gearbox_run( + task, + task_type, + requested_gear, + allowed_context, + output_contract, + authority, + budget +) +``` + +### Irreducible mechanism + +1. Validate the task, authority, context manifest, requested gear, model and + effort limits, output contract, and literal budget. +2. Admit deterministic software whenever cognition is unnecessary. +3. Admit one fresh bounded helper only when cognition is necessary. +4. Resolve only explicitly allowed, content-addressed context. +5. Execute through a deterministic or bounded cognitive gear. +6. Block at the OS or transport layer until the bounded execution terminates; + the primary model does not poll. +7. Keep raw logs outside primary model context and addressable. +8. Return one compact structured success, failure, indeterminate, or escalation + result. +9. Terminate temporary helpers, reconcile residue, and fail closed if bounded + cleanup cannot be established. +10. Emit exact or observed Visible Value telemetry without inventing token, + cost, latency, or causal savings. + +### Current minimal package boundary + +```text +src/ +├── contract/ request, result, gear, budget, and error schemas +├── admission/ deterministic-vs-cognitive and requested-gear validation +├── context/ content-addressed allow-manifest resolution +├── gears/ +│ ├── deterministic/ +│ └── bounded-helper/ +├── transport/ passive blocking and terminal-result capture +├── budget/ literal provider/session/tool/time accounting +├── cleanup/ helper termination and residue reconciliation +├── result/ compact result assembly and fail-closed escalation +└── telemetry/ Visible Value and trajectory emission hooks +``` + +The public reference core currently implements these responsibilities in a +small Python package rather than one module per diagram node. The diagram is a +responsibility map, not a requirement to split packages or repositories. + +### External dependencies and profiles + +- **Context Firewall** adapts raw tool or helper evidence for primary context. + Gearbox does not own reducer policies. +- **Decision Evidence Protocol** defines or validates compact result facts. + Gearbox owns execution, not general evidence-protocol governance. +- **Agent Trajectory Profiler** consumes execution and Visible Value telemetry. + It never selects a gear. +- **Agent Routing Policy** may supply one Gearbox-specific model, effort, and + provider route after Gearbox admits cognitive execution. +- **Agent Execution Authorization** may validate a capability-style authority + envelope. Reduction and redaction never become authorization. +- **Agent Resource Claims** is optional when shared mutable resources require + leases and fences; it is not needed for every bounded call. +- **Ephemeral Agent Workers** is optional execution/isolation infrastructure for + risky helpers; deterministic local gears need not use it. +- **Verifiable Agent Handoff** is optional when a helper source environment will + be destroyed before independent verification. + +### Explicit non-responsibilities + +Gearbox must not absorb: + +- durable objective ownership, cross-activation reconstruction, queues, + schedules, Global Pause, or autonomous discovery; +- autonomous retry, consultation, alternate-provider fallback, or recovery; +- a persistent hierarchy of agents or a second permanent supervisor; +- exact-session resume or replacement of ChatGPT Remote; +- generic worker isolation, handoff storage, or every Opsle protocol; +- final integration or completion authority from the primary developer. + +### Visible Value contract + +A Gearbox run should expose, when directly observed: + +- requested, admitted, and executed gear; +- deterministic operation identity or helper model/effort identity; +- content manifest identity and permitted bytes/artifacts; +- provider sessions actually started, with missing counts left missing; +- passive-wait duration only when directly measured outside deterministic + semantic output; +- helper termination and residue state; +- raw and returned evidence bytes from compatible Context Firewall receipts; +- exact budget use and typed rejection/escalation reasons. + +It may report a deterministic task completed with zero provider sessions when +that fact is directly recorded. It may not call that value *tokens saved* or +*cost avoided* without a controlled baseline and provider-recorded usage. + +## Gearbox versus Durable Supervisor + +| Dimension | Agent Gearbox | Durable Supervisor / autonomous orchestration | +|---|---|---| +| Primary owner | One continuously responsible powerful primary developer | A durable objective-owning supervisor or orchestrator | +| Persistence | Bounded call state; durable receipts/artifacts as needed | Durable objective, work, event, attempt, and reconstruction state | +| Session lifecycle | One primary-developer activation delegates and receives one terminal result | May stop and resume reasoning across multiple activations | +| Child purpose | Perform one routine operation or bounded cognitive assignment cheaply | Advance autonomous objective state between supervisor decisions | +| Waiting | Synchronous OS/transport blocking without primary-model polling | Persisted wait registration and event-driven reactivation | +| State | Content manifest, request, result, budget, and cleanup state for one call | Ledger projection spanning objectives, tasks, attempts, schedules, and events | +| Retry/recovery | None by default; return failure or indeterminate and fail closed | May use a separately authorized bounded recovery policy across attempts | +| Scheduling | None beyond executing the current bounded call | Queues, dependencies, timeouts, schedules, pause, and readiness transitions | +| Completion | Returns a compact result; the primary developer decides project completion | Records durable task/objective transitions under evidence gates | +| Context ownership | Primary developer owns project context; Gearbox sees only an allowlisted slice | Supervisor reconstructs the model-readable state required for the next activation | +| Intended use | Save the primary developer's intelligence and context on routine bounded work | Operate long-horizon autonomous work despite inactive or interrupted reasoning sessions | +| Failure model | One typed terminal result, escalation, literal budget, helper termination, no hidden fallback | Crash/restart, duplicate/lost events, stale projections, retries, recovery, and terminal orchestration states | + +Canonical invariant: + +> Gearbox enhances a primary developer. Durable orchestration owns progress +> across activations without requiring that developer to remain continuously +> active. + +The earlier wording “can operate across multiple activations” is necessary but +not sufficient. The decisive distinction is **objective ownership**: Gearbox +never takes the primary developer's end-to-end ownership, whereas a Durable +Supervisor owns durable autonomous progress between decisions. + +The prior drift occurred because Durable Supervisor and Event Wakeup also use a +fresh child, bounded assignment, and no model inference while waiting. Their +durable objective, reconstruction, scheduler, and reactivation semantics are the +parts that make them a different theory. + +## Context Firewall + +> Context Firewall is a deterministic boundary that keeps operational noise out +> of an AI agent's context while preserving the compact evidence, provenance, +> and escalation path the agent needs to make correct decisions. + +```text +tool or process + ↓ +raw addressable evidence + ↓ +deterministic versioned adapter and policy + ↓ +compact evidence packet + ↓ +model context +``` + +The packet carries the functional equivalent of schema/policy version, +producer/tool class, source identity and SHA-256, exit state, compact facts, +retained/suppressed accounting, completeness, raw locator, escalation state and +reason, and a deterministic packet identity. + +Context Firewall is not an AI summarizer, sandbox, authorization system, +redaction system, evidence deletion mechanism, supervisor, or proof that model +correctness was preserved. It fails closed or escalates for unsupported formats, +incomplete parsing, uncertain identity, unexpected truncation, unsafe evidence +ceilings, unclassifiable evidence, missing or hash-invalid raw artifacts, and +source/policy incompatibility. + +The current repository implements one strict flat TAP-compatible adapter. That +adapter supports `PROTOTYPED` for its scoped implementation; it does not imply +the adapter families below exist. + +### Intended adapter families + +| Family | Likely compact facts | Unsafe to hide | Escalation conditions | Raw artifact handling | +|---|---|---|---|---| +| Tests | runner/profile, exit/interruption, pass/fail/skip totals, failed identities and complete failure regions, duration when supplied | every failure, crash, timeout, contradiction, flaky/retry marker that changes verdict | unsupported dialect, nested structure not understood, truncated failure, contradictory counts, unexplained exit, ceiling cannot retain failures | retain exact stdout/stderr bytes, framed hash, runner identity, and immutable locator | +| Lint | tool/config/version, file/rule/severity counts, affected paths, fixability, exit state | errors, parser/config failures, suppressed-rule changes, unsafe autofix warnings | unknown formatter, omitted diagnostics, path ambiguity, config mismatch, truncated output | retain complete diagnostics and config identity; packet links line/rule facts to raw bytes | +| Typecheck | compiler/project version, diagnostic codes, files/locations, error counts, build mode | all type errors, project-reference failures, emit-on-error behavior, config resolution defects | unrecognized diagnostic grammar, incremental cache uncertainty, truncated chains, project/config drift | retain compiler streams, project graph/config hashes, and artifact locator | +| Git | repository/worktree identity, HEAD/base/target, status counts, changed paths, diff/stat/ref outcomes | conflicts, unmerged entries, rejected refs, dirty-overlap evidence, signature or object failures, destructive target ambiguity | ambiguous repository, stale refs, truncated diff, unsafe path encoding, missing objects, hash mismatch | retain exact command output and required diff/bundle/object evidence with repository identity | +| Build/compiler | toolchain/config/target, exit state, produced artifact identities, error/warning counts, cache state | compile/link errors, missing outputs, nondeterminism warnings, unsafe fallback, signing or packaging defects | unsupported formatter, missing artifact, cache provenance uncertain, truncated diagnostic, artifact hash mismatch | retain logs plus content-addressed outputs, manifests, and toolchain/config identity | +| Processes/services | command/unit identity, exit/signal, health state, revision, bounded recent error facts | crash loops, unhealthy identity/revision, permission failures, timeouts, resource exhaustion, cleanup residue | uncertain process identity, log gap, unexpected truncation, health/workload disagreement, missing start/end evidence | retain bounded journal/process artifacts with cursor/time window, command/revision identity, and hashes | +| Helper-agent results | task/context/authority/budget/model identity, terminal status, changed artifacts, verification, limitations, cleanup | any failure, uncertainty, unauthorized access, budget breach, incomplete verification, residue, hidden retry/fallback | malformed result contract, context mismatch, missing raw transcript/artifact, provider identity uncertainty, truncated result, helper not terminated | retain raw transcript/logs outside primary context, content-addressed result artifacts, request/response hashes, and cleanup receipt | + +## Seventeen repository classifications + +The classification is conceptual. The disposition is a future recommendation, +not an executed repository action. + +| Repository | Primary classification | Recommended disposition | Confidence | One-sentence rationale | +|---|---|---|---|---| +| `gearbox` | `GEARBOX_CORE` | `KEEP_STANDALONE` | `HIGH` | One bounded primary-developer transmission operation now has a narrow public home without absorbing durable orchestration or external policies and protocols. | +| `affected-verification` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | Minimum-defensible verification planning composes native impact evidence, catalogs, and risk policy without owning check execution or model-context reduction. | +| `agent-trajectory-profiler` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | Reusable correctness-gated telemetry is external to both Gearbox and durable orchestration, although generic Visible Value scope needs ownership cleanup. | +| `semantic-edit-protocol` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | Bounded structural editing is independently useful and can be selected as a deterministic Gearbox tool without becoming Gearbox core. | +| `durable-supervisor` | `DURABLE_ORCHESTRATION` | `KEEP_AS_RESEARCH` | `HIGH` | Durable objective ownership and reconstruction form a distinct autonomous-orchestration hypothesis whose package boundary remains unproven. | +| `event-driven-agent-wakeup` | `DURABLE_ORCHESTRATION` | `CONSOLIDATE_WITH_OTHER` | `MEDIUM_HIGH` | Durable wait registration and event reactivation share the supervisor/scheduler state boundary and are not Gearbox's synchronous passive wait. | +| `context-firewall` | `INDEPENDENT_OPSLE_TOOL` | `KEEP_STANDALONE` | `HIGH` | Deterministic evidence adaptation has standalone value; the current TAP reducer is one adapter, not the whole concept. | +| `decision-evidence-protocol` | `CROSS_CUTTING_PROTOCOL` | `KEEP_AS_PROTOCOL` | `HIGH` | Vendor-neutral evidence representation and independent conformance are reused by many producers and consumers. | +| `agent-state-ledger` | `DURABLE_ORCHESTRATION` | `CONSOLIDATE_WITH_OTHER` | `MEDIUM_HIGH` | The ledger is the Durable Supervisor family's authoritative history/projection module rather than a second product boundary. | +| `agent-scheduler-runtime` | `DURABLE_ORCHESTRATION` | `CONSOLIDATE_WITH_OTHER` | `HIGH` | Readiness, time, pause, and queue transitions belong to the durable-orchestration runtime and currently overlap claims, routing, and recovery. | +| `verifiable-agent-handoff` | `CROSS_CUTTING_PROTOCOL` | `KEEP_AS_PROTOCOL` | `HIGH` | Durable result transfer across source destruction is reusable, while the current HMAC code proves only a seal subcomponent. | +| `agent-routing-policy` | `GEARBOX_SUPPORTING_POLICY` | `FUTURE_GEARBOX_POLICY` | `MEDIUM_HIGH` | Gearbox needs a scoped model/effort/provider route policy after gear admission, but reviewer and fallback routing remain external concerns. | +| `agent-resource-claims` | `GEARBOX_SUPPORTING_POLICY` | `KEEP_AS_RESEARCH` | `MEDIUM` | Leases and fencing can support Gearbox and other systems, but a portable policy boundary and independent packaging value are not yet proven. | +| `agent-discovery-control` | `DURABLE_ORCHESTRATION` | `CONSOLIDATE_WITH_OTHER` | `HIGH` | Autonomous discovered-work admission depends on the supervisor's durable objective graph, ledger, and budgets and is outside explicit-task Gearbox. | +| `agent-execution-authorization` | `GEARBOX_SUPPORTING_POLICY` | `FUTURE_GEARBOX_POLICY` | `HIGH` | Exact capability-style authority is required before Gearbox execution but remains reusable by orchestrators and isolation infrastructure. | +| `controlled-agent-acceptance` | `RESEARCH_ONLY_HYPOTHESIS` | `KEEP_AS_RESEARCH` | `MEDIUM_HIGH` | One-shot autonomy acceptance is an experiment-control method, not a production runtime or Gearbox component. | +| `agent-recovery-policy` | `DURABLE_ORCHESTRATION` | `CONSOLIDATE_WITH_OTHER` | `HIGH` | Recovery permission relies on durable attempts and shared budgets; canonical Gearbox explicitly performs no autonomous retry or fallback. | +| `ephemeral-agent-workers` | `EXECUTION_ISOLATION_INFRASTRUCTURE` | `KEEP_STANDALONE` | `HIGH` | Brokered disposable execution and destruction proof are reusable infrastructure independent of reasoning and orchestration. | + +Detailed original problems, implementation fidelity, drift status, gains, risks, +provenance concerns, relationships, dependencies, consumers, evidence, and +unresolved questions are machine-readable in `theory-registry.json`. + +## Accidental duplication and exact boundaries + +### Gearbox selection versus routing + +Gearbox decides whether the explicit task should use deterministic software or +bounded cognition and validates the requested gear. Routing selects one eligible +model/provider/profile after cognitive admission. Routing does not admit +cognition and neither decision authorizes autonomous fallback. + +### Discovery versus routing + +Discovery decides whether newly proposed autonomous work should exist. Routing +decides where an already admitted cognitive attempt may run. Gearbox accepts an +explicit primary-developer task and creates no recursive work. + +### Authorization versus resource claims + +Resource Claims establishes current ownership of a conflicting resource set +through leases and fences. Execution Authorization decides whether a subject may +perform one exact action and binds current claim identity plus immutable source, +target, purpose, and route authority. Authorization references claims; it does +not reimplement claim lifecycle. + +### Scheduler versus resource claims and routing + +Scheduler owns **when** durable work is ready. Claims owns **who currently +controls** a resource. Routing owns **where/by which eligible executor** a +cognitive attempt may run. Provider availability is route evidence; retry timing +is scheduler/recovery state. + +### Recovery versus routing and acceptance + +Recovery decides whether another attempt can add information or change +conditions and which recovery class is allowed. Routing selects the exact route +only for that separately authorized attempt. Controlled Acceptance preserves +first-failure behavior and cannot silently enable recovery unless its immutable +manifest says so. + +### State Ledger versus Durable Supervisor + +The ledger records immutable facts and projects current state. Durable +Supervisor owns objective decisions based on that projection. No supervisor, +scheduler, or discovery module may maintain a competing authoritative history. + +### Event wakeup versus scheduler + +The scheduler owns wait readiness and timeout policy; the wakeup module persists +registrations and delivers idempotent decision-relevant activation events. One +shared event identity and store contract is required. + +### Handoff versus Decision Evidence and workers + +Decision Evidence owns generic fact/receipt conformance. Handoff adds exact +artifact identity, durable-before-destroy ordering, and fresh reconstruction. +Worker infrastructure enforces containment and lifecycle, then consumes Handoff +for result transfer. A worker must not mint its own authorization or duplicate +the handoff protocol. + +### Context Firewall versus Decision Evidence and Profiler + +Context Firewall reduces and emits. Decision Evidence validates the packet and +its claims independently. Profiler records exposure and value observations. The +Decision Evidence validator may reclassify source bytes for a pinned conformance +profile, but it must not evolve into a competing reduction policy. Profiler must +not become the normative owner of every receipt protocol merely because it +aggregates them. + +### Affected Verification versus selectors, Gearbox, and Context Firewall + +Native affected/related systems remain authoritative evidence providers for +their own graphs and test frameworks. Affected Verification composes that +evidence with a complete verification catalog and explicit risk policy to +produce selected checks, explained skips, and a sufficiency or uncertainty +state. It does not reimplement those selectors, execute CI, or claim global +mathematical minimality. + +Gearbox may consume a plan and choose an execution gear, but does not own the +verification theory or planner. After checks execute, Context Firewall decides +which results enter model context; Affected Verification decides only what +verification should execute. Decision Evidence may validate plan provenance, +and Agent Trajectory Profiler may measure planned versus actual or shadow work, +without initial package coupling. + +## Restored home and remaining gaps + +`opsle/gearbox` now coherently owns deterministic task admission, requested-gear +validation before model/provider routing, content-addressed staged helper +context, exact model and reasoning-effort profiles, one injected bounded helper +transport, OS-level blocking wait, compact result and failure contracts, +literal command/provider/context/output/time budgets, termination checks, raw +artifact accounting, and directly observed provider-session counts. + +The portable core deliberately leaves production helper isolation, provider +routing, general execution authorization, Context Firewall adapters, Decision +Evidence conformance, and trajectory aggregation outside its ownership. A +production helper transport and controlled provider-session/correctness/context +comparison remain implementation and evidence gaps. No concept now lacks a +machine-recorded home, and these responsibilities do not justify additional +repositories. + +## Implementation fidelity and lifecycle scope + +- Agent Gearbox is prototyped for its provider-free public core: deterministic + execution and injected-helper contract paths are runnable and tested, while + no production helper transport, Context Firewall integration, comparative + benchmark, or model/provider subject exists. +- Agent Trajectory Profiler is verified for its implemented metrics, Context + Firewall profile, run records, and Visible Value summaries; predictive or + causal benefit remains unmeasured. +- Context Firewall is prototyped for the TAP-subset adapter. No other adapter + family exists, and no safe correctness frontier is known. +- Decision Evidence is verified for the Context Firewall and Visible Value + profiles. Its generic multi-tool envelope remains narrow. +- Event-Driven Agent Wakeup is prototyped only as an in-memory transition + function; durability, restart, event delivery, and timeout behavior are not + implemented. +- Verifiable Agent Handoff is prototyped only for HMAC manifest binding and + caller-supplied destruction state; publication, transport, destruction proof, + reconstruction, and independent verification are not implemented. +- Affected Verification is verified for its narrow deterministic planner and + the revision-bound AV-EXP-001 Zustand/Vitest shadow calibration. Its frozen + corpus observed 8/8 relevant checks selected by both AV arms, including + conservative full broadening under incomplete impact evidence. The result is + still one repository, one ecosystem, and mostly synthetic change/fault + shapes; no production adapter, independent qualifying replication, or trusted + selective-verification class exists. +- The remaining theory repositories contain coherent falsifiable narratives, + but their `SPEC.md` files are generic templates and their source/tests are + placeholders. Their current lifecycle remains `THEORY`. + +Conceptual reclassification neither promotes nor demotes any repository. A +future consolidation does not combine lifecycle stages arithmetically: evidence +continues to support only its exact historical concept, revision, and scope. + +## EXP-001 reconciliation + +`EXP-001` remains the correct first controlled Context Firewall experiment: + +> How much context can an AI coding agent safely not see? + +Three layers must remain separate: + +1. **Concept definition:** deterministic evidence boundary with provenance, + completeness, raw addressability, and fail-closed escalation. +2. **Current prototype:** one strict TAP-subset adapter and its packet/value + contracts. +3. **Experimental hypothesis:** reduced context preserves task correctness under + equal fixtures, models, prompts, and correctness gates. + +The conformance corpora and Prompt 005 dogfood evidence validate implementation +behavior; they are not an EXP-001 result. The content-addressed task corpus, +correctness oracle, baseline and arms, offline harness, exact model/provider +configuration, blinded allocation, coordinator, and provider-free live +preflight are frozen or qualified at their recorded revisions. `EXP-001` stays +`PLANNED` with zero run identities and zero result artifacts. + +EXP-001 has no technical dependency on Gearbox, and no provider/model subject is +authorized by the provider-free preparation. Controlled experiments are now +LATER program work rather than the immediate execution; the authoritative +current priority is the machine state in `program/registry.json`. + +## Public Opsle and predecessor boundary + +The 2026-08-25 extraction snapshot is historical provenance, not a dependency +direction. The retired predecessor demonstrated that several mechanisms were +possible, but does not establish their general primitives, repository +boundaries, or comparative benefit. + +The long-term direction is: + +```text +public Opsle mechanisms and protocols + ↓ +Opsle Tasks and other products consume pinned versions/adapters +``` + +It is not: + +```text +the retired predecessor remains the canonical implementation + ↓ +public concepts are repeatedly rediscovered after production coupling +``` + +The separately authorized 2026-08-29 extraction adapted the canonical portable +Gearbox core into a public AGPL repository from Taslos Tasks revision +`7734caf208366a0515cf4d78efc17a86363f2238`. The public provenance file records +the exact source and introduction commits. It copied no credentials, private +evidence, host paths, provider configuration, services, databases, product +state, or Durable Supervisor machinery, and the source repository remained +unchanged. Future product adoption still requires public contracts, independent +evidence, explicit versioned adapters, and separate authorization. + +## Future consolidation and provenance policy + +A later consolidation may occur only after a reviewed plan identifies the exact +source repository, target module, preserved commit/history path, lifecycle and +benchmark evidence mapping, and public-link redirect strategy. + +Minimum requirements: + +1. Never delete or silently replace the source repository. +2. Preserve commit attribution and cite the original extraction revision. +3. Preserve theory, specification, benchmark, experiment, negative-result, and + lifecycle history under stable public references. +4. Map each old artifact to a named target module; do not claim the target's + broader lifecycle from narrow source evidence. +5. Publish an archive/redirect note only after the target is available and the + mapping is independently checked. +6. Keep citations and public links working or provide explicit redirects. +7. Record rejected consolidation proposals as research history. +8. Execute repository transfer, archive, rename, or deletion only under new + explicit authority. + +## Repository anti-forgetting invariant + +The authoritative program registry tracks exactly 21 repositories: 18 concept +repositories, `research`, `site`, and `.github`. Agent Gearbox and Affected +Verification are each mapped exactly once to their public repositories. Concept +tracking does not replace repository tracking. diff --git a/program/history/pre-consolidation/registry.json b/program/history/pre-consolidation/registry.json new file mode 100644 index 0000000..cfdc129 --- /dev/null +++ b/program/history/pre-consolidation/registry.json @@ -0,0 +1,766 @@ +{ + "schema_version": 1, + "program": "Opsle Research", + "authoritative_repository_count": 21, + "lifecycle_model": "program/LIFECYCLE.md", + "experiment_registry": "program/experiments.json", + "theory_registry": "program/theory-registry.json", + "theory_map": "program/THEORY_MAP.md", + "theory_reconciliation": { + "status": "IMPLEMENTED_HOME_REGISTERED", + "verified_at": "2026-08-29T02:53:08Z", + "canonical_gearbox_definition": "Agent Gearbox lets a powerful primary developer delegate routine operations and bounded work to deterministic software or less expensive models, then receive only the compact result needed to continue.", + "canonical_context_firewall_definition": "Context Firewall is a deterministic boundary that keeps operational noise out of an AI agent's context while preserving the compact evidence, provenance, and escalation path the agent needs to make correct decisions.", + "gearbox_vs_durable_supervisor": "Gearbox enhances a primary developer; durable orchestration owns autonomous objective progress across durable state and multiple activations without requiring that developer to remain continuously active.", + "gearbox_repository_status": "CREATED_PROTOTYPED", + "repository_topology_operations_executed": true, + "lifecycle_changes_executed": true, + "existing_repository_lifecycle_changes_executed": false, + "existing_repository_dispositions_executed": false, + "model_provider_runs_added": 0 + }, + "gearbox_publication": { + "status": "RELEASED", + "verified_at": "2026-08-29T02:53:08Z", + "repository": "gearbox", + "github_url": "https://github.com/opsle/gearbox", + "pull_request": "https://github.com/opsle/gearbox/pull/1", + "ci_run": "https://github.com/opsle/gearbox/actions/runs/33229696388", + "ci_status": "SUCCESS", + "final_main_sha": "f3fab9f292cf4eabd7200615d444f98881f57d55", + "implementation_revision": "6005b340a6f6fb3f8683439d6b5fd154e1fd253f", + "provider_model_runs": 0, + "repository_consolidations": 0 + }, + "program_control": { + "schema_version": 1, + "priority_order": ["NOW", "NEXT", "THEN", "LATER", "PARKED"], + "current_lane": "NOW", + "operating_question": "What prevents Durable Supervisor from successfully finishing Opsle Tasks with less intelligence, less context, less human involvement, and defensible evidence?", + "exact_next_execution": "In opsle/durable-supervisor, make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening Durable Supervisor v0.1.", + "work_item_admission": { + "rule": "Do not create a work item solely because an implementation can be improved.", + "qualifying_reasons": ["violated invariant", "demonstrated defect", "measured inefficiency", "missing capability blocking the current program objective", "experiment requirement", "security or safety issue", "externally required release condition"], + "parked_by_default": ["cosmetic cleanup", "architectural taste", "hypothetical robustness", "speculative future requirement"] + }, + "lanes": [ + { + "name": "NOW", + "objective": "Finish Durable Supervisor v0.1 as a bounded measured system, then freeze feature work.", + "repositories": ["durable-supervisor", "research"], + "entry_condition": "Current program lane.", + "exit_condition": "Every Durable Supervisor v0.1 stopping criterion is satisfied and the release is explicitly declared and frozen." + }, + { + "name": "NEXT", + "objective": "Use Opsle Tasks as the primary real-world workload and collect integrated measurements from the current implementation.", + "repositories": ["gearbox", "context-firewall", "decision-evidence-protocol", "agent-trajectory-profiler", "affected-verification"], + "entry_condition": "Durable Supervisor v0.1 is declared and feature-frozen.", + "exit_condition": "Durable Supervisor has driven the remaining authorized Opsle Tasks readiness work and produced defensible integrated measurements." + }, + { + "name": "THEN", + "objective": "Advance an individual concept only when Durable Supervisor or Opsle Tasks evidence demonstrates a concrete need.", + "repositories": ["semantic-edit-protocol", "event-driven-agent-wakeup", "agent-state-ledger", "agent-scheduler-runtime", "verifiable-agent-handoff", "agent-routing-policy", "agent-resource-claims", "agent-discovery-control", "agent-execution-authorization", "controlled-agent-acceptance", "agent-recovery-policy", "ephemeral-agent-workers"], + "entry_condition": "A qualifying work-item reason and real workload evidence identify the smallest relevant concept boundary.", + "exit_condition": "The demonstrated deficiency is resolved or falsified at its narrowest justified boundary; do not march mechanically through the portfolio." + }, + { + "name": "LATER", + "objective": "Run controlled experiments, freeze a real-workload benchmark corpus, seek independent replication, and only then consider public product and research-site releases.", + "repositories": ["site"], + "entry_condition": "NOW, NEXT, and evidence-triggered THEN work establish a defensible need and release prerequisites are separately authorized.", + "exit_condition": "Applicable evidence and separate release authorization exist." + }, + { + "name": "PARKED", + "objective": "Retain useful non-priority ideas without turning them into active work.", + "repositories": [".github"], + "entry_condition": "The idea is useful but lacks a qualifying reason to compete with the current objective.", + "exit_condition": "New evidence supplies a qualifying work-item reason." + } + ], + "durable_supervisor_v0_1": { + "status": "IN_PROGRESS", + "durability_foundation": "VERIFIED_ENOUGH_TO_STOP_NITPICKING_UNLESS_REAL_WORKLOAD_EXPOSES_A_DEFECT", + "verified_main_sha": "1b5ab7631ba651a32592bbbdab8001865a3baf3d", + "verified_runtime_state": "PAUSED_NO_ACTIVE_TASK_OR_ATTEMPT", + "runtime_state_recorded_at": "2026-09-05T12:20:01.551Z", + "runtime_state_source": "Read-only .opsle/state.json authority showed supervisor_state PAUSED with null active_task_id and active_attempt_id; .opsle/supervisor.json showed AUTHORITATIVE authority.", + "stopping_criteria": [ + {"id": "DS-V0.1-01", "status": "OPEN", "criterion": "Expose the selected Gearbox route, routing rationale, and Context Firewall accounting in an operator-visible per-child receipt."}, + {"id": "DS-V0.1-02", "status": "OPEN", "criterion": "Expose each child's model and reasoning effort."}, + {"id": "DS-V0.1-03", "status": "OPEN", "criterion": "Record actual tokens and cost when provider evidence supplies them, and clearly labeled estimates otherwise."}, + {"id": "DS-V0.1-04", "status": "OPEN", "criterion": "Measure raw evidence or context versus Context Firewall retained context and report the reduction."}, + {"id": "DS-V0.1-05", "status": "OPEN", "criterion": "Report first-pass success rate."}, + {"id": "DS-V0.1-06", "status": "OPEN", "criterion": "Report repair-child or retry rate and the token cost of retries."}, + {"id": "DS-V0.1-07", "status": "OPEN", "criterion": "Estimate avoidable intelligence consumption by identifying failures whose needed facts were discoverable through deterministic preflight before model execution."}, + {"id": "DS-V0.1-08", "status": "OPEN", "criterion": "Implement only the smallest useful deterministic preflight or reconnaissance mechanism shown necessary by the avoidable-failure evidence."}, + {"id": "DS-V0.1-09", "status": "OPEN", "criterion": "Complete another real foreign-repository workload that exercises and preserves the new measurements."}, + {"id": "DS-V0.1-10", "status": "OPEN", "criterion": "Declare Durable Supervisor v0.1 and freeze feature work except for defects exposed by real workloads."} + ] + }, + "opsle_tasks": { + "current_repository": "opsle/tasks", + "future_name": "Opsle Tasks", + "role": "NEXT primary real-world workload after Durable Supervisor v0.1", + "measurements": ["Gearbox", "Context Firewall", "Decision Evidence Protocol", "Agent Trajectory Profiler", "Affected Verification"], + "prohibited_without_separate_authorization": ["public release", "DNS or TLS changes", "launch provider work"] + }, + "concept_activation": [ + {"deficiency": "routing", "repositories": ["agent-routing-policy"]}, + {"deficiency": "durable state", "repositories": ["agent-state-ledger"]}, + {"deficiency": "retry or deterministic preflight", "repositories": ["gearbox"]}, + {"deficiency": "context overload", "repositories": ["context-firewall"]}, + {"deficiency": "verification selection", "repositories": ["affected-verification"]}, + {"deficiency": "recovery", "repositories": ["agent-recovery-policy"]} + ], + "later_items": ["controlled empirical experiments", "frozen real-workload benchmark corpus", "independent replication", "Opsle Tasks public and self-hosted release", "opsle.com research and public site", "hosted Opsle offering"], + "parked_items": ["Durable Supervisor src/cli.js is 1,570 lines at the verified main SHA and may eventually warrant decomposition; this is a maintenance smell, not a current objective, unless measured work shows material reliability or efficiency harm.", "Background projection repair or reconciliation is not justified merely because explicit retry exists.", "Historical pre-fix cleanup or migration requires evidence of need.", "General architectural polishing remains parked unless a real workload exposes a concrete defect."] + }, + "visible_value": { + "baseline": "The configured primary supervisor model and reasoning effort performs all child work itself and receives raw unfiltered evidence.", + "baseline_rule": "Savings require an inspectable same-work baseline or a clearly labeled estimate derived from that baseline; marketing counterfactuals are prohibited.", + "value_kinds": [ + {"name": "MEASURED", "definition": "The value comes directly from deterministic artifacts or provider usage records and still carries the applicable Visible Value evidence class, normally EXACT or OBSERVED."}, + {"name": "DERIVED", "definition": "The value is a reproducible calculation over identified measured inputs and records its method and applicable Visible Value evidence class."}, + {"name": "ESTIMATED", "definition": "The value carries the ESTIMATED evidence class and names its method, assumptions, pricing source, and uncertainty."}, + {"name": "UNAVAILABLE", "definition": "The value is not currently supported and must remain absent rather than being rendered as zero."} + ], + "per_child_receipt_fields": ["child/task identity", "model", "reasoning effort", "Gearbox route", "routing rationale", "input tokens", "output tokens", "raw evidence/context size", "Context Firewall retained size", "reduction percentage", "estimated tokens avoided", "estimated cost avoided", "duration", "attempt number", "success/failure", "escalation/retry reason", "retry potentially avoidable through deterministic preflight"], + "supervisor_summary_fields": ["total children", "model/effort distribution", "total model tokens consumed", "work completed deterministically without model use", "Context Firewall reduction", "estimated token/cost savings", "first-pass success rate", "repair-child rate", "tokens spent on retries", "avoidable-intelligence estimate"] + }, + "last_verified_at": "2026-09-05T15:28:58Z", + "repositories": [ + { + "name": "agent-trajectory-profiler", + "github_url": "https://github.com/opsle/agent-trajectory-profiler", + "default_branch": "main", + "last_verified_head_sha": "0a89661640721d6a39f127514b993d29bd728d47", + "project_type": "concept", + "purpose": "Measure correctness-gated mutation, edit payload, and semantic-region revisits from observable agent trajectories.", + "lifecycle_stage": "VERIFIED", + "implementation_status": "dependency-free JavaScript trajectory profiler with value-receipt ingestion, observational run records, class/unit/trust-safe per-run and cumulative summaries, Context Firewall adapter, canonical CLI/API, and content-addressed conformance corpus", + "implementation_requirement": "A runnable profiler with documented metric semantics is required.", + "specification_status": "versioned trajectory, measurement, observational run-record, value-receipt ingestion, and deterministic value-summary contracts with trust, class, unit, escalation, missing-data, deduplication, and invalid-state semantics", + "test_status": "74 of 74 automated tests and 14 of 14 content-addressed measurement fixtures passed locally at the verified HEAD; 5 of 5 pinned exact-revision interoperability cases and 5 determinism tests passed; PR #2 CI passed", + "benchmark_status": "14 deterministic measurement fixtures, a public-synthetic two-receipt interoperability chain, one observational implementation-run record, the research-owned frozen EXP-001 corpus/oracle/arm/provider-free trajectory harness at 04234a65bf36192d63f1dd173c440d45a6604d2b, launch preregistration at 31848c3f25ff9371055932657e8e2f8ad54cc8c7, and eight coordinator qualification profiles at 9ee43197880c18d4e185cf7e29e02a151d22a12e; no model run", + "measured_experiment_status": "EXP-001 remains planned; zero model/provider experiment runs", + "reproducibility_status": "tests, canonical conformance, byte-identical fixture regeneration, value receipt/run-record validation, deterministic summaries, and exact-revision Context Firewall/Decision Evidence interoperability are reproducible locally and in CI; comparative benefit and model correctness are untested", + "documentation_status": "public theory, versioned specification, architecture, byte/event/token semantics, Visible Value ingestion and aggregation rules, observational corpus fields, operator channel separation, fixture provenance, EXP-001 boundary, and limitations documented", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Receipt aggregation requires explicit safe declarations and exact compatibility partitions; ordinary records reject EXPERIMENTAL measurements, only Context Firewall packet-v1 has a trajectory adapter, and model correctness, token/cost savings, causal benefit, a safe reduction frontier, and replication remain unmeasured."], + "dependencies": [], + "dependents": ["context-firewall", "gearbox", "semantic-edit-protocol"], + "active_experiment_ids": ["EXP-001"], + "blockers": ["EXP-001 exact budgeted authorization, catalogue/pricing preflight, measured experiment, and independent replication remain missing; the one-block coordinator is provider-free qualified and zero model/provider subjects have run."], + "next_task": "Wait for a Durable Supervisor or Opsle Tasks workload to require trajectory measurement; do not advance the profiler merely to polish it.", + "evidence": ["https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/value-summary.js", "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/run-record.js", "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/tests/value-summary.test.js", "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/tools/verify-context-firewall-interop.js", "https://github.com/opsle/agent-trajectory-profiler/pull/2", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/profile.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", "https://github.com/opsle/research/pull/11", "https://github.com/opsle/research/pull/13"], + "completion_criteria": ["Publish metric semantics and executable profiler.", "Run correctness-gated benchmarks across multiple trajectories.", "Replicate predictive-value claims and document failure modes."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "semantic-edit-protocol", + "github_url": "https://github.com/opsle/semantic-edit-protocol", + "default_branch": "main", + "last_verified_head_sha": "29caad5c03827cde17aabd71c38bc25899413a33", + "project_type": "concept", + "purpose": "Apply bounded semantic edit operations with structural and hash preconditions.", + "lifecycle_stage": "THEORY", + "implementation_status": "none; placeholder source directory only", + "implementation_requirement": "A reference language adapter and transactional validator are required.", + "specification_status": "experimental theory contract; concept-specific operation schema remains incomplete", + "test_status": "placeholder only; no automated tests", + "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "none", + "reproducibility_status": "not available", + "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Language adapters, conflict semantics, anchor durability measurements, and controlled patch comparisons are missing."], + "dependencies": ["agent-resource-claims", "agent-trajectory-profiler"], + "dependents": [], + "active_experiment_ids": [], + "blockers": ["No executable semantic operation, validator, tests, or benchmark fixtures."], + "next_task": "Wait until a real Durable Supervisor or Opsle Tasks workload demonstrates a semantic-edit deficiency that simpler bounded edits cannot satisfy.", + "evidence": ["https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/THEORY.md", "https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/SPEC.md"], + "completion_criteria": ["Publish a precise operation contract and executable reference adapter.", "Benchmark against patch and whole-file baselines under identical correctness gates.", "Replicate results and document conflicts and unsupported syntax."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "durable-supervisor", + "github_url": "https://github.com/opsle/durable-supervisor", + "default_branch": "main", + "last_verified_head_sha": "1b5ab7631ba651a32592bbbdab8001865a3baf3d", + "project_type": "concept", + "purpose": "Own autonomous objective progress through durable authority, bounded child execution, evidence reduction, evaluation, pause, wake, and reconstruction across model activations.", + "lifecycle_stage": "VERIFIED", + "implementation_status": "dependency-free Node.js Durable Supervisor v0.1 with persistent supervisor authority, task handoff, discovery, exact Gearbox routing, claims and fencing, detached Runner ownership, Context Firewall packets, acceptance and immutable evaluation, opsled wake delivery, bounded reconstruction, policy controls, and runtime-release fencing", + "implementation_requirement": "A durable supervisor/runner reference implementation is required.", + "specification_status": "public v0.1 contract defines durable authority, immutable history, task and attempt state, exact routes, claims and fencing, Runner ownership, Context Firewall and acceptance boundaries, evaluation, wake, reconstruction, pause, and runtime compatibility", + "test_status": "PR #2 recorded 141 full-suite tests and canonical validation passing; PR #3 recorded 176 passing with 1 intentional skip; D5 PR #4 passed 6 focused atomicity, operator, and wake checks; projection-reconciliation PR #5 passed 76 focused evaluation, invariant, operator, recovery, and wake checks plus syntax, release-manifest, and canonical validation at the exact head whose tree is main", + "benchmark_status": "bounded self-hosting evidence records meaningful child work, exact route and acceptance artifacts, and measured Context Firewall byte reduction; PR #2 additionally records disposable and real foreign-repository portability plus a live automatic wake, but no frozen comparative baseline or controlled token/cost benchmark exists", + "measured_experiment_status": "operational self-hosting and foreign-repository observations only; no controlled comparative experiment", + "reproducibility_status": "provider-free tests, schema and release checks, reconstruction, portability fixtures, concurrency campaigns, and explicit evaluation replay are reproducible from the public source; latest PRs used focused affected verification rather than a new full-suite run, and no independent replication exists", + "documentation_status": "public v0.1 README, normative specification, architecture, operations guidance, and bounded self-hosting proof document supported behavior and explicit claim limits", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Operator receipts do not yet expose the full Gearbox route/rationale, child model/effort, actual token/cost, retry cost, or avoidable-intelligence accounting required for v0.1 completion; existing Context Firewall evidence is byte-level, the latest repairs used focused rather than full-suite verification, and comparative benefit and independent replication remain unproven."], + "dependencies": ["agent-state-ledger", "agent-scheduler-runtime", "event-driven-agent-wakeup", "decision-evidence-protocol"], + "dependents": [], + "active_experiment_ids": [], + "blockers": ["Durable Supervisor v0.1 still lacks the operator-visible measurement surface, avoidable-intelligence classification, smallest evidence-driven preflight, and another measured foreign-repository workload required by the program stopping criteria."], + "next_task": "Make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening v0.1.", + "evidence": ["https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/README.md", "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/SPEC.md", "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/ARCHITECTURE.md", "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/docs/SELF_HOSTING_PROOF.md", "https://github.com/opsle/durable-supervisor/pull/2", "https://github.com/opsle/durable-supervisor/pull/3", "https://github.com/opsle/durable-supervisor/pull/4", "https://github.com/opsle/durable-supervisor/pull/5"], + "completion_criteria": ["Satisfy the ten bounded Durable Supervisor v0.1 stopping criteria in the program-control registry.", "Prove the measurement surface on another real foreign-repository workload.", "Declare v0.1 and freeze feature work except for workload-exposed defects."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "active", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "event-driven-agent-wakeup", + "github_url": "https://github.com/opsle/event-driven-agent-wakeup", + "default_branch": "main", + "last_verified_head_sha": "a6209860c2151450cc28ed648bc8c2631c8db7ef", + "project_type": "concept", + "purpose": "Suspend model activity during waits and wake idempotently on durable decision-relevant events.", + "lifecycle_stage": "PROTOTYPED", + "implementation_status": "dependency-free JavaScript state-machine prototype", + "implementation_requirement": "A durable adapter plus restart-safe reference implementation is required.", + "specification_status": "experimental prototype contract", + "test_status": "2 of 2 automated tests passed locally at the verified HEAD", + "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "none", + "reproducibility_status": "prototype tests reproducible locally; inference-avoidance benefit not reproduced", + "documentation_status": "public theory, specification, architecture, and benchmark plan present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Restart campaigns, duplicate-event campaigns, wake-latency baselines, and multi-runtime adapters are missing."], + "dependencies": ["agent-state-ledger"], + "dependents": ["durable-supervisor"], + "active_experiment_ids": [], + "blockers": ["Current prototype is in-memory and has no restart or durable-store harness."], + "next_task": "Wait for a real Durable Supervisor or Opsle Tasks wake deficiency; advance the standalone concept only if the integrated runtime exposes one.", + "evidence": ["https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/src/runtime.js", "https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/tests/runtime.test.js"], + "completion_criteria": ["Implement restart-safe durable wait registration and wakeup.", "Benchmark polling and model-turn baselines.", "Replicate event-loss and duplicate-delivery correctness."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "context-firewall", + "github_url": "https://github.com/opsle/context-firewall", + "default_branch": "main", + "last_verified_head_sha": "953c48f1cfd154d6b7ed10b51b87fe54e4df45f2", + "project_type": "concept", + "purpose": "Reduce operational payload deterministically while preserving provenance and safe raw-evidence escalation.", + "lifecycle_stage": "PROTOTYPED", + "implementation_status": "dependency-free deterministic JavaScript TAP-subset reducer with compact evidence packets, separate opsle.value-receipt.v1 sidecars, named operator indicators, payload ceilings, raw escalation, and synthetic conformance corpus", + "implementation_requirement": "An executable reducer and conformance suite are sufficient; a full agent runtime is not required.", + "specification_status": "experimental prototype contract for packet v1, deterministic TAP-subset policy v1, Visible Value receipt profile, deterministic sidecar, and operator/model channel separation", + "test_status": "42 of 42 automated tests and 30 of 30 synthetic conformance fixtures passed locally at the verified HEAD; deterministic stdout/sidecar separation and operator stderr tests passed; PR #2 CI passed", + "benchmark_status": "30 synthetic conformance fixtures, one observational dogfood reduction attempt, the research-owned frozen six-task EXP-001 corpus/oracle plus raw and three strict-TAP arm harness at 04234a65bf36192d63f1dd173c440d45a6604d2b, launch preregistration at 31848c3f25ff9371055932657e8e2f8ad54cc8c7, and six coordinator qualification reductions at 9ee43197880c18d4e185cf7e29e02a151d22a12e; no model run", + "measured_experiment_status": "EXP-001 remains planned; zero model/provider experiment runs", + "reproducibility_status": "prototype tests, receipt validation, canonical output checks, and conformance are reproducible locally and in CI; model correctness and a safe reduction frontier are untested", + "documentation_status": "public theory, packet/input contract, parser limits, retention/suppression policy, exact Visible Value measurements, operator channel, escalation, ceilings, CLI, conformance, claim limits, and EXP-001 boundary documented", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["The parser is a strict flat TAP-compatible subset; actual Node checkmark-formatted test output dogfood expanded from 3,425 to 13,027 bytes and correctly escalated, caller raw references are not externally verified, and no evidence establishes model correctness or a safe reduction frontier."], + "dependencies": ["decision-evidence-protocol", "agent-trajectory-profiler"], + "dependents": ["gearbox"], + "active_experiment_ids": ["EXP-001"], + "blockers": ["EXP-001 exact budgeted authorization, catalogue/pricing preflight, measured correctness experiment, and replication remain missing; the one-block coordinator is provider-free qualified and zero model/provider subjects have run."], + "next_task": "Wait for Durable Supervisor or Opsle Tasks evidence of context overload, unsafe omission, or unsupported input before advancing the standalone reducer.", + "evidence": ["https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/src/value-receipt.js", "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/tests/reducer.test.js", "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/fixtures/corpus.js", "https://github.com/opsle/context-firewall/pull/2", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/context-firewall-value-receipt.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", "https://github.com/opsle/research/pull/11", "https://github.com/opsle/research/pull/13"], + "completion_criteria": ["Publish reducer policy and executable conformance validator.", "Run correctness-gated measured context-reduction experiments.", "Replicate the safe frontier and publish omission failure modes."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "decision-evidence-protocol", + "github_url": "https://github.com/opsle/decision-evidence-protocol", + "default_branch": "main", + "last_verified_head_sha": "b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09", + "project_type": "concept", + "purpose": "Represent bounded vendor-neutral decision facts with provenance and explicit raw-output escalation.", + "lifecycle_stage": "VERIFIED", + "implementation_status": "dependency-free generic value-receipt validator plus independent Context Firewall packet/source/value-receipt validator, Decision Evidence validation receipts, canonical CLI with named operator indicators, and self-contained conformance suite", + "implementation_requirement": "As a protocol, an executable validator and conformance suite can satisfy implementation; a full runtime is not required.", + "specification_status": "versioned normative experimental Context Firewall packet-v1 and value-receipt validation profiles plus generic opsle.value-receipt.v1 semantic validation", + "test_status": "60 of 60 automated tests and 24 of 24 public-safe conformance vectors passed locally at the verified HEAD; 5 of 5 exact-revision interoperability cases, 8 determinism tests, operator separation, source-unverified, and tamper paths passed; PR #2 CI passed", + "benchmark_status": "24 protocol conformance vectors, a 5-case exact-revision interoperability proof, one source-backed observational dogfood validation, 36 source-backed validations in the research-owned frozen EXP-001 offline harness at 04234a65bf36192d63f1dd173c440d45a6604d2b, launch preregistration at 31848c3f25ff9371055932657e8e2f8ad54cc8c7, and six coordinator qualification validations at 9ee43197880c18d4e185cf7e29e02a151d22a12e; no correctness experiment or model run", + "measured_experiment_status": "none", + "reproducibility_status": "validator tests, self-contained vectors, deterministic CLI/sidecar output, value-receipt validation, and exact-revision Context Firewall interoperability are reproducible locally and in CI; other tool classes and independent replication are missing", + "documentation_status": "public theory, packet-v1 and value-receipt profiles, enforced invariants, API/CLI, named operator channel, structural-versus-cryptographic verification, sufficiency/escalation semantics, claim limits, interoperability, and limitations documented", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Only Context Firewall's strict TAP-subset packet-v1 producer is independently validated; receipt-only source claims and caller-owned raw locator existence remain unverified without external evidence, and no result establishes model correctness, comparative benefit, or independent replication."], + "dependencies": [], + "dependents": ["agent-recovery-policy", "context-firewall", "durable-supervisor", "gearbox", "verifiable-agent-handoff"], + "active_experiment_ids": ["EXP-001"], + "blockers": ["Additional real tool classes, measured subject decision adequacy, EXP-001 exact budgeted authorization and catalogue/pricing preflight, a controlled comparative result, and independent replication are missing."], + "next_task": "Wait for an integrated Durable Supervisor or Opsle Tasks receipt to expose a decision-evidence gap that blocks a defensible decision.", + "evidence": ["https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/context-firewall-value.js", "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/value-receipt-v1.js", "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/tests/value-receipt.test.js", "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/docs/value-receipt-v1.md", "https://github.com/opsle/decision-evidence-protocol/pull/2", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/decision-evidence-validation.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", "https://github.com/opsle/research/pull/11", "https://github.com/opsle/research/pull/13"], + "completion_criteria": ["Publish a versioned protocol and comprehensive executable conformance suite.", "Measure adequacy across multiple real tool classes.", "Replicate interoperability and document loss/escalation failure modes."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "agent-state-ledger", + "github_url": "https://github.com/opsle/agent-state-ledger", + "default_branch": "main", + "last_verified_head_sha": "acab03b1ff7168222552050e21e7553b07d00e7c", + "project_type": "concept", + "purpose": "Record append-only agent work evidence and derive deterministic current-state projections.", + "lifecycle_stage": "THEORY", + "implementation_status": "none; placeholder source directory only", + "implementation_requirement": "A portable append-only store and projection validator are required.", + "specification_status": "experimental theory contract; schema and contradiction semantics remain incomplete", + "test_status": "placeholder only; no automated tests", + "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "none", + "reproducibility_status": "not available", + "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Portable schema, contradiction semantics, projection correctness, and reconstruction-size benchmarks are missing."], + "dependencies": [], + "dependents": ["agent-discovery-control", "agent-execution-authorization", "agent-recovery-policy", "agent-scheduler-runtime", "durable-supervisor", "event-driven-agent-wakeup"], + "active_experiment_ids": [], + "blockers": ["No portable schema, implementation, projection oracle, or replay fixtures."], + "next_task": "Wait for a real Durable Supervisor or Opsle Tasks workload to expose a portable durable-state deficiency before advancing this repository.", + "evidence": ["https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/THEORY.md", "https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/SPEC.md"], + "completion_criteria": ["Implement append-only storage and deterministic projection.", "Prove replay/reconstruction against failure fixtures.", "Measure and replicate reconstruction size and correctness."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "agent-scheduler-runtime", + "github_url": "https://github.com/opsle/agent-scheduler-runtime", + "default_branch": "main", + "last_verified_head_sha": "d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0", + "project_type": "concept", + "purpose": "Own queues, dependencies, schedules, leases, cancellation, and pause as deterministic runtime mechanics.", + "lifecycle_stage": "THEORY", + "implementation_status": "none; placeholder source directory only", + "implementation_requirement": "A portable deterministic scheduler core and store adapter are required.", + "specification_status": "experimental theory contract; portable store and fairness semantics remain incomplete", + "test_status": "placeholder only; no automated tests", + "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "none", + "reproducibility_status": "not available", + "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Portable store interface, fairness/latency evidence, reasoning boundary, and external adapters are missing."], + "dependencies": ["agent-resource-claims", "agent-state-ledger"], + "dependents": ["controlled-agent-acceptance", "durable-supervisor"], + "active_experiment_ids": [], + "blockers": ["Ledger and claim interfaces are not portable or executable in these repositories."], + "next_task": "Wait for a real workload to expose a scheduler-runtime deficiency not already satisfied by Durable Supervisor's local implementation.", + "evidence": ["https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/THEORY.md", "https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/BENCHMARK.md"], + "completion_criteria": ["Implement the portable deterministic core and store contract.", "Pass restart, lease, pause, cancellation, duplicate, and fairness campaigns.", "Benchmark and replicate latency and correctness."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "verifiable-agent-handoff", + "github_url": "https://github.com/opsle/verifiable-agent-handoff", + "default_branch": "main", + "last_verified_head_sha": "399e5cfae94345affa3f087f0f6eb9e77669d33c", + "project_type": "concept", + "purpose": "Seal exact execution results before source destruction so a fresh verifier can reconstruct and validate them.", + "lifecycle_stage": "PROTOTYPED", + "implementation_status": "dependency-free JavaScript HMAC manifest prototype", + "implementation_requirement": "As a protocol, an executable sealing/verifying library and adversarial conformance suite can satisfy implementation.", + "specification_status": "experimental prototype contract", + "test_status": "3 of 3 automated tests passed locally at the verified HEAD", + "benchmark_status": "prose plan only; no artifact reconstruction harness or frozen fixtures", + "measured_experiment_status": "EXP-001 potential support only; no runs", + "reproducibility_status": "prototype tests reproducible locally; cross-environment handoff not reproduced", + "documentation_status": "public theory, specification, architecture, and benchmark plan present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Independent implementations, cryptographic agility, cross-VCS artifacts, byte ceilings, and correlated-error studies are missing."], + "dependencies": ["decision-evidence-protocol"], + "dependents": ["controlled-agent-acceptance", "ephemeral-agent-workers"], + "active_experiment_ids": ["EXP-001"], + "blockers": ["Prototype authenticates a manifest but does not build, transport, or reconstruct artifacts."], + "next_task": "Wait for a real workload to require evidence survival across an isolation or source-destruction boundary before advancing this protocol.", + "evidence": ["https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/src/seal.js", "https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/tests/seal.test.js"], + "completion_criteria": ["Publish a versioned seal/reconstruction contract and executable conformance suite.", "Benchmark full artifact handoff and failure cases.", "Replicate with an independent implementation and environment."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "agent-routing-policy", + "github_url": "https://github.com/opsle/agent-routing-policy", + "default_branch": "main", + "last_verified_head_sha": "43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8", + "project_type": "concept", + "purpose": "Select purpose-bound agent routes under exact capability, access, independence, budget, and availability constraints.", + "lifecycle_stage": "THEORY", + "implementation_status": "none; placeholder source directory only", + "implementation_requirement": "An executable deterministic policy evaluator and trace validator are sufficient.", + "specification_status": "experimental theory contract; route schema and ordering semantics remain incomplete", + "test_status": "placeholder only; no automated tests", + "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "none", + "reproducibility_status": "not available", + "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Quality-adjusted price evidence, provider comparability, portability, and stability/fairness analysis are missing."], + "dependencies": [], + "dependents": ["agent-recovery-policy", "controlled-agent-acceptance", "gearbox"], + "active_experiment_ids": [], + "blockers": ["No normative route schema, evaluator, fixtures, or provider-independent quality evidence."], + "next_task": "Wait for measured Durable Supervisor routing deficiencies before advancing a standalone routing policy.", + "evidence": ["https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/THEORY.md", "https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/SPEC.md"], + "completion_criteria": ["Publish route schema and executable policy evaluator.", "Measure strict and adaptive routing against real baselines.", "Replicate quality/cost claims and document unstable routes."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "agent-resource-claims", + "github_url": "https://github.com/opsle/agent-resource-claims", + "default_branch": "main", + "last_verified_head_sha": "dfe0fbc90c67ce5ef4256354bb62f1d511b1304c", + "project_type": "concept", + "purpose": "Coordinate atomic resource claim sets with leases and fencing so stale actors lose authority.", + "lifecycle_stage": "THEORY", + "implementation_status": "none; placeholder source directory only", + "implementation_requirement": "A reference claim-set evaluator with lease/fence state machine is required.", + "specification_status": "experimental theory contract; resource identity and conflict semantics remain incomplete", + "test_status": "placeholder only; no automated tests", + "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "none", + "reproducibility_status": "not available", + "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Deadlock analysis, semantic-region adapters, fairness benchmarks, and interoperability are missing."], + "dependencies": [], + "dependents": ["agent-execution-authorization", "agent-scheduler-runtime", "ephemeral-agent-workers", "semantic-edit-protocol"], + "active_experiment_ids": [], + "blockers": ["No portable resource catalog, claim-set state machine, concurrency tests, or fairness fixtures."], + "next_task": "Wait for a workload-exposed resource-claim or fencing deficiency not already covered by Durable Supervisor's local claims.", + "evidence": ["https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/THEORY.md", "https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/ARCHITECTURE.md"], + "completion_criteria": ["Implement atomic claims, expiry, renewal, takeover, and fencing.", "Pass deadlock, crash, stale-actor, and fairness campaigns.", "Replicate across at least two store/runtime adapters."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "agent-discovery-control", + "github_url": "https://github.com/opsle/agent-discovery-control", + "default_branch": "main", + "last_verified_head_sha": "926547dd9fd713990b1d6f1f2e650aa6c0883564", + "project_type": "concept", + "purpose": "Admit, merge, ignore, or escalate agent-discovered work under provenance and storm budgets.", + "lifecycle_stage": "THEORY", + "implementation_status": "none; placeholder source directory only", + "implementation_requirement": "An executable admission policy and convergence simulator are sufficient.", + "specification_status": "experimental theory contract; proposal similarity and budget semantics remain incomplete", + "test_status": "placeholder only; no automated tests", + "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "none", + "reproducibility_status": "not available", + "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Semantic duplicate benchmarks, multi-agent convergence, calibrated confidence, and complexity budgets are missing."], + "dependencies": ["agent-state-ledger"], + "dependents": [], + "active_experiment_ids": [], + "blockers": ["No durable proposal dataset, similarity oracle, policy evaluator, or convergence harness."], + "next_task": "Wait for real duplicate-discovery or already-satisfied work evidence before advancing this repository.", + "evidence": ["https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/THEORY.md", "https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/BENCHMARK.md"], + "completion_criteria": ["Implement deterministic proposal admission and convergence.", "Benchmark duplicates, storms, and already-satisfied proofs.", "Replicate calibrated policies across objective sizes."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "agent-execution-authorization", + "github_url": "https://github.com/opsle/agent-execution-authorization", + "default_branch": "main", + "last_verified_head_sha": "8fb02e943c83fae121c63185d2f0d0dde8c4260a", + "project_type": "concept", + "purpose": "Bind execution authority to exact jobs, leases, targets, resources, evidence, routes, and purposes.", + "lifecycle_stage": "THEORY", + "implementation_status": "none; placeholder source directory only", + "implementation_requirement": "As an authorization abstraction, an executable grant validator and revocation conformance suite are sufficient.", + "specification_status": "experimental theory contract; portable grant and composition semantics remain incomplete", + "test_status": "placeholder only; no automated tests", + "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "none", + "reproducibility_status": "not available", + "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Portable grant format, revocation latency, cross-orchestrator subjects, and authority composition are missing."], + "dependencies": ["agent-resource-claims", "agent-state-ledger"], + "dependents": ["controlled-agent-acceptance", "ephemeral-agent-workers", "gearbox"], + "active_experiment_ids": [], + "blockers": ["No grant schema, current-authority oracle, revocation model, or adversarial fixtures."], + "next_task": "Wait for a real execution-authorization gap that blocks the current Durable Supervisor or Opsle Tasks objective.", + "evidence": ["https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/THEORY.md", "https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/SPEC.md"], + "completion_criteria": ["Publish a portable grant schema and executable validator.", "Pass stale, revoked, drifted, replayed, and composed-authority cases.", "Measure and replicate revocation and interoperability behavior."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "controlled-agent-acceptance", + "github_url": "https://github.com/opsle/controlled-agent-acceptance", + "default_branch": "main", + "last_verified_head_sha": "2d652adf56e53953327d09b1ba9c4a9c3445f052", + "project_type": "concept", + "purpose": "Run exact one-shot autonomy acceptance under immutable authority, provider budgets, pause restoration, and evidence retention.", + "lifecycle_stage": "THEORY", + "implementation_status": "none; placeholder source directory only", + "implementation_requirement": "A provider-independent controller simulator and manifest validator must precede any real-provider harness.", + "specification_status": "experimental theory contract; manifest and interruption semantics remain incomplete", + "test_status": "placeholder only; no automated tests", + "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "none in this repository", + "reproducibility_status": "not available", + "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Independent harness, interruption model, non-reusable grants, and cross-runtime replication are missing."], + "dependencies": ["agent-execution-authorization", "agent-routing-policy", "agent-scheduler-runtime", "verifiable-agent-handoff"], + "dependents": [], + "active_experiment_ids": [], + "blockers": ["Authorization, routing, scheduling, and handoff contracts are not yet executable together."], + "next_task": "Wait for real acceptance evidence to show a missing portable capability before advancing the standalone concept.", + "evidence": ["https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/THEORY.md", "https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/ARCHITECTURE.md"], + "completion_criteria": ["Implement and verify a one-shot fail-closed acceptance controller.", "Run bounded acceptance experiments only under separately authorized exact manifests.", "Replicate interruption, restoration, and residue reconciliation behavior."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "agent-recovery-policy", + "github_url": "https://github.com/opsle/agent-recovery-policy", + "default_branch": "main", + "last_verified_head_sha": "1b733a111e26e0a409fee3b96f627048531daefe", + "project_type": "concept", + "purpose": "Choose bounded retries, consultations, alternate routes, fallbacks, or terminal outcomes from typed failure evidence.", + "lifecycle_stage": "THEORY", + "implementation_status": "none; placeholder source directory only", + "implementation_requirement": "An executable deterministic policy evaluator and failure-fixture suite are sufficient.", + "specification_status": "experimental theory contract; failure taxonomy and information-gain semantics remain incomplete", + "test_status": "placeholder only; no automated tests", + "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "none", + "reproducibility_status": "not available", + "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Outcome calibration, cross-provider taxonomy, information-gain estimation, and comparative experiments are missing."], + "dependencies": ["agent-routing-policy", "agent-state-ledger", "decision-evidence-protocol"], + "dependents": [], + "active_experiment_ids": [], + "blockers": ["No shared failure schema, attempt ledger, route evaluator, or comparative fixture set."], + "next_task": "Wait for repeated real recovery failures to demonstrate a policy deficiency; do not create speculative recovery work.", + "evidence": ["https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md", "https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/BENCHMARK.md"], + "completion_criteria": ["Publish a typed failure schema and executable bounded policy.", "Compare retry, consultation, alternate route, and terminal baselines.", "Replicate outcome and convergence claims across failure classes."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "ephemeral-agent-workers", + "github_url": "https://github.com/opsle/ephemeral-agent-workers", + "default_branch": "main", + "last_verified_head_sha": "ad96fcfdfac06d340b5e96d369634980cee78ef4", + "project_type": "concept", + "purpose": "Execute in disposable bounded workers through a narrow broker, seal results, and prove destruction.", + "lifecycle_stage": "THEORY", + "implementation_status": "none; placeholder source directory only", + "implementation_requirement": "A portable worker/broker reference adapter and containment test harness are required.", + "specification_status": "experimental theory contract; broker and destruction-proof formats remain incomplete", + "test_status": "placeholder only; no automated tests", + "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "none in this repository", + "reproducibility_status": "not available", + "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["Portable isolation adapters, quantitative containment, kernel threat model, and destruction-proof interoperability are missing."], + "dependencies": ["agent-execution-authorization", "agent-resource-claims", "verifiable-agent-handoff"], + "dependents": [], + "active_experiment_ids": [], + "blockers": ["Portable authority, claim, and handoff contracts are not ready; no safe synthetic containment harness exists."], + "next_task": "Wait for a real workload to require an ephemeral-worker boundary before advancing this repository.", + "evidence": ["https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md", "https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/ARCHITECTURE.md"], + "completion_criteria": ["Implement a portable bounded broker/worker adapter and destruction proof.", "Pass containment, credential, process, network, seal, and cleanup campaigns.", "Replicate across isolation technologies with a published threat model."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "gearbox", + "github_url": "https://github.com/opsle/gearbox", + "default_branch": "main", + "last_verified_head_sha": "f3fab9f292cf4eabd7200615d444f98881f57d55", + "project_type": "concept", + "purpose": "Let one powerful primary developer execute deterministic operations or one bounded cognitive assignment through an authority-, context-, result-, and budget-constrained transmission boundary.", + "lifecycle_stage": "PROTOTYPED", + "implementation_status": "provider-free Python reference core with strict authority-policy admission, exact deterministic argv execution, content-addressed staged helper context, injected one-shot helper transport, passive process waiting, compact results, raw-artifact accounting, fail-closed budgets, and Visible Value receipts", + "implementation_requirement": "A runnable provider-free core plus deterministic and injected-helper conformance paths is sufficient for the prototype gate; production provider integration and comparative benefit require later evidence.", + "specification_status": "versioned request, policy, result, context-selection, output-contract, budget, helper-transport, raw-artifact, cleanup, and Visible Value boundaries are public at the verified revision", + "test_status": "19 of 19 provider-free automated tests passed locally, in PR #1 CI, and in final-main CI; ruff, shellcheck, actionlint, gitleaks, wheel build, receipt validation, and public raw-locator/hash checks passed", + "benchmark_status": "one revision-bound deterministic dogfood fixture with public compact result, raw artifacts, and value receipt plus exact-revision Decision Evidence validation and Trajectory Profiler ingestion; benchmark plan only, with no comparative baseline or model/provider subject", + "measured_experiment_status": "none; zero model/provider experiment runs", + "reproducibility_status": "tools/verify, package build, provider-free dogfood, receipt validation, and artifact hash checks are reproducible from the public revision; cognitive transport and comparative claims remain untested", + "documentation_status": "canonical theory, normative draft specification, architecture, security boundary, known limitations, benchmark plan, provenance, usage, and exact release evidence are public", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["No production helper transport is bundled; transport isolation is external, Context Firewall reduction is not integrated, Python is the only symbol selector, raw deterministic byte ceilings are post-process, distributed locking and source-write application are absent, and no controlled evidence establishes correctness preservation or intelligence, context, token, latency, cost, or provider-session savings."], + "dependencies": ["context-firewall", "decision-evidence-protocol", "agent-trajectory-profiler", "agent-routing-policy", "agent-execution-authorization"], + "dependents": [], + "active_experiment_ids": [], + "blockers": ["A production-quality bounded helper transport, independently verified isolation and termination, full Context Firewall integration, and a frozen comparative benchmark remain missing."], + "next_task": "Support Durable Supervisor receipt visibility and Opsle Tasks workload measurement; advance the standalone Gearbox only if integrated evidence exposes a routing or preflight deficiency.", + "evidence": ["https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py", "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/tests/test_core.py", "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/evidence/release-001/verification.json", "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/PROVENANCE.md", "program/evidence/gearbox-publication/README.md", "https://github.com/opsle/gearbox/pull/1", "https://github.com/opsle/gearbox/actions/runs/33229696388"], + "completion_criteria": ["Publish and verify a production-quality bounded helper transport without absorbing durable orchestration.", "Benchmark deterministic and bounded-cognitive gears against real direct-execution baselines under equal correctness gates.", "Replicate bounded value claims and document transport, isolation, cleanup, escalation, and unsupported-task failure modes."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "affected-verification", + "github_url": "https://github.com/opsle/affected-verification", + "default_branch": "main", + "last_verified_head_sha": "97f490a67337552fee25757266f3dc034660dca0", + "project_type": "concept", + "purpose": "Select the smallest verification workload whose sufficiency can be defended from available change-impact, dependency, coverage, policy, and risk evidence.", + "lifecycle_stage": "VERIFIED", + "implementation_status": "dependency-free Node.js deterministic plan-v2 planner with check-level dependency-completeness states, mechanism and boundary evidence, check-local fail-closed forced selection, explainable skips, bounded deterministic Python boundary inspection, identity-bound SHADOW validation, frozen-oracle repair/replay harnesses, and opsle.value-receipt.v1 safety-cost telemetry", + "implementation_requirement": "A runnable deterministic planner, input validator, plan contract, conformance fixtures, and shadow classifier are sufficient for the prototype gate; real adapters and comparative evidence are later gates.", + "specification_status": "versioned input-v2 and opsle.affected-verification.plan.v2 contracts define check-level dependency mechanisms, completeness states, opaque-boundary provenance, forced selection, evidence coverage, policy matching, selection and skip reasons, deterministic identities, Visible Value safety additions, and SHADOW observations; plan v1 remains immutable historical evidence", + "test_status": "107 of 107 automated tests, 15 of 15 conformance scenarios, and 10 of 10 determinism checks passed locally, in PR #4 CI, in exact final-main CI, and from a fresh detached worktree at 97f490a67337552fee25757266f3dc034660dca0; the AV-EXP-003 verifier and full deterministic repair reproduction also passed with identical result and regression-matrix identities", + "benchmark_status": "AV-EXP-003 preregistered and recorded an opaque-boundary SHADOW repair: the permanently preserved AV-EXP-002 AV2-006 miss is selected in both repaired AV arms for generalized subprocess/child-interpreter completeness evidence; ten adversarial cases have zero misses, 7 added checks, 13 checks still skipped, 10/10 targeted scenarios, and zero FULL escalations; frozen AV-EXP-001 adds 0 test executions and has 0 repaired misses, while AV-EXP-002 adds 6 AV_CORE and 7 AV_WITH_SELECTOR executions and has 0 repaired misses", + "measured_experiment_status": "AV-EXP-001, AV-EXP-002, and AV-EXP-003 RECORDED; AV-EXP-002 permanently remains FAIL with one miss in each AV arm; AV-EXP-003 result sha256:03b2f7d6a380c84f6a1749531067cf8b87404c879f42380de8f07cce48251519 and regression matrix sha256:7260c2d3476a6e78323e75d36c54c8409ea4cb18fa3a8f9a76b5533e1df08615 measure repair selection and exact precision cost; zero provider/model runs", + "reproducibility_status": "the public AV-EXP-003 harness deterministically replays frozen AV-EXP-001/002 evidence, reruns ten adversarial full-oracle SHADOW cases, inspects the pinned Click source, validates ten Visible Value receipts, and reproduces identical result and regression-matrix identities from a fresh exact-main worktree on this host; no independent qualifying replication exists", + "documentation_status": "public canonical definition, normative specification, source-linked prior-art audit, architecture and independence boundaries, verification catalog and policy semantics, trust ramp, benchmark plan, limitations, usage, security, schema, and fixtures are present", + "site_publication_status": "GitHub documentation only; no evidence-backed site publication", + "known_limitations": ["AV-EXP-002 permanently remains a FAIL. AV-EXP-003 repairs the known skip in frozen replay but its bounded static inspector does not trace child-process imports, close arbitrary plugin/reflection behavior, solve dynamic Python dependencies, provide a production adapter, add historical real-change evidence, establish independent replication, prove general safety or correctness equivalence, claim causal savings, or authorize production trust."], + "dependencies": [], + "dependents": [], + "active_experiment_ids": ["AV-EXP-001", "AV-EXP-002", "AV-EXP-003"], + "blockers": ["The bounded repair does not close opaque boundaries or constitute dynamic analysis; historical real-change replay, a production-quality evidence adapter, and independent qualifying replication remain missing. AV remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized."], + "next_task": "Remain OBSERVE/SHADOW and wait for Durable Supervisor-driven Opsle Tasks work to expose a concrete verification-selection deficiency before adding another experiment.", + "evidence": ["https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/REPORT.md", "https://github.com/opsle/affected-verification/blob/7aa4d13e42d6a547973d7f2a6b330821145cedc2/benchmark/av-exp-003/preregistration-v1/preregistration.json", "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/summary.json", "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/repair-regression-matrix.json", "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/evidence-manifest.json", "https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/REPORT.md", "https://github.com/opsle/affected-verification/pull/4", "https://github.com/opsle/affected-verification/actions/runs/33412942072"], + "completion_criteria": ["Publish production evidence adapters and independently validate their completeness boundaries.", "Run correctness-first comparisons against full verification and native selectors with durable shadow miss evidence.", "Replicate bounded workload and safety claims on independent repositories or environments."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "research", + "github_url": "https://github.com/opsle/research", + "default_branch": "main", + "last_verified_head_sha": "e8a2c36678ac4d4f72a79a56e03e9f896811be02", + "project_type": "program infrastructure", + "purpose": "Own public research methodology, portfolio reconciliation, experiment records, and the authoritative program ledger.", + "lifecycle_stage": "PROTOTYPED", + "implementation_status": "authoritative 21-repository portfolio and priority ledger, machine-readable 18-concept theory registry, five-experiment registry, lifecycle and anti-nitpick controls, normative Visible Value target, deterministic generated status and priority views, provider-free EXP-001 preparation, and recorded AV-EXP-001/002/003 shadow evidence with integrity CI", + "implementation_requirement": "Program infrastructure requires a validated registry, generated dashboard, experiment ledger, operating rules, and CI.", + "specification_status": "canonical lifecycle, theory classifications and dispositions, Gearbox and Context Firewall definitions, Gearbox-versus-Durable boundary, Visible Value receipt and measurement classes, operator/model channels, observational corpus, shadow/replay attachment, NOW/NEXT/THEN/LATER/PARKED lanes, work-item admission, registry, experiment, and deterministic generated-view controls are specified", + "test_status": "90 of 90 repository tests passed locally at the verified default-branch HEAD; the 21-repository and 18-concept registries, five experiment records, generated dashboard, and deterministic experiment evidence checks passed", + "benchmark_status": "EXP-001 provider-free components remain frozen and unconsumed; AV-EXP-001/002/003 are recorded correctness-first shadow calibrations, including the preserved AV-EXP-002 miss and bounded AV-EXP-003 repair evidence; no provider/model subject ran", + "measured_experiment_status": "AV-EXP-001, AV-EXP-002, and AV-EXP-003 are recorded shadow experiments; EXP-001 remains planned and unconsumed; no provider/model experiment run exists", + "reproducibility_status": "program and receipt validation, both generated views, public dogfood artifacts, exact-revision EXP-001 provider-free qualification, and AV-EXP-001/002/003 deterministic shadow evidence are reproducible; full secret-backed coordinator replay still requires the external seed, and no provider/model subject result exists", + "documentation_status": "public canonical theory map, project reconciliation, generated priority and portfolio views, registered Gearbox boundary and prototype scope, Context Firewall adapter scope, consolidation provenance policy, program documentation, Visible Value semantics, machine controls, and evidence are present", + "site_publication_status": "GitHub research hub only", + "known_limitations": ["The registry cannot self-reference the commit that contains its own SHA, default-branch HEAD verification remains operator-driven, EXP-001 is still unconsumed, AV remains OBSERVE/SHADOW, the site is not registry-derived, and integrated Durable Supervisor/Opsle Tasks token, cost, first-pass, retry, and avoidable-intelligence measurements do not yet exist."], + "dependencies": [], + "dependents": [".github", "site"], + "active_experiment_ids": ["EXP-001", "AV-EXP-001", "AV-EXP-002", "AV-EXP-003", "LEGACY-001"], + "blockers": ["Durable Supervisor v0.1 measurement and foreign-workload stopping criteria remain open; Opsle Tasks cannot become the primary workload until v0.1 is declared and frozen."], + "next_task": "Keep the authoritative priority and portfolio views current while Durable Supervisor completes its ten bounded v0.1 stopping criteria.", + "evidence": ["program/registry.json", "PROGRAM_STATUS.md", "program/PRIORITY.md", "program/THEORY_MAP.md", "program/theory-registry.json", "program/evidence/gearbox-publication/README.md", "program/evidence/exp-001-offline-freeze/README.md", "program/evidence/exp-001-preregistration/README.md", "program/evidence/exp-001-block-coordinator/README.md", "program/evidence/exp-001-live-preflight/README.md", "https://github.com/opsle/research/pull/19", "https://github.com/opsle/research/pull/20", "https://github.com/opsle/research/pull/21", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/experiments/exp-001/benchmark.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/experiments/exp-001/preregistration-v1/preregistration.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/experiments/exp-001/coordinator-v1/coordinator.py", "https://github.com/opsle/research/blob/72f9e4a0326d68d3870e2e79ce4e351acb1d8ffa/program/VISIBLE_VALUE_CONTRACT.md"], + "completion_criteria": ["Keep the 21-repository ledger and dashboard mechanically consistent.", "Retain immutable experiment evidence and lifecycle promotion proof.", "Publish program documentation without stale or unsupported claims."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "active", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": "site", + "github_url": "https://github.com/opsle/site", + "default_branch": "main", + "last_verified_head_sha": "28ad65be4750dc849976fbf5c9eae9501c6bbb25", + "project_type": "program infrastructure", + "purpose": "Provide source for the future public research, benchmark, failure, documentation, and product-gateway site.", + "lifecycle_stage": "PROTOTYPED", + "implementation_status": "React/Vinext source implementation with content routes", + "implementation_requirement": "Program infrastructure requires tested source, registry-derived content, accessibility, and separately authorized publication proof.", + "specification_status": "README and route structure define the current source scope", + "test_status": "automated build/render tests present; not rerun because this reconciliation kept other repositories read-only", + "benchmark_status": "not applicable to the source shell; evidence-backed benchmark content is absent", + "measured_experiment_status": "none", + "reproducibility_status": "build instructions present; current test result unverified in this run", + "documentation_status": "public source README and content routes present", + "site_publication_status": "source only; explicitly not deployed", + "known_limitations": ["Content is not registry-generated, evidence-backed research results are absent, and deployment is intentionally unauthorized."], + "dependencies": ["research"], + "dependents": [], + "active_experiment_ids": [], + "blockers": ["Wait for validated registry data and measured research; deployment requires separate authorization."], + "next_task": "Remain later until measured research and separate site-release authorization justify registry-derived public content.", + "evidence": ["https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md", "https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/tests/rendered-html.test.mjs"], + "completion_criteria": ["Render registry-derived public state without unsupported claims.", "Pass build, render, accessibility, and content-consistency checks.", "Publish only under a separate authorized release with exact revision evidence."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + }, + { + "name": ".github", + "github_url": "https://github.com/opsle/.github", + "default_branch": "main", + "last_verified_head_sha": "01c38e726db7c3e45059d25fccce55e071e35938", + "project_type": "program infrastructure", + "purpose": "Publish the Opsle organization profile and entry links to the research program.", + "lifecycle_stage": "THEORY", + "implementation_status": "documentation-only organization profile", + "implementation_requirement": "Program infrastructure may remain documentation-only but must be mechanically checked against the authoritative registry before completion.", + "specification_status": "profile purpose is documented; no synchronization contract", + "test_status": "not applicable to current single Markdown profile; consistency is unverified", + "benchmark_status": "not applicable", + "measured_experiment_status": "none", + "reproducibility_status": "not applicable", + "documentation_status": "public organization profile present and lists all 16 concepts", + "site_publication_status": "published as the GitHub organization profile", + "known_limitations": ["Repository list is manually maintained and can drift from the authoritative registry."], + "dependencies": ["research"], + "dependents": [], + "active_experiment_ids": [], + "blockers": ["No mechanical registry consistency check exists in this repository."], + "next_task": "Remain parked until a broken organization-profile link or external release condition creates a concrete need.", + "evidence": ["https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md"], + "completion_criteria": ["Keep organization identity and repository links consistent with the authoritative registry.", "Automate or verify consistency at exact revisions.", "Document ownership and update workflow."], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-05T15:28:58Z" + } + ] +} diff --git a/program/history/pre-consolidation/theory-registry.json b/program/history/pre-consolidation/theory-registry.json new file mode 100644 index 0000000..dd58e53 --- /dev/null +++ b/program/history/pre-consolidation/theory-registry.json @@ -0,0 +1,744 @@ +{ + "schema_version": 1, + "registry_id": "opsle.theory-registry.v1", + "verified_at": "2026-08-31T02:59:01Z", + "source_repository_count": 21, + "current_concept_repository_count": 18, + "canonical_definitions": { + "gearbox": "Agent Gearbox lets a powerful primary developer delegate routine operations and bounded work to deterministic software or less expensive models, then receive only the compact result needed to continue.", + "context_firewall": "Context Firewall is a deterministic boundary that keeps operational noise out of an AI agent's context while preserving the compact evidence, provenance, and escalation path the agent needs to make correct decisions.", + "durable_orchestration": "Durable orchestration owns autonomous objective progress across durable state and multiple activations without requiring a continuously active primary developer.", + "affected_verification": "Affected Verification deterministically selects the smallest verification workload whose sufficiency can be defended from the available change-impact, dependency, coverage, policy, and risk evidence." + }, + "canonical_invariants": [ + "Gearbox saves intelligence; Context Firewall saves context.", + "Gearbox determines where bounded work executes; Context Firewall determines what information returns.", + "Gearbox enhances a primary developer; durable orchestration owns progress across activations without that developer remaining continuously active.", + "Research-hypothesis granularity does not imply repository granularity.", + "Repository tracking and concept tracking are independent anti-forgetting controls.", + "Affected Verification decides what verification should execute; Context Firewall decides what resulting evidence enters model context." + ], + "concepts": [ + { + "id": "gearbox", + "canonical_concept_name": "Agent Gearbox", + "one_sentence_definition": "Agent Gearbox lets a powerful primary developer delegate routine operations and bounded work to deterministic software or less expensive models, then receive only the compact result needed to continue.", + "original_problem": "A powerful primary developer spends scarce reasoning and context on deterministic operations or bounded work that cheaper software or models could perform.", + "primary_classification": "GEARBOX_CORE", + "current_repository": "gearbox", + "recommended_disposition": "KEEP_STANDALONE", + "disposition_rationale": "The released repository now owns one coherent primary-developer transmission boundary while keeping Context Firewall, evidence protocols, routing, authorization, and isolation external.", + "disposition_gain": "The public home reunifies admission, gear selection, bounded helper execution, passive waiting, compact return, budget enforcement, cleanup, and exact provenance.", + "disposition_risk": "A poorly bounded repository could absorb durable orchestration, isolation infrastructure, or every supporting policy and become a general autonomous-agent platform.", + "provenance_concerns": "The public AGPL repository records the exact Taslos Tasks source and introduction revisions; future migrations must preserve that attribution and must not silently absorb other Opsle repositories or their lifecycle evidence.", + "relationship_to_gearbox": "Canonical concept and current public implementation boundary.", + "relationship_to_context_firewall": "Independent and complementary: Gearbox chooses the execution gear; Context Firewall adapts the evidence returned from that execution.", + "dependencies": [ + "context-firewall", + "decision-evidence-protocol", + "agent-trajectory-profiler", + "agent-routing-policy", + "agent-execution-authorization" + ], + "consumers": [], + "current_implementation_fidelity": { + "status": "NARROW_PROTOTYPE", + "assessment": "The public provider-free core implements strict request and policy admission, exact deterministic execution, content-addressed staged helper context, one injected helper transport, blocking wait, literal budgets, raw-artifact locators, compact results, termination checks, and Visible Value receipts. No production provider transport, Context Firewall integration, comparative benchmark, or autonomous orchestration exists." + }, + "drift_status": "RESTORED_PUBLIC_HOME", + "current_name_accuracy": "Canonical name and repository boundary now match the restored primary-developer theory.", + "evidence_references": [ + "program/THEORY_MAP.md", + "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/THEORY.md", + "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/SPEC.md", + "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py", + "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/evidence/release-001/verification.json", + "program/evidence/gearbox-publication/README.md", + "https://github.com/opsle/gearbox/pull/1" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "Which production helper transport can independently prove filesystem, credential, network, provider, and termination bounds?", + "Which external routing and authorization profiles should receive formal Gearbox adapters without moving their ownership into this repository?", + "What comparative provider-session, correctness, and primary-context effects survive a controlled benchmark without inventing counterfactual savings?" + ] + }, + { + "id": "agent-trajectory-profiler", + "canonical_concept_name": "Agent Trajectory Profiler", + "one_sentence_definition": "Agent Trajectory Profiler reconstructs correctness-gated mutation, edit, revisit, and evidence-exposure measurements from observable execution artifacts.", + "original_problem": "Final correctness, latency, token, and cost scores hide discarded mutation, repeated editing, semantic revisits, and excess evidence exposure.", + "primary_classification": "INDEPENDENT_OPSLE_TOOL", + "current_repository": "agent-trajectory-profiler", + "recommended_disposition": "KEEP_STANDALONE", + "disposition_rationale": "The measurement mechanism is reusable across Gearbox, Context Firewall, editing research, and durable orchestration.", + "disposition_gain": "Independent telemetry and benchmark instrumentation remain installable and falsifiable without an orchestrator.", + "disposition_risk": "Generic Visible Value validation and aggregation may continue to expand beyond trajectory profiling unless protocol ownership is clarified.", + "provenance_concerns": "Retain the initial metric implementation, later Context Firewall adapter, and Visible Value aggregation commits as distinct evidence scopes.", + "relationship_to_gearbox": "External telemetry consumer for gear choice, helper execution, passive wait, provider-session, cleanup, and result measurements; it never selects or runs a gear.", + "relationship_to_context_firewall": "Consumes validated packet, suppression, and escalation observations to measure initial and final model-visible evidence; it is not a Context Firewall dependency.", + "dependencies": [ + "decision-evidence-protocol" + ], + "consumers": [ + "gearbox", + "semantic-edit-protocol" + ], + "current_implementation_fidelity": { + "status": "SCOPED_VERIFIED", + "assessment": "The executable metrics, Context Firewall adapter, run records, and value summaries are verified at the recorded revision; predictive value, causal efficiency, and generic protocol ownership are unproven." + }, + "drift_status": "SCOPE_EXPANDED", + "current_name_accuracy": "Accurate for the original metrics; generic Visible Value aggregation is a broader responsibility.", + "evidence_references": [ + "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/README.md", + "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/metrics.js", + "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/value-summary.js" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "Should generic Visible Value receipt validation remain in this repository?", + "Which metrics predict useful agent outcomes across models and tasks?", + "How should semantic-region ground truth be established?" + ] + }, + { + "id": "semantic-edit-protocol", + "canonical_concept_name": "Semantic Edit Protocol", + "one_sentence_definition": "Semantic Edit Protocol applies bounded structural edits under explicit semantic-region and hash preconditions with atomic rollback and concise receipts.", + "original_problem": "Whole-file transport and repeated text patches can inflate payloads, overwrite concurrent work, and mutate stale structure.", + "primary_classification": "INDEPENDENT_OPSLE_TOOL", + "current_repository": "semantic-edit-protocol", + "recommended_disposition": "KEEP_STANDALONE", + "disposition_rationale": "A bounded semantic editing primitive has independent use outside Gearbox and any autonomous orchestrator.", + "disposition_gain": "Language adapters and edit semantics can be benchmarked against patch and whole-file baselines as one focused tool.", + "disposition_risk": "Repository overhead may not be justified if no portable operation schema or adapter emerges.", + "provenance_concerns": "Preserve the extraction commit and negative benchmark evidence if the language-independent boundary fails.", + "relationship_to_gearbox": "A possible deterministic execution tool selected by Gearbox, not a Gearbox core module.", + "relationship_to_context_firewall": "Independent; a Context Firewall adapter may compact edit diagnostics or receipts without defining edit semantics.", + "dependencies": [ + "agent-resource-claims", + "agent-trajectory-profiler" + ], + "consumers": [ + "gearbox" + ], + "current_implementation_fidelity": { + "status": "THEORY_ONLY", + "assessment": "The theory matches the original problem, but the generic specification does not define a concrete operation and no adapter, validator, fixture, or automated test exists." + }, + "drift_status": "ALIGNED_THEORY_ONLY", + "current_name_accuracy": "Accurate.", + "evidence_references": [ + "https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/README.md", + "https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/SPEC.md" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "What is the first concrete operation and language adapter?", + "Can anchors, transactions, and rollback be made portable?", + "Does the protocol outperform ordinary patching under equal correctness gates?" + ] + }, + { + "id": "durable-supervisor", + "canonical_concept_name": "Durable Supervisor", + "one_sentence_definition": "Durable Supervisor owns autonomous objective progress, reconstructs durable state, and resumes decision-making across activations while a non-reasoning runner handles lifecycle mechanics.", + "original_problem": "Long-horizon autonomous work loses state across interruptions and wastes cognition on waiting, lifecycle mechanics, and reconstruction.", + "primary_classification": "DURABLE_ORCHESTRATION", + "current_repository": "durable-supervisor", + "recommended_disposition": "KEEP_AS_RESEARCH", + "disposition_rationale": "The autonomous cross-activation hypothesis is distinct from Gearbox, but the installable boundary and family packaging are not yet proven.", + "disposition_gain": "The separate research home prevents durable objective ownership and reconstruction from being conflated with bounded primary-developer delegation.", + "disposition_risk": "Remaining an umbrella research repository may defer decisions about ledger, scheduler, and wakeup module ownership.", + "provenance_concerns": "Retain the extraction commit and explicitly attribute any future consolidation of ledger, scheduler, wakeup, discovery, or recovery research.", + "relationship_to_gearbox": "Formally separate: it may operate across activations without a continuously active primary developer; Gearbox remains subordinate to one primary developer's bounded call.", + "relationship_to_context_firewall": "May consume compact evidence packets from child or runner activity, but Context Firewall provides no persistence, scheduling, or objective ownership.", + "dependencies": [ + "agent-state-ledger", + "agent-scheduler-runtime", + "event-driven-agent-wakeup", + "decision-evidence-protocol" + ], + "consumers": [], + "current_implementation_fidelity": { + "status": "THEORY_ONLY", + "assessment": "The theory is coherent, but no ledger, runner, reconstruction envelope, fixture, or automated test exists. Shared fresh-child and passive-wait language caused external conflation with Gearbox." + }, + "drift_status": "BOUNDARY_CONFLATION", + "current_name_accuracy": "Accurate when durable objective ownership and cross-activation reconstruction remain explicit.", + "evidence_references": [ + "https://github.com/opsle/durable-supervisor/blob/555ebedb992ac74236bb7da8230b4d6b0489830b/README.md", + "https://github.com/opsle/durable-supervisor/blob/555ebedb992ac74236bb7da8230b4d6b0489830b/ARCHITECTURE.md" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "What is the minimum durable reconstruction envelope?", + "Which family mechanisms should share one implementation repository?", + "Can correctness match a continuously active supervisor under restart and duplicate-event faults?" + ] + }, + { + "id": "event-driven-agent-wakeup", + "canonical_concept_name": "Event-Driven Agent Wakeup", + "one_sentence_definition": "Event-Driven Agent Wakeup durably registers a wait, suspends model activity, and reactivates an orchestrator exactly once on a decision-relevant event.", + "original_problem": "Polling and waiting were implemented as repeated model turns even though they are runtime concerns.", + "primary_classification": "DURABLE_ORCHESTRATION", + "current_repository": "event-driven-agent-wakeup", + "recommended_disposition": "CONSOLIDATE_WITH_OTHER", + "disposition_rationale": "Durable wait registration, event delivery, scheduling, and reconstruction share one state boundary with the Durable Supervisor family.", + "disposition_gain": "One durable event and scheduling contract avoids duplicate persistence and wakeup semantics.", + "disposition_risk": "Independent polling-versus-wakeup benchmarks and reusable adapters could become less visible.", + "provenance_concerns": "Preserve the repository commit, two prototype tests, lifecycle record, citations, and a future redirect before consolidation.", + "relationship_to_gearbox": "Not Gearbox's passive wait: Gearbox blocks synchronously at the OS or transport layer inside one bounded call; this concept persists a wait and reactivates an orchestrator.", + "relationship_to_context_firewall": "May use adapters for event/process chatter, but wake identity, timeout, failure, approval, and loss signals are unsafe to hide.", + "dependencies": [ + "agent-state-ledger", + "agent-scheduler-runtime" + ], + "consumers": [ + "durable-supervisor" + ], + "current_implementation_fidelity": { + "status": "NARROW_PROTOTYPE", + "assessment": "The eight-line in-memory transition function exercises duplicate and wrong-wait behavior only; it does not implement durable registration, restart recovery, event delivery, or timeouts." + }, + "drift_status": "PROTOTYPE_SCOPE_OVERSTATED", + "current_name_accuracy": "Accurate for the theory; the current implementation is not durable.", + "evidence_references": [ + "https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/README.md", + "https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/src/runtime.js", + "https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/tests/runtime.test.js" + ], + "confidence": "MEDIUM_HIGH", + "unresolved_questions": [ + "Who owns timeout production and late-event retention?", + "What proves exactly-once decision activation across restart?", + "Do reusable wait adapters justify independent packaging?" + ] + }, + { + "id": "context-firewall", + "canonical_concept_name": "Context Firewall", + "one_sentence_definition": "Context Firewall is a deterministic boundary that keeps operational noise out of an AI agent's context while preserving the compact evidence, provenance, and escalation path the agent needs to make correct decisions.", + "original_problem": "Tool and process output became model-visible merely because it was emitted, consuming context with operational noise.", + "primary_classification": "INDEPENDENT_OPSLE_TOOL", + "current_repository": "context-firewall", + "recommended_disposition": "KEEP_STANDALONE", + "disposition_rationale": "Deterministic evidence adaptation is independently useful across agents, Gearbox, CI, operations, and orchestrators.", + "disposition_gain": "Adapter policies, packet conformance, and raw-evidence escalation can evolve without being coupled to an executor.", + "disposition_risk": "The mature TAP-subset adapter may be mistaken for the complete multi-adapter concept or its lifecycle stage may be overgeneralized.", + "provenance_concerns": "Preserve TAP reducer, Visible Value, conformance, and EXP-001 evidence as adapter-specific evidence rather than rewriting them as proof of the whole concept.", + "relationship_to_gearbox": "Independent complement: it adapts deterministic-tool or helper evidence after Gearbox chooses and executes a gear.", + "relationship_to_context_firewall": "Canonical home.", + "dependencies": [], + "consumers": [ + "gearbox", + "decision-evidence-protocol", + "agent-trajectory-profiler", + "durable-supervisor" + ], + "intended_adapter_families": [ + "tests", + "lint", + "typecheck", + "git", + "build_compiler", + "process_service", + "helper_agent_result" + ], + "current_implementation_fidelity": { + "status": "ADAPTER_PROTOTYPED", + "assessment": "The strict TAP-compatible reducer faithfully implements one deterministic adapter and packet profile. Other adapter families and the safe correctness frontier do not exist." + }, + "drift_status": "PARTIAL_ADAPTER_SCOPE", + "current_name_accuracy": "Accurate for the concept; implementation maturity must remain adapter-scoped.", + "evidence_references": [ + "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/README.md", + "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/SPEC.md", + "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/src/reducer.js" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "What common adapter interface preserves tool-specific safety?", + "How are raw artifact availability and hash validity independently proven?", + "Where is the experimentally safe reduction frontier by task and adapter family?" + ] + }, + { + "id": "decision-evidence-protocol", + "canonical_concept_name": "Decision Evidence Protocol", + "one_sentence_definition": "Decision Evidence Protocol defines small versioned vendor-neutral envelopes for decision facts, verification, uncertainty, provenance, artifacts, and explicit raw-evidence escalation.", + "original_problem": "Agents repeatedly parsed provider-specific prose and tool chatter to recover the same small set of decision facts.", + "primary_classification": "CROSS_CUTTING_PROTOCOL", + "current_repository": "decision-evidence-protocol", + "recommended_disposition": "KEEP_AS_PROTOCOL", + "disposition_rationale": "Evidence representation and conformance are reusable across Gearbox, Context Firewall, handoff, and orchestration.", + "disposition_gain": "Independent protocol validation prevents producers from self-certifying their own compact evidence.", + "disposition_risk": "The Context Firewall-specific validator and duplicated Visible Value semantics may dominate or distort the generic protocol.", + "provenance_concerns": "Keep generic-envelope, Context Firewall profile, and Visible Value validation evidence scoped to their exact commits and conformance vectors.", + "relationship_to_gearbox": "External result and evidence protocol; Gearbox owns admission and helper lifecycle, while this protocol defines and validates the compact facts returned.", + "relationship_to_context_firewall": "Context Firewall produces reduced packets; Decision Evidence independently validates their structure, hashes, source-derived claims, and sufficiency without becoming a reducer.", + "dependencies": [], + "consumers": [ + "gearbox", + "context-firewall", + "agent-trajectory-profiler", + "verifiable-agent-handoff", + "durable-supervisor", + "agent-recovery-policy" + ], + "current_implementation_fidelity": { + "status": "PROFILE_VERIFIED", + "assessment": "The Context Firewall packet/value profiles and generic value receipt validator are verified; the broader multi-tool decision envelope remains a minimal prototype." + }, + "drift_status": "PROFILE_DOMINATES_CORE", + "current_name_accuracy": "Accurate, with maturity scoped to implemented profiles.", + "evidence_references": [ + "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/README.md", + "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/context-firewall-v1.js", + "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/validate.js" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "What is the canonical generic envelope independent of one producer profile?", + "Should Decision Evidence be the sole normative Visible Value validator?", + "Which tool-class profiles demonstrate adequate decision coverage?" + ] + }, + { + "id": "agent-state-ledger", + "canonical_concept_name": "Agent State Ledger", + "one_sentence_definition": "Agent State Ledger records immutable orchestration facts and derives a deterministic stale-aware projection of current work, evidence, authority, contradictions, and uncertainty.", + "original_problem": "Conversation history could not reliably reconstruct what autonomous work was done, verified, failed, authorized, uncertain, or next after interruption.", + "primary_classification": "DURABLE_ORCHESTRATION", + "current_repository": "agent-state-ledger", + "recommended_disposition": "CONSOLIDATE_WITH_OTHER", + "disposition_rationale": "The ledger is the authoritative history and projection module of the Durable Supervisor family, not a second independent state owner.", + "disposition_gain": "One durable state contract prevents supervisor, scheduler, discovery, and recovery from maintaining competing histories.", + "disposition_risk": "Potential reuse as a generic audit/projection library may be obscured.", + "provenance_concerns": "Preserve the current repository, extraction commit, benchmark theory, and future negative evidence through an attributable module and redirect.", + "relationship_to_gearbox": "Normally unnecessary for one bounded Gearbox call; absorbing it would pull Gearbox toward persistent supervision.", + "relationship_to_context_firewall": "May record packet hashes, raw locators, and escalation facts, but compact packets are not a replacement for authoritative history.", + "dependencies": [ + "decision-evidence-protocol" + ], + "consumers": [ + "durable-supervisor", + "event-driven-agent-wakeup", + "agent-scheduler-runtime", + "agent-discovery-control", + "agent-recovery-policy" + ], + "current_implementation_fidelity": { + "status": "THEORY_ONLY", + "assessment": "The theory is coherent, but the generic specification defines no event schema, projection contract, contradiction rules, fixture, implementation, or test." + }, + "drift_status": "OVER_EXTRACTED_REPOSITORY_BOUNDARY", + "current_name_accuracy": "Accurate, though the mechanism is not intrinsically agent-specific.", + "evidence_references": [ + "https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/README.md", + "https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/SPEC.md" + ], + "confidence": "MEDIUM_HIGH", + "unresolved_questions": [ + "Do non-orchestration consumers justify independent packaging?", + "What event and projection schemas preserve contradictions and uncertainty?", + "What compaction is safe without losing reconstructability?" + ] + }, + { + "id": "agent-scheduler-runtime", + "canonical_concept_name": "Agent Scheduler Runtime", + "one_sentence_definition": "Agent Scheduler Runtime performs deterministic readiness, queue, time, cancellation, pause, and bounded-concurrency transitions over durable execution identities.", + "original_problem": "Queues, dependency readiness, retry timing, timeouts, and schedules were mechanical decisions left inside model loops.", + "primary_classification": "DURABLE_ORCHESTRATION", + "current_repository": "agent-scheduler-runtime", + "recommended_disposition": "CONSOLIDATE_WITH_OTHER", + "disposition_rationale": "Scheduling is the deterministic runtime module of Durable Supervisor and shares persistence and event boundaries with the ledger and wakeup mechanism.", + "disposition_gain": "One orchestration family can own readiness, time, pause, durable event, and reconstruction semantics without duplicating state.", + "disposition_risk": "Independent scheduler reuse may be lost and the Durable Supervisor family could become too monolithic.", + "provenance_concerns": "Retain the existing theory and benchmark plan as a named module with attributable history if folded.", + "relationship_to_gearbox": "Not Gearbox core: a bounded Gearbox call does not require queues, schedules, Global Pause, autonomous retries, or restart recovery.", + "relationship_to_context_firewall": "May expose scheduler and process evidence through adapters, but Context Firewall cannot decide readiness or mutate schedules.", + "dependencies": [ + "agent-state-ledger", + "agent-resource-claims" + ], + "consumers": [ + "durable-supervisor", + "event-driven-agent-wakeup", + "controlled-agent-acceptance" + ], + "current_implementation_fidelity": { + "status": "THEORY_ONLY", + "assessment": "No queue state machine, store, fake clock, pause contract, fixture, implementation, or automated test exists; the initial theory also overlaps claims, recovery, and routing." + }, + "drift_status": "OVER_BROAD_INITIAL_SCOPE", + "current_name_accuracy": "Accurate if claims, routing, and recovery remain external.", + "evidence_references": [ + "https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/README.md", + "https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/THEORY.md" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "What exact transitions belong to scheduler versus claims and recovery?", + "Who owns timeout and provider-availability events?", + "Can the scheduler remain a reusable module inside a consolidated repository?" + ] + }, + { + "id": "verifiable-agent-handoff", + "canonical_concept_name": "Verifiable Agent Handoff", + "one_sentence_definition": "Verifiable Agent Handoff seals exact source, result, artifact, and authority identities before source cleanup so a fresh verifier can reconstruct and validate without fallback.", + "original_problem": "A verifier could not establish an executor's exact result when evidence remained only in the executor's mutable or disposable environment.", + "primary_classification": "CROSS_CUTTING_PROTOCOL", + "current_repository": "verifiable-agent-handoff", + "recommended_disposition": "KEEP_AS_PROTOCOL", + "disposition_rationale": "The durable result-transfer and independent-verification boundary is reusable across Gearbox, durable orchestration, CI, and isolation systems.", + "disposition_gain": "Artifact identity, publication-before-destruction, and reconstruction semantics remain independently conformable.", + "disposition_risk": "Another evidence envelope may duplicate Decision Evidence unless layering is explicit.", + "provenance_concerns": "Preserve the result-loss failure narrative and scope the current HMAC prototype to manifest authentication rather than claiming full handoff.", + "relationship_to_gearbox": "Optional external protocol when a bounded helper's source environment will be destroyed; not gear selection, admission, or lifecycle ownership.", + "relationship_to_context_firewall": "Handoff preserves exact durable artifacts; Context Firewall decides which verified facts enter model context and must escalate rather than hide required handoff evidence.", + "dependencies": [ + "decision-evidence-protocol" + ], + "consumers": [ + "gearbox", + "ephemeral-agent-workers", + "controlled-agent-acceptance" + ], + "current_implementation_fidelity": { + "status": "SUBCOMPONENT_PROTOTYPED", + "assessment": "The HMAC manifest prototype validates required bindings and a caller-supplied destruction assertion; it does not publish, transport, prove destruction, reconstruct, or independently verify artifacts." + }, + "drift_status": "PROTOTYPE_SCOPE_OVERSTATED", + "current_name_accuracy": "Accurate for the theory; current code is a seal subcomponent.", + "evidence_references": [ + "https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/README.md", + "https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/src/seal.js" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "How does the handoff profile layer over Decision Evidence?", + "What independently proves source destruction and publication durability?", + "Which non-Git artifact formats are supported?" + ] + }, + { + "id": "agent-routing-policy", + "canonical_concept_name": "Agent Routing Policy", + "one_sentence_definition": "Agent Routing Policy filters and selects a purpose-bound model, provider, profile, or executor route under exact capability, access, availability, independence, and budget constraints with a durable reason.", + "original_problem": "Opaque routing violated provider budgets, access boundaries, exact model or profile constraints, and reviewer independence.", + "primary_classification": "GEARBOX_SUPPORTING_POLICY", + "current_repository": "agent-routing-policy", + "recommended_disposition": "FUTURE_GEARBOX_POLICY", + "disposition_rationale": "Gearbox needs bounded model, effort, and provider route selection after deterministic-versus-cognitive admission, while broader review and recovery routing must remain explicit extensions.", + "disposition_gain": "A versioned Gearbox-facing selection policy would replace ad hoc model and cost choice.", + "disposition_risk": "Absorption could erase reviewer-independence or autonomous routing research and accidentally authorize fallback retries.", + "provenance_concerns": "Migrate only the Gearbox-facing route profile; retain the original general routing theory, commit, and later experiments with attribution.", + "relationship_to_gearbox": "Supporting policy: Gearbox first admits deterministic versus cognitive work; routing then selects one allowed cognitive route and never autonomously retries or falls back.", + "relationship_to_context_firewall": "May consume compact availability or performance evidence, but incomplete or uncertain packets must make selection fail closed.", + "dependencies": [ + "decision-evidence-protocol" + ], + "consumers": [ + "gearbox", + "agent-recovery-policy", + "controlled-agent-acceptance" + ], + "current_implementation_fidelity": { + "status": "THEORY_ONLY", + "assessment": "The initial theory mixes gear/model selection, provider routing, authorization inputs, reviewer independence, and fallback eligibility; no evaluator or normative precedence exists." + }, + "drift_status": "MIXED_DECISION_OWNERSHIP", + "current_name_accuracy": "Accurate for broad routing; not a substitute name for Gearbox.", + "evidence_references": [ + "https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/README.md", + "https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/THEORY.md" + ], + "confidence": "MEDIUM_HIGH", + "unresolved_questions": [ + "Where exactly does gear selection stop and route selection begin?", + "Who owns fallback eligibility and reviewer routing?", + "What deterministic precedence applies to price, quality, access, and availability?" + ] + }, + { + "id": "agent-resource-claims", + "canonical_concept_name": "Agent Resource Claims", + "one_sentence_definition": "Agent Resource Claims establishes current concurrent ownership through canonical resource identities, atomic shared or exclusive claim sets, expiring leases, and fencing tokens.", + "original_problem": "A process could remain alive after losing authority, while independent locks could deadlock or admit partially conflicting work.", + "primary_classification": "GEARBOX_SUPPORTING_POLICY", + "current_repository": "agent-resource-claims", + "recommended_disposition": "KEEP_AS_RESEARCH", + "disposition_rationale": "The policy can support Gearbox writes, durable orchestration, semantic editing, and isolation, but a portable state machine and independent packaging value are not yet proven.", + "disposition_gain": "Research remains broadly reusable and avoids prematurely coupling lease and fence semantics to Gearbox.", + "disposition_risk": "Repository granularity and overlap with scheduler and authorization remain unresolved.", + "provenance_concerns": "Preserve the fenced-claim failure narrative and original extraction before any later module migration.", + "relationship_to_gearbox": "Optional policy for bounded shared-resource or write authority; unnecessary when a Gearbox task touches no shared mutable resource.", + "relationship_to_context_firewall": "Claim observations may be compacted, but fence mismatch, expiry, and takeover evidence are unsafe to hide and reduction never grants authority.", + "dependencies": [], + "consumers": [ + "semantic-edit-protocol", + "agent-scheduler-runtime", + "agent-execution-authorization", + "ephemeral-agent-workers" + ], + "current_implementation_fidelity": { + "status": "THEORY_ONLY", + "assessment": "The useful invariants are present, but no catalog schema, conflict lattice, lease state machine, store adapter, concurrency fixture, or automated test exists." + }, + "drift_status": "BOUNDARY_OVERLAP", + "current_name_accuracy": "Substantially accurate; the mechanism is generic concurrency authority rather than intrinsically agent-specific.", + "evidence_references": [ + "https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/README.md", + "https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/THEORY.md" + ], + "confidence": "MEDIUM", + "unresolved_questions": [ + "Does a portable resource identity and conflict lattice justify a standalone package?", + "Which time, fairness, and takeover semantics are required?", + "Where does resource authority stop and execution authorization begin?" + ] + }, + { + "id": "agent-discovery-control", + "canonical_concept_name": "Agent Discovery Control", + "one_sentence_definition": "Agent Discovery Control admits, merges, ignores, or escalates provenance-bearing discovered-work proposals under similarity, depth, existing-work, and storm-budget constraints.", + "original_problem": "Autonomous discovery created duplicate proposals, recursive work trees, and churn for already-satisfied work.", + "primary_classification": "DURABLE_ORCHESTRATION", + "current_repository": "agent-discovery-control", + "recommended_disposition": "CONSOLIDATE_WITH_OTHER", + "disposition_rationale": "Autonomous discovered-work admission depends on the Durable Supervisor family's objective graph, ledger, and shared budgets.", + "disposition_gain": "One durable proposal and convergence model avoids a standalone shell with duplicate state.", + "disposition_risk": "Independent admission-policy research and non-agent issue-triage reuse could become less visible.", + "provenance_concerns": "Preserve the repository commit, benchmark plan, extraction citation, and future fixtures through an attributable module and redirect.", + "relationship_to_gearbox": "Outside Gearbox: canonical Gearbox starts from an explicit primary-developer task and must not create recursive autonomous work.", + "relationship_to_context_firewall": "May receive compact proposal evidence, but Context Firewall cannot decide semantic duplication or replace proposal provenance.", + "dependencies": [ + "agent-state-ledger", + "decision-evidence-protocol" + ], + "consumers": [ + "durable-supervisor" + ], + "current_implementation_fidelity": { + "status": "THEORY_ONLY", + "assessment": "The theory is coherent, but the generic spec does not define proposal dispositions, similarity evidence, budgets, convergence, or already-satisfied proofs; no implementation exists." + }, + "drift_status": "OVER_EXTRACTED_REPOSITORY_BOUNDARY", + "current_name_accuracy": "Accurate for an orchestration admission policy, not a standalone agent product.", + "evidence_references": [ + "https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/README.md", + "https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/THEORY.md" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "What semantic duplicate oracle is safe?", + "How do storm budgets scale with objective complexity?", + "Does a reusable admission library exist independently of a supervisor?" + ] + }, + { + "id": "agent-execution-authorization", + "canonical_concept_name": "Agent Execution Authorization", + "one_sentence_definition": "Agent Execution Authorization validates capability-style grants that bind immutable source authority and current execution authority to one exact task, target, purpose, resource, route, and expiry.", + "original_problem": "Machine actors received broad or caller-asserted authority instead of server-derived grants bound to exact work and current state.", + "primary_classification": "GEARBOX_SUPPORTING_POLICY", + "current_repository": "agent-execution-authorization", + "recommended_disposition": "FUTURE_GEARBOX_POLICY", + "disposition_rationale": "Gearbox must fail closed before deterministic or cognitive execution, while the grant contract should remain reusable by durable orchestration and isolation infrastructure.", + "disposition_gain": "Gearbox obtains an exact authority envelope instead of ad hoc caller assertions.", + "disposition_risk": "Narrow absorption could discard cross-orchestrator composition or duplicate claim, routing, and handoff state.", + "provenance_concerns": "Preserve a separately versioned grant schema and conformance history, plus the immutable-source/current-continuation failure narrative.", + "relationship_to_gearbox": "Supporting policy that supplies and validates authority; Gearbox owns applying the result to its requested gear, context, model, budget, and task.", + "relationship_to_context_firewall": "Independent: reduced or redacted evidence is never authorization; a firewall may compact an authorization receipt only after authoritative validation.", + "dependencies": [ + "agent-resource-claims", + "decision-evidence-protocol", + "verifiable-agent-handoff" + ], + "consumers": [ + "gearbox", + "ephemeral-agent-workers", + "controlled-agent-acceptance" + ], + "current_implementation_fidelity": { + "status": "THEORY_ONLY", + "assessment": "The source/current authority distinction is preserved, but no portable grant, issuance, revocation, replay, composition, validator, or conformance suite exists." + }, + "drift_status": "BOUNDARY_OVERLAP", + "current_name_accuracy": "Accurate.", + "evidence_references": [ + "https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/README.md", + "https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/THEORY.md" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "What is the portable grant schema and revocation oracle?", + "Are claims and routes embedded or content-addressed references?", + "How are source and continuation authority composed without replay?" + ] + }, + { + "id": "controlled-agent-acceptance", + "canonical_concept_name": "Controlled Agent Acceptance", + "one_sentence_definition": "Controlled Agent Acceptance is an experiment-control method for one-shot real-agent runs under an immutable manifest, literal budgets, fail-closed activation, restoration, evidence retention, and residue reconciliation.", + "original_problem": "Real-agent acceptance could exceed provider budgets, target the wrong revision, conflict with active work, leave reusable authority, fail to restore pause state, or leave residue.", + "primary_classification": "RESEARCH_ONLY_HYPOTHESIS", + "current_repository": "controlled-agent-acceptance", + "recommended_disposition": "KEEP_AS_RESEARCH", + "disposition_rationale": "This is a provider-independent acceptance methodology for testing autonomous systems, not a production runtime primitive or Gearbox component.", + "disposition_gain": "Harness and system-under-test verdicts, first-failure behavior, budgets, and restoration remain independently testable.", + "disposition_risk": "A separate repository may be excessive if the eventual controller belongs beside each system under test.", + "provenance_concerns": "Retain predecessor run citations and future immutable manifests without promoting private operational evidence into generalized proof.", + "relationship_to_gearbox": "May later test one-shot Gearbox budget and cleanup behavior, but does not define or implement Gearbox.", + "relationship_to_context_firewall": "Can use adapters for captured logs; preflight defects, first failures, budget use, restoration failures, residue, and dual verdicts are unsafe to hide.", + "dependencies": [ + "agent-execution-authorization", + "agent-routing-policy", + "agent-scheduler-runtime", + "verifiable-agent-handoff" + ], + "consumers": [], + "current_implementation_fidelity": { + "status": "THEORY_ONLY", + "assessment": "The theory is coherent, but no manifest schema, state machine, CAS semantics, restoration protocol, residue model, controller, fixture, or test exists." + }, + "drift_status": "REPOSITORY_BOUNDARY_OVERSTATED", + "current_name_accuracy": "Accurate for a methodology or harness.", + "evidence_references": [ + "https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/README.md", + "https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/ARCHITECTURE.md" + ], + "confidence": "MEDIUM_HIGH", + "unresolved_questions": [ + "Which manifest and interruption semantics generalize beyond the predecessor?", + "Where should the eventual harness implementation live?", + "How are non-reusable grants and residue proved across runtimes?" + ] + }, + { + "id": "agent-recovery-policy", + "canonical_concept_name": "Agent Recovery Policy", + "one_sentence_definition": "Agent Recovery Policy permits bounded recovery only when typed failure evidence shows that another authorized attempt can add information or change conditions, otherwise converging to a terminal or human-input state.", + "original_problem": "Retries, consultations, fallbacks, and verification failures bypassed shared budgets or repeated the same failure indefinitely.", + "primary_classification": "DURABLE_ORCHESTRATION", + "current_repository": "agent-recovery-policy", + "recommended_disposition": "CONSOLIDATE_WITH_OTHER", + "disposition_rationale": "Recovery permission depends on the Durable Supervisor family's attempt ledger, global budgets, scheduling, and terminal-state model.", + "disposition_gain": "One authoritative attempt and budget model prevents recovery paths from bypassing orchestration limits.", + "disposition_risk": "Independent deterministic failure-policy testing and reuse could become less visible.", + "provenance_concerns": "Retain extraction history, failure narratives, benchmark plan, and future policy fixtures as an attributable module.", + "relationship_to_gearbox": "Outside canonical Gearbox: a Gearbox run returns one failed or indeterminate result and does not autonomously retry, consult, route, or fall back.", + "relationship_to_context_firewall": "May consume typed failure packets, but incomplete, truncated, hash-invalid, or unavailable raw evidence must terminate or escalate rather than authorize recovery.", + "dependencies": [ + "agent-state-ledger", + "agent-routing-policy", + "decision-evidence-protocol" + ], + "consumers": [ + "durable-supervisor" + ], + "current_implementation_fidelity": { + "status": "THEORY_ONLY", + "assessment": "No failure taxonomy, attempt ledger, information-gain rule, budget transition, evaluator, fixture, or automated test exists; alternate-route ownership overlaps routing." + }, + "drift_status": "BOUNDARY_OVERLAP", + "current_name_accuracy": "Accurate, with routing and scheduling boundaries made explicit.", + "evidence_references": [ + "https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/README.md", + "https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "What portable failure taxonomy and information-gain rule are adequate?", + "Does recovery choose an alternate-route class while routing chooses the exact route?", + "Which retry timing decisions belong to the scheduler?" + ] + }, + { + "id": "ephemeral-agent-workers", + "canonical_concept_name": "Ephemeral Agent Workers", + "one_sentence_definition": "Ephemeral Agent Workers provides restricted-broker execution in disposable bounded environments with scoped resources, sealed result capture, termination, destruction, and independent destruction verification.", + "original_problem": "Autonomous execution exposed host mounts, production secrets, broad network access, lingering processes, and ambiguous cleanup.", + "primary_classification": "EXECUTION_ISOLATION_INFRASTRUCTURE", + "current_repository": "ephemeral-agent-workers", + "recommended_disposition": "KEEP_STANDALONE", + "disposition_rationale": "Isolation, containment, worker lifecycle, and destruction proof have independent use for Gearbox, durable orchestration, CI, and other untrusted workloads.", + "disposition_gain": "A separate threat model, broker contract, portability layer, and containment harness can mature independently.", + "disposition_risk": "The implementation may duplicate authorization, claims, or handoff schemas unless those remain explicit external contracts.", + "provenance_concerns": "Preserve the Incus-derived predecessor evidence while keeping Incus an optional adapter and not copying private infrastructure.", + "relationship_to_gearbox": "Optional external executor for bounded cognitive helpers; deterministic local gears need not use it, and it never selects a gear.", + "relationship_to_context_firewall": "May expose helper logs and results through adapters, but Context Firewall proves neither containment, result sealing, nor destruction.", + "dependencies": [ + "agent-execution-authorization", + "agent-resource-claims", + "verifiable-agent-handoff" + ], + "consumers": [ + "gearbox", + "durable-supervisor" + ], + "current_implementation_fidelity": { + "status": "THEORY_ONLY", + "assessment": "No portable broker or worker adapter, resource profile, threat model, containment harness, destruction receipt, fixture, or test exists." + }, + "drift_status": "ALIGNED_THEORY_ONLY", + "current_name_accuracy": "Reasonably accurate, though the isolation primitive is not inherently agent-specific.", + "evidence_references": [ + "https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/README.md", + "https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/ARCHITECTURE.md" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "What portable isolation and destruction-proof contracts work across technologies?", + "What kernel and broker threat model is in scope?", + "How much result sealing belongs solely to Verifiable Handoff?" + ] + }, + { + "id": "affected-verification", + "canonical_concept_name": "Affected Verification", + "one_sentence_definition": "Affected Verification deterministically selects the smallest verification workload whose sufficiency can be defended from the available change-impact, dependency, coverage, policy, and risk evidence.", + "original_problem": "Existing affected-test and affected-target selectors do not by themselves establish which heterogeneous verification evidence is sufficient to accept a change, why every omitted check is irrelevant, or when uncertainty and policy must broaden the workload.", + "primary_classification": "INDEPENDENT_OPSLE_TOOL", + "current_repository": "affected-verification", + "recommended_disposition": "KEEP_STANDALONE", + "disposition_rationale": "Verification sufficiency, catalogs, risk policy, fail-closed escalation, and skip arguments are independently reusable by humans, agents, CI, Gearbox, and other developer tooling without belonging to an execution engine.", + "disposition_gain": "A standalone evidence-composition boundary can reuse native affected selectors while remaining verification-class-neutral and independently benchmarkable.", + "disposition_risk": "The project could duplicate mature test-impact systems, overclaim global minimality or safety, or drift into a CI scheduler unless adapters, claim ceilings, and execution boundaries remain explicit.", + "provenance_concerns": "The initial public prototype was created directly in opsle/affected-verification; preserve PR #1 at revision 12076522c9b82501794d816f1fcc0b7775fad6e1, the AV-EXP-001 preregistration at 0544362d7659093b7f0b4f89ee8f68023fd269c3, PR #2 and final revision 641aee9d29a89e2a8819f00817ccee8e5d234dcb, source-linked prior art, and the one-repository synthetic-corpus claim boundary without retroactively claiming novelty or general safety.", + "relationship_to_gearbox": "Independent provider: Gearbox may request a minimum defensible verification plan and choose execution gears, but it does not own the verification theory, catalog, policy, or implementation.", + "relationship_to_context_firewall": "Complementary and sequential: Affected Verification decides what checks should execute; Context Firewall decides what check results should enter model context after execution.", + "dependencies": [], + "consumers": [ + "gearbox", + "decision-evidence-protocol", + "agent-trajectory-profiler" + ], + "current_implementation_fidelity": { + "status": "VERIFIED_NARROW_PROTOTYPE", + "assessment": "The public dependency-free Node.js core validates normalized evidence/catalog/policy input, computes reverse impact, selects and explains checks, fails closed on uncertainty, emits canonical plans and Visible Value receipts, and validates identity-bound SHADOW results. AV-EXP-001 adds one benchmark-only Zustand/Vitest adapter and a frozen full-catalog oracle: both AV arms selected all 8 relevant checks observed across ten scenarios, while the uncertainty case forced full verification. It has no production adapter, independent qualifying replication, or trusted selective-verification class." + }, + "drift_status": "NEW_ALIGNED_PUBLIC_HOME", + "current_name_accuracy": "Accurate for the implemented verification-planning boundary and explicitly broader than affected-test selection.", + "evidence_references": [ + "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/SPEC.md", + "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/PRIOR_ART.md", + "https://github.com/opsle/affected-verification/blob/0544362d7659093b7f0b4f89ee8f68023fd269c3/benchmark/av-exp-001/preregistration-v1/preregistration.json", + "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/REPORT.md", + "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/summary.json", + "https://github.com/opsle/affected-verification/pull/2" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "Does the zero-observed-miss result persist in a second public repository and different ecosystem?", + "Which benchmark-only evidence adapter is valuable enough to harden into a production-quality adapter?", + "Which evidence-based promotion criteria justify TRUSTED_BOUNDED authority for a narrowly defined change class?" + ] + } + ] +} diff --git a/program/registry.json b/program/registry.json index cfdc129..aeaed02 100644 --- a/program/registry.json +++ b/program/registry.json @@ -1,22 +1,17 @@ { - "schema_version": 1, + "schema_version": 2, "program": "Opsle Research", - "authoritative_repository_count": 21, + "authoritative_repository_count": 23, "lifecycle_model": "program/LIFECYCLE.md", "experiment_registry": "program/experiments.json", "theory_registry": "program/theory-registry.json", "theory_map": "program/THEORY_MAP.md", "theory_reconciliation": { - "status": "IMPLEMENTED_HOME_REGISTERED", - "verified_at": "2026-08-29T02:53:08Z", + "status": "POST_CONSOLIDATION_RECONCILED", + "verified_at": "2026-09-09T00:00:00Z", "canonical_gearbox_definition": "Agent Gearbox lets a powerful primary developer delegate routine operations and bounded work to deterministic software or less expensive models, then receive only the compact result needed to continue.", "canonical_context_firewall_definition": "Context Firewall is a deterministic boundary that keeps operational noise out of an AI agent's context while preserving the compact evidence, provenance, and escalation path the agent needs to make correct decisions.", - "gearbox_vs_durable_supervisor": "Gearbox enhances a primary developer; durable orchestration owns autonomous objective progress across durable state and multiple activations without requiring that developer to remain continuously active.", - "gearbox_repository_status": "CREATED_PROTOTYPED", - "repository_topology_operations_executed": true, - "lifecycle_changes_executed": true, - "existing_repository_lifecycle_changes_executed": false, - "existing_repository_dispositions_executed": false, + "historical_reconciliation": "program/history/pre-consolidation/registry.json", "model_provider_runs_added": 0 }, "gearbox_publication": { @@ -30,107 +25,220 @@ "final_main_sha": "f3fab9f292cf4eabd7200615d444f98881f57d55", "implementation_revision": "6005b340a6f6fb3f8683439d6b5fd154e1fd253f", "provider_model_runs": 0, - "repository_consolidations": 0 + "repository_consolidations": 0, + "scope": "Historical initial publication; final_main_sha is the publication revision, not current HEAD." }, "program_control": { - "schema_version": 1, - "priority_order": ["NOW", "NEXT", "THEN", "LATER", "PARKED"], + "schema_version": 2, + "priority_order": [ + "NOW", + "NEXT", + "THEN", + "LATER", + "PARKED" + ], "current_lane": "NOW", - "operating_question": "What prevents Durable Supervisor from successfully finishing Opsle Tasks with less intelligence, less context, less human involvement, and defensible evidence?", - "exact_next_execution": "In opsle/durable-supervisor, make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening Durable Supervisor v0.1.", + "operating_question": "How can current Opsle Tasks workloads use less intelligence and context while preserving correctness and defensible evidence?", + "exact_next_execution": "Select the next authorized task from Opsle Tasks; use its project scope and workload evidence to identify the smallest justified change. This ledger does not create a parallel queue.", "work_item_admission": { "rule": "Do not create a work item solely because an implementation can be improved.", - "qualifying_reasons": ["violated invariant", "demonstrated defect", "measured inefficiency", "missing capability blocking the current program objective", "experiment requirement", "security or safety issue", "externally required release condition"], - "parked_by_default": ["cosmetic cleanup", "architectural taste", "hypothetical robustness", "speculative future requirement"] + "qualifying_reasons": [ + "violated invariant", + "demonstrated defect", + "measured inefficiency", + "missing capability blocking the current program objective", + "experiment requirement", + "security or safety issue", + "externally required release condition" + ], + "parked_by_default": [ + "cosmetic cleanup", + "architectural taste", + "hypothetical robustness", + "speculative future requirement" + ] }, "lanes": [ { "name": "NOW", - "objective": "Finish Durable Supervisor v0.1 as a bounded measured system, then freeze feature work.", - "repositories": ["durable-supervisor", "research"], - "entry_condition": "Current program lane.", - "exit_condition": "Every Durable Supervisor v0.1 stopping criterion is satisfied and the release is explicitly declared and frozen." + "repositories": [ + "tasks", + "research" + ], + "objective": "Maintain the current Tasks workload and evidence-backed program state.", + "entry_condition": "An authorized Tasks work item and a qualifying evidence-backed reason are required.", + "exit_condition": "The scoped defect or question is resolved with bounded evidence; release requirements remain separate." }, { "name": "NEXT", - "objective": "Use Opsle Tasks as the primary real-world workload and collect integrated measurements from the current implementation.", - "repositories": ["gearbox", "context-firewall", "decision-evidence-protocol", "agent-trajectory-profiler", "affected-verification"], - "entry_condition": "Durable Supervisor v0.1 is declared and feature-frozen.", - "exit_condition": "Durable Supervisor has driven the remaining authorized Opsle Tasks readiness work and produced defensible integrated measurements." + "repositories": [ + "gearbox", + "context-firewall", + "visible-value", + "affected-verification", + "decision-evidence-protocol", + "agent-trajectory-profiler" + ], + "objective": "Address demonstrated integration or measurement deficiencies from Tasks workloads.", + "entry_condition": "An authorized Tasks work item and a qualifying evidence-backed reason are required.", + "exit_condition": "The scoped defect or question is resolved with bounded evidence; release requirements remain separate." }, { "name": "THEN", - "objective": "Advance an individual concept only when Durable Supervisor or Opsle Tasks evidence demonstrates a concrete need.", - "repositories": ["semantic-edit-protocol", "event-driven-agent-wakeup", "agent-state-ledger", "agent-scheduler-runtime", "verifiable-agent-handoff", "agent-routing-policy", "agent-resource-claims", "agent-discovery-control", "agent-execution-authorization", "controlled-agent-acceptance", "agent-recovery-policy", "ephemeral-agent-workers"], - "entry_condition": "A qualifying work-item reason and real workload evidence identify the smallest relevant concept boundary.", - "exit_condition": "The demonstrated deficiency is resolved or falsified at its narrowest justified boundary; do not march mechanically through the portfolio." + "repositories": [ + "semantic-edit-protocol" + ], + "objective": "Investigate semantic edits only when workload evidence establishes a need.", + "entry_condition": "An authorized Tasks work item and a qualifying evidence-backed reason are required.", + "exit_condition": "The scoped defect or question is resolved with bounded evidence; release requirements remain separate." }, { "name": "LATER", - "objective": "Run controlled experiments, freeze a real-workload benchmark corpus, seek independent replication, and only then consider public product and research-site releases.", - "repositories": ["site"], - "entry_condition": "NOW, NEXT, and evidence-triggered THEN work establish a defensible need and release prerequisites are separately authorized.", - "exit_condition": "Applicable evidence and separate release authorization exist." + "repositories": [ + "site" + ], + "objective": "Prepare controlled research and public material when evidence and separate release authority support it.", + "entry_condition": "An authorized Tasks work item and a qualifying evidence-backed reason are required.", + "exit_condition": "The scoped defect or question is resolved with bounded evidence; release requirements remain separate." }, { "name": "PARKED", - "objective": "Retain useful non-priority ideas without turning them into active work.", - "repositories": [".github"], - "entry_condition": "The idea is useful but lacks a qualifying reason to compete with the current objective.", - "exit_condition": "New evidence supplies a qualifying work-item reason." + "repositories": [ + ".github" + ], + "objective": "Retain organization housekeeping separately from workload eligibility.", + "entry_condition": "An authorized Tasks work item and a qualifying evidence-backed reason are required.", + "exit_condition": "The scoped defect or question is resolved with bounded evidence; release requirements remain separate." } ], - "durable_supervisor_v0_1": { - "status": "IN_PROGRESS", - "durability_foundation": "VERIFIED_ENOUGH_TO_STOP_NITPICKING_UNLESS_REAL_WORKLOAD_EXPOSES_A_DEFECT", - "verified_main_sha": "1b5ab7631ba651a32592bbbdab8001865a3baf3d", - "verified_runtime_state": "PAUSED_NO_ACTIVE_TASK_OR_ATTEMPT", - "runtime_state_recorded_at": "2026-09-05T12:20:01.551Z", - "runtime_state_source": "Read-only .opsle/state.json authority showed supervisor_state PAUSED with null active_task_id and active_attempt_id; .opsle/supervisor.json showed AUTHORITATIVE authority.", - "stopping_criteria": [ - {"id": "DS-V0.1-01", "status": "OPEN", "criterion": "Expose the selected Gearbox route, routing rationale, and Context Firewall accounting in an operator-visible per-child receipt."}, - {"id": "DS-V0.1-02", "status": "OPEN", "criterion": "Expose each child's model and reasoning effort."}, - {"id": "DS-V0.1-03", "status": "OPEN", "criterion": "Record actual tokens and cost when provider evidence supplies them, and clearly labeled estimates otherwise."}, - {"id": "DS-V0.1-04", "status": "OPEN", "criterion": "Measure raw evidence or context versus Context Firewall retained context and report the reduction."}, - {"id": "DS-V0.1-05", "status": "OPEN", "criterion": "Report first-pass success rate."}, - {"id": "DS-V0.1-06", "status": "OPEN", "criterion": "Report repair-child or retry rate and the token cost of retries."}, - {"id": "DS-V0.1-07", "status": "OPEN", "criterion": "Estimate avoidable intelligence consumption by identifying failures whose needed facts were discoverable through deterministic preflight before model execution."}, - {"id": "DS-V0.1-08", "status": "OPEN", "criterion": "Implement only the smallest useful deterministic preflight or reconnaissance mechanism shown necessary by the avoidable-failure evidence."}, - {"id": "DS-V0.1-09", "status": "OPEN", "criterion": "Complete another real foreign-repository workload that exercises and preserves the new measurements."}, - {"id": "DS-V0.1-10", "status": "OPEN", "criterion": "Declare Durable Supervisor v0.1 and freeze feature work except for defects exposed by real workloads."} - ] - }, "opsle_tasks": { "current_repository": "opsle/tasks", - "future_name": "Opsle Tasks", - "role": "NEXT primary real-world workload after Durable Supervisor v0.1", - "measurements": ["Gearbox", "Context Firewall", "Decision Evidence Protocol", "Agent Trajectory Profiler", "Affected Verification"], - "prohibited_without_separate_authorization": ["public release", "DNS or TLS changes", "launch provider work"] + "role": "current workload and task-management authority", + "measurements": [ + "Gearbox", + "Context Firewall", + "Visible Value", + "Affected Verification" + ], + "prohibited_without_separate_authorization": [ + "public release", + "DNS or TLS changes", + "launch provider work" + ], + "name": "Opsle Tasks", + "capability_contract": "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/CAPABILITIES.md", + "capability_ids": [ + "opsle.gearbox", + "opsle.context-firewall", + "opsle.affected-verification", + "opsle.visible-value" + ] }, "concept_activation": [ - {"deficiency": "routing", "repositories": ["agent-routing-policy"]}, - {"deficiency": "durable state", "repositories": ["agent-state-ledger"]}, - {"deficiency": "retry or deterministic preflight", "repositories": ["gearbox"]}, - {"deficiency": "context overload", "repositories": ["context-firewall"]}, - {"deficiency": "verification selection", "repositories": ["affected-verification"]}, - {"deficiency": "recovery", "repositories": ["agent-recovery-policy"]} + { + "deficiency": "routing", + "repositories": [ + "gearbox" + ] + }, + { + "deficiency": "durable state", + "repositories": [ + "tasks" + ] + }, + { + "deficiency": "retry or deterministic preflight", + "repositories": [ + "gearbox" + ] + }, + { + "deficiency": "context overload", + "repositories": [ + "context-firewall" + ] + }, + { + "deficiency": "verification selection", + "repositories": [ + "affected-verification" + ] + }, + { + "deficiency": "recovery", + "repositories": [ + "tasks" + ] + } + ], + "later_items": [ + "controlled empirical experiments", + "frozen real-workload benchmark corpus", + "independent replication", + "Opsle Tasks public and self-hosted release", + "opsle.com research and public site", + "hosted Opsle offering" + ], + "parked_items": [ + "Cosmetic cleanup, architectural polishing, hypothetical robustness and speculative requirements without qualifying evidence." ], - "later_items": ["controlled empirical experiments", "frozen real-workload benchmark corpus", "independent replication", "Opsle Tasks public and self-hosted release", "opsle.com research and public site", "hosted Opsle offering"], - "parked_items": ["Durable Supervisor src/cli.js is 1,570 lines at the verified main SHA and may eventually warrant decomposition; this is a maintenance smell, not a current objective, unless measured work shows material reliability or efficiency harm.", "Background projection repair or reconciliation is not justified merely because explicit retry exists.", "Historical pre-fix cleanup or migration requires evidence of need.", "General architectural polishing remains parked unless a real workload exposes a concrete defect."] + "workload_authority": "opsle/tasks" }, "visible_value": { - "baseline": "The configured primary supervisor model and reasoning effort performs all child work itself and receives raw unfiltered evidence.", + "baseline": "The configured primary model and reasoning effort performs the same task work and receives raw unfiltered evidence.", "baseline_rule": "Savings require an inspectable same-work baseline or a clearly labeled estimate derived from that baseline; marketing counterfactuals are prohibited.", "value_kinds": [ - {"name": "MEASURED", "definition": "The value comes directly from deterministic artifacts or provider usage records and still carries the applicable Visible Value evidence class, normally EXACT or OBSERVED."}, - {"name": "DERIVED", "definition": "The value is a reproducible calculation over identified measured inputs and records its method and applicable Visible Value evidence class."}, - {"name": "ESTIMATED", "definition": "The value carries the ESTIMATED evidence class and names its method, assumptions, pricing source, and uncertainty."}, - {"name": "UNAVAILABLE", "definition": "The value is not currently supported and must remain absent rather than being rendered as zero."} + { + "name": "MEASURED", + "definition": "The value comes directly from deterministic artifacts or provider usage records and still carries the applicable Visible Value evidence class, normally EXACT or OBSERVED." + }, + { + "name": "DERIVED", + "definition": "The value is a reproducible calculation over identified measured inputs and records its method and applicable Visible Value evidence class." + }, + { + "name": "ESTIMATED", + "definition": "The value carries the ESTIMATED evidence class and names its method, assumptions, pricing source, and uncertainty." + }, + { + "name": "UNAVAILABLE", + "definition": "The value is not currently supported and must remain absent rather than being rendered as zero." + } ], - "per_child_receipt_fields": ["child/task identity", "model", "reasoning effort", "Gearbox route", "routing rationale", "input tokens", "output tokens", "raw evidence/context size", "Context Firewall retained size", "reduction percentage", "estimated tokens avoided", "estimated cost avoided", "duration", "attempt number", "success/failure", "escalation/retry reason", "retry potentially avoidable through deterministic preflight"], - "supervisor_summary_fields": ["total children", "model/effort distribution", "total model tokens consumed", "work completed deterministically without model use", "Context Firewall reduction", "estimated token/cost savings", "first-pass success rate", "repair-child rate", "tokens spent on retries", "avoidable-intelligence estimate"] + "per_child_receipt_fields": [ + "child/task identity", + "model", + "reasoning effort", + "Gearbox route", + "routing rationale", + "input tokens", + "output tokens", + "raw evidence/context size", + "Context Firewall retained size", + "reduction percentage", + "estimated tokens avoided", + "estimated cost avoided", + "duration", + "attempt number", + "success/failure", + "escalation/retry reason", + "retry potentially avoidable through deterministic preflight" + ], + "run_summary_fields": [ + "total children", + "model/effort distribution", + "total model tokens consumed", + "work completed deterministically without model use", + "Context Firewall reduction", + "estimated token/cost savings", + "first-pass success rate", + "repair-child rate", + "tokens spent on retries", + "avoidable-intelligence estimate" + ] }, - "last_verified_at": "2026-09-05T15:28:58Z", + "last_verified_at": "2026-09-09T00:00:00Z", "repositories": [ { "name": "agent-trajectory-profiler", @@ -143,24 +251,51 @@ "implementation_status": "dependency-free JavaScript trajectory profiler with value-receipt ingestion, observational run records, class/unit/trust-safe per-run and cumulative summaries, Context Firewall adapter, canonical CLI/API, and content-addressed conformance corpus", "implementation_requirement": "A runnable profiler with documented metric semantics is required.", "specification_status": "versioned trajectory, measurement, observational run-record, value-receipt ingestion, and deterministic value-summary contracts with trust, class, unit, escalation, missing-data, deduplication, and invalid-state semantics", - "test_status": "74 of 74 automated tests and 14 of 14 content-addressed measurement fixtures passed locally at the verified HEAD; 5 of 5 pinned exact-revision interoperability cases and 5 determinism tests passed; PR #2 CI passed", + "test_status": "Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation.", "benchmark_status": "14 deterministic measurement fixtures, a public-synthetic two-receipt interoperability chain, one observational implementation-run record, the research-owned frozen EXP-001 corpus/oracle/arm/provider-free trajectory harness at 04234a65bf36192d63f1dd173c440d45a6604d2b, launch preregistration at 31848c3f25ff9371055932657e8e2f8ad54cc8c7, and eight coordinator qualification profiles at 9ee43197880c18d4e185cf7e29e02a151d22a12e; no model run", "measured_experiment_status": "EXP-001 remains planned; zero model/provider experiment runs", "reproducibility_status": "tests, canonical conformance, byte-identical fixture regeneration, value receipt/run-record validation, deterministic summaries, and exact-revision Context Firewall/Decision Evidence interoperability are reproducible locally and in CI; comparative benefit and model correctness are untested", "documentation_status": "public theory, versioned specification, architecture, byte/event/token semantics, Visible Value ingestion and aggregation rules, observational corpus fields, operator channel separation, fixture provenance, EXP-001 boundary, and limitations documented", "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Receipt aggregation requires explicit safe declarations and exact compatibility partitions; ordinary records reject EXPERIMENTAL measurements, only Context Firewall packet-v1 has a trajectory adapter, and model correctness, token/cost savings, causal benefit, a safe reduction frontier, and replication remain unmeasured."], + "known_limitations": [ + "Receipt aggregation requires explicit safe declarations and exact compatibility partitions; ordinary records reject EXPERIMENTAL measurements, only Context Firewall packet-v1 has a trajectory adapter, and model correctness, token/cost savings, causal benefit, a safe reduction frontier, and replication remain unmeasured." + ], "dependencies": [], - "dependents": ["context-firewall", "gearbox", "semantic-edit-protocol"], - "active_experiment_ids": ["EXP-001"], - "blockers": ["EXP-001 exact budgeted authorization, catalogue/pricing preflight, measured experiment, and independent replication remain missing; the one-block coordinator is provider-free qualified and zero model/provider subjects have run."], - "next_task": "Wait for a Durable Supervisor or Opsle Tasks workload to require trajectory measurement; do not advance the profiler merely to polish it.", - "evidence": ["https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/value-summary.js", "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/run-record.js", "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/tests/value-summary.test.js", "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/tools/verify-context-firewall-interop.js", "https://github.com/opsle/agent-trajectory-profiler/pull/2", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/profile.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", "https://github.com/opsle/research/pull/11", "https://github.com/opsle/research/pull/13"], - "completion_criteria": ["Publish metric semantics and executable profiler.", "Run correctness-gated benchmarks across multiple trajectories.", "Replicate predictive-value claims and document failure modes."], + "dependents": [], + "active_experiment_ids": [ + "EXP-001" + ], + "blockers": [ + "EXP-001 is prepared but unconsumed; measured subject decision adequacy, broader workload coverage and independent replication remain absent." + ], + "next_task": "Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository.", + "evidence": [ + "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/README.md", + "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/value-summary.js", + "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/run-record.js", + "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/tests/value-summary.test.js", + "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/tools/verify-context-firewall-interop.js", + "https://github.com/opsle/agent-trajectory-profiler/pull/2", + "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/profile.json", + "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", + "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", + "https://github.com/opsle/research/pull/11", + "https://github.com/opsle/research/pull/13" + ], + "completion_criteria": [ + "Publish metric semantics and executable profiler.", + "Run correctness-gated benchmarks across multiple trajectories.", + "Replicate predictive-value claims and document failure modes." + ], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "ACTIVE", + "consolidated_into": null, + "workload_eligible": true, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "semantic-edit-protocol", @@ -173,114 +308,191 @@ "implementation_status": "none; placeholder source directory only", "implementation_requirement": "A reference language adapter and transactional validator are required.", "specification_status": "experimental theory contract; concept-specific operation schema remains incomplete", - "test_status": "placeholder only; no automated tests", + "test_status": "Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation.", "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", "measured_experiment_status": "none", "reproducibility_status": "not available", "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Language adapters, conflict semantics, anchor durability measurements, and controlled patch comparisons are missing."], - "dependencies": ["agent-resource-claims", "agent-trajectory-profiler"], + "known_limitations": [ + "Language adapters, conflict semantics, anchor durability measurements, and controlled patch comparisons are missing." + ], + "dependencies": [], "dependents": [], "active_experiment_ids": [], - "blockers": ["No executable semantic operation, validator, tests, or benchmark fixtures."], - "next_task": "Wait until a real Durable Supervisor or Opsle Tasks workload demonstrates a semantic-edit deficiency that simpler bounded edits cannot satisfy.", - "evidence": ["https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/THEORY.md", "https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/SPEC.md"], - "completion_criteria": ["Publish a precise operation contract and executable reference adapter.", "Benchmark against patch and whole-file baselines under identical correctness gates.", "Replicate results and document conflicts and unsupported syntax."], + "blockers": [ + "No executable semantic operation, validator, tests, or benchmark fixtures." + ], + "next_task": "Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository.", + "evidence": [ + "https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/README.md", + "https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/THEORY.md", + "https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/SPEC.md" + ], + "completion_criteria": [ + "Publish a precise operation contract and executable reference adapter.", + "Benchmark against patch and whole-file baselines under identical correctness gates.", + "Replicate results and document conflicts and unsupported syntax." + ], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "ACTIVE", + "consolidated_into": null, + "workload_eligible": true, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "durable-supervisor", "github_url": "https://github.com/opsle/durable-supervisor", "default_branch": "main", - "last_verified_head_sha": "1b5ab7631ba651a32592bbbdab8001865a3baf3d", + "last_verified_head_sha": "a69f3b1ce523912a1f840cea305db008bab8e03f", "project_type": "concept", "purpose": "Own autonomous objective progress through durable authority, bounded child execution, evidence reduction, evaluation, pause, wake, and reconstruction across model activations.", "lifecycle_stage": "VERIFIED", - "implementation_status": "dependency-free Node.js Durable Supervisor v0.1 with persistent supervisor authority, task handoff, discovery, exact Gearbox routing, claims and fencing, detached Runner ownership, Context Firewall packets, acceptance and immutable evaluation, opsled wake delivery, bounded reconstruction, policy controls, and runtime-release fencing", + "implementation_status": "Historical source preserved; runtime and repository retired; findings preserved in Tasks.", "implementation_requirement": "A durable supervisor/runner reference implementation is required.", - "specification_status": "public v0.1 contract defines durable authority, immutable history, task and attempt state, exact routes, claims and fencing, Runner ownership, Context Firewall and acceptance boundaries, evaluation, wake, reconstruction, pause, and runtime compatibility", - "test_status": "PR #2 recorded 141 full-suite tests and canonical validation passing; PR #3 recorded 176 passing with 1 intentional skip; D5 PR #4 passed 6 focused atomicity, operator, and wake checks; projection-reconciliation PR #5 passed 76 focused evaluation, invariant, operator, recovery, and wake checks plus syntax, release-manifest, and canonical validation at the exact head whose tree is main", - "benchmark_status": "bounded self-hosting evidence records meaningful child work, exact route and acceptance artifacts, and measured Context Firewall byte reduction; PR #2 additionally records disposable and real foreign-repository portability plus a live automatic wake, but no frozen comparative baseline or controlled token/cost benchmark exists", - "measured_experiment_status": "operational self-hosting and foreign-repository observations only; no controlled comparative experiment", - "reproducibility_status": "provider-free tests, schema and release checks, reconstruction, portability fixtures, concurrency campaigns, and explicit evaluation replay are reproducible from the public source; latest PRs used focused affected verification rather than a new full-suite run, and no independent replication exists", - "documentation_status": "public v0.1 README, normative specification, architecture, operations guidance, and bounded self-hosting proof document supported behavior and explicit claim limits", - "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Operator receipts do not yet expose the full Gearbox route/rationale, child model/effort, actual token/cost, retry cost, or avoidable-intelligence accounting required for v0.1 completion; existing Context Firewall evidence is byte-level, the latest repairs used focused rather than full-suite verification, and comparative benefit and independent replication remain unproven."], - "dependencies": ["agent-state-ledger", "agent-scheduler-runtime", "event-driven-agent-wakeup", "decision-evidence-protocol"], + "specification_status": "Historical observation before retirement: public v0.1 contract defines durable authority, immutable history, task and attempt state, exact routes, claims and fencing, Runner ownership, Context Firewall and acceptance boundaries, evaluation, wake, reconstruction, pause, and runtime compatibility", + "test_status": "Historical observation before retirement: PR #2 recorded 141 full-suite tests and canonical validation passing; PR #3 recorded 176 passing with 1 intentional skip; D5 PR #4 passed 6 focused atomicity, operator, and wake checks; projection-reconciliation PR #5 passed 76 focused evaluation, invariant, operator, recovery, and wake checks plus syntax, release-manifest, and canonical validation at the exact head whose tree is main", + "benchmark_status": "Historical observation before retirement: bounded self-hosting evidence records meaningful child work, exact route and acceptance artifacts, and measured Context Firewall byte reduction; PR #2 additionally records disposable and real foreign-repository portability plus a live automatic wake, but no frozen comparative baseline or controlled token/cost benchmark exists", + "measured_experiment_status": "Historical observation before retirement: operational self-hosting and foreign-repository observations only; no controlled comparative experiment", + "reproducibility_status": "Historical observation before retirement: provider-free tests, schema and release checks, reconstruction, portability fixtures, concurrency campaigns, and explicit evaluation replay are reproducible from the public source; latest PRs used focused affected verification rather than a new full-suite run, and no independent replication exists", + "documentation_status": "Historical observation before retirement: public v0.1 README, normative specification, architecture, operations guidance, and bounded self-hosting proof document supported behavior and explicit claim limits", + "site_publication_status": "Historical observation before retirement: GitHub documentation only; no evidence-backed site publication", + "known_limitations": [ + "Operator receipts do not yet expose the full Gearbox route/rationale, child model/effort, actual token/cost, retry cost, or avoidable-intelligence accounting required for v0.1 completion; existing Context Firewall evidence is byte-level, the latest repairs used focused rather than full-suite verification, and comparative benefit and independent replication remain unproven." + ], + "dependencies": [], "dependents": [], "active_experiment_ids": [], - "blockers": ["Durable Supervisor v0.1 still lacks the operator-visible measurement surface, avoidable-intelligence classification, smallest evidence-driven preflight, and another measured foreign-repository workload required by the program stopping criteria."], - "next_task": "Make the selected Gearbox route, routing rationale, and Context Firewall raw-versus-retained accounting operator-visible in each child receipt without broadening v0.1.", - "evidence": ["https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/README.md", "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/SPEC.md", "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/ARCHITECTURE.md", "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/docs/SELF_HOSTING_PROOF.md", "https://github.com/opsle/durable-supervisor/pull/2", "https://github.com/opsle/durable-supervisor/pull/3", "https://github.com/opsle/durable-supervisor/pull/4", "https://github.com/opsle/durable-supervisor/pull/5"], - "completion_criteria": ["Satisfy the ten bounded Durable Supervisor v0.1 stopping criteria in the program-control registry.", "Prove the measurement surface on another real foreign-repository workload.", "Declare v0.1 and freeze feature work except for workload-exposed defects."], + "blockers": [], + "next_task": "None; historical source is retired from workload selection.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/README.md", + "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/SPEC.md", + "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/ARCHITECTURE.md", + "https://github.com/opsle/durable-supervisor/blob/1b5ab7631ba651a32592bbbdab8001865a3baf3d/docs/SELF_HOSTING_PROOF.md", + "https://github.com/opsle/durable-supervisor/pull/2", + "https://github.com/opsle/durable-supervisor/pull/3", + "https://github.com/opsle/durable-supervisor/pull/4", + "https://github.com/opsle/durable-supervisor/pull/5" + ], + "completion_criteria": [ + "No current completion gate: retired source; highest evidenced maturity and original criteria preserved in historical_evidence." + ], "completion_evidence": [], - "completion_status": "INCOMPLETE", - "program_state": "active", - "last_verified_at": "2026-09-05T15:28:58Z" + "completion_status": "SUPERSEDED", + "program_state": "historical", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "RETIRED", + "consolidated_into": null, + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "event-driven-agent-wakeup", "github_url": "https://github.com/opsle/event-driven-agent-wakeup", "default_branch": "main", - "last_verified_head_sha": "a6209860c2151450cc28ed648bc8c2631c8db7ef", + "last_verified_head_sha": "8278bfb064eba50ce05aeaf9c1b7cc54a748c210", "project_type": "concept", "purpose": "Suspend model activity during waits and wake idempotently on durable decision-relevant events.", "lifecycle_stage": "PROTOTYPED", - "implementation_status": "dependency-free JavaScript state-machine prototype", + "implementation_status": "Historical source preserved; concept contracts consolidated into tasks.", "implementation_requirement": "A durable adapter plus restart-safe reference implementation is required.", - "specification_status": "experimental prototype contract", - "test_status": "2 of 2 automated tests passed locally at the verified HEAD", - "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", - "measured_experiment_status": "none", - "reproducibility_status": "prototype tests reproducible locally; inference-avoidance benefit not reproduced", - "documentation_status": "public theory, specification, architecture, and benchmark plan present", - "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Restart campaigns, duplicate-event campaigns, wake-latency baselines, and multi-runtime adapters are missing."], - "dependencies": ["agent-state-ledger"], - "dependents": ["durable-supervisor"], + "specification_status": "Historical observation before retirement: experimental prototype contract", + "test_status": "Historical observation before retirement: 2 of 2 automated tests passed locally at the verified HEAD", + "benchmark_status": "Historical observation before retirement: prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "Historical observation before retirement: none", + "reproducibility_status": "Historical observation before retirement: prototype tests reproducible locally; inference-avoidance benefit not reproduced", + "documentation_status": "Historical observation before retirement: public theory, specification, architecture, and benchmark plan present", + "site_publication_status": "Historical observation before retirement: GitHub documentation only; no evidence-backed site publication", + "known_limitations": [ + "Restart campaigns, duplicate-event campaigns, wake-latency baselines, and multi-runtime adapters are missing." + ], + "dependencies": [], + "dependents": [], "active_experiment_ids": [], - "blockers": ["Current prototype is in-memory and has no restart or durable-store harness."], - "next_task": "Wait for a real Durable Supervisor or Opsle Tasks wake deficiency; advance the standalone concept only if the integrated runtime exposes one.", - "evidence": ["https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/src/runtime.js", "https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/tests/runtime.test.js"], - "completion_criteria": ["Implement restart-safe durable wait registration and wakeup.", "Benchmark polling and model-turn baselines.", "Replicate event-loss and duplicate-delivery correctness."], + "blockers": [], + "next_task": "None; historical source is retired from workload selection.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/src/runtime.js", + "https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/tests/runtime.test.js" + ], + "completion_criteria": [ + "No current completion gate: retired source; highest evidenced maturity and original criteria preserved in historical_evidence." + ], "completion_evidence": [], - "completion_status": "INCOMPLETE", - "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "completion_status": "SUPERSEDED", + "program_state": "historical", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "CONSOLIDATED", + "consolidated_into": "tasks", + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "context-firewall", "github_url": "https://github.com/opsle/context-firewall", "default_branch": "main", - "last_verified_head_sha": "953c48f1cfd154d6b7ed10b51b87fe54e4df45f2", + "last_verified_head_sha": "6dd6e5fdf21f28dc5ebfa07954aaa9bed2dbcc32", "project_type": "concept", "purpose": "Reduce operational payload deterministically while preserving provenance and safe raw-evidence escalation.", "lifecycle_stage": "PROTOTYPED", - "implementation_status": "dependency-free deterministic JavaScript TAP-subset reducer with compact evidence packets, separate opsle.value-receipt.v1 sidecars, named operator indicators, payload ceilings, raw escalation, and synthetic conformance corpus", + "implementation_status": "v0.5.0 reducer for flat TAP and Node spec/dot reporters, semantic-only model projection, canonical audit packet and separate value receipt.", "implementation_requirement": "An executable reducer and conformance suite are sufficient; a full agent runtime is not required.", "specification_status": "experimental prototype contract for packet v1, deterministic TAP-subset policy v1, Visible Value receipt profile, deterministic sidecar, and operator/model channel separation", - "test_status": "42 of 42 automated tests and 30 of 30 synthetic conformance fixtures passed locally at the verified HEAD; deterministic stdout/sidecar separation and operator stderr tests passed; PR #2 CI passed", + "test_status": "Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation.", "benchmark_status": "30 synthetic conformance fixtures, one observational dogfood reduction attempt, the research-owned frozen six-task EXP-001 corpus/oracle plus raw and three strict-TAP arm harness at 04234a65bf36192d63f1dd173c440d45a6604d2b, launch preregistration at 31848c3f25ff9371055932657e8e2f8ad54cc8c7, and six coordinator qualification reductions at 9ee43197880c18d4e185cf7e29e02a151d22a12e; no model run", "measured_experiment_status": "EXP-001 remains planned; zero model/provider experiment runs", "reproducibility_status": "prototype tests, receipt validation, canonical output checks, and conformance are reproducible locally and in CI; model correctness and a safe reduction frontier are untested", "documentation_status": "public theory, packet/input contract, parser limits, retention/suppression policy, exact Visible Value measurements, operator channel, escalation, ceilings, CLI, conformance, claim limits, and EXP-001 boundary documented", "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["The parser is a strict flat TAP-compatible subset; actual Node checkmark-formatted test output dogfood expanded from 3,425 to 13,027 bytes and correctly escalated, caller raw references are not externally verified, and no evidence establishes model correctness or a safe reduction frontier."], - "dependencies": ["decision-evidence-protocol", "agent-trajectory-profiler"], - "dependents": ["gearbox"], - "active_experiment_ids": ["EXP-001"], - "blockers": ["EXP-001 exact budgeted authorization, catalogue/pricing preflight, measured correctness experiment, and replication remain missing; the one-block coordinator is provider-free qualified and zero model/provider subjects have run."], - "next_task": "Wait for Durable Supervisor or Opsle Tasks evidence of context overload, unsafe omission, or unsupported input before advancing the standalone reducer.", - "evidence": ["https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/src/value-receipt.js", "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/tests/reducer.test.js", "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/fixtures/corpus.js", "https://github.com/opsle/context-firewall/pull/2", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/context-firewall-value-receipt.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", "https://github.com/opsle/research/pull/11", "https://github.com/opsle/research/pull/13"], - "completion_criteria": ["Publish reducer policy and executable conformance validator.", "Run correctness-gated measured context-reduction experiments.", "Replicate the safe frontier and publish omission failure modes."], + "known_limitations": [ + "Not arbitrary-log parsing or a production security boundary. Packet byte metrics do not measure downstream semantic projection delivery; Tasks records actual delivery separately. No token, cost or causal savings established." + ], + "dependencies": [], + "dependents": [ + "tasks" + ], + "active_experiment_ids": [ + "EXP-001" + ], + "blockers": [ + "EXP-001 remains PLANNED with zero consumed authorizations and no subject runs; comparative correctness and replication remain absent." + ], + "next_task": "Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository.", + "evidence": [ + "https://github.com/opsle/context-firewall/blob/6dd6e5fdf21f28dc5ebfa07954aaa9bed2dbcc32/README.md", + "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/src/value-receipt.js", + "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/tests/reducer.test.js", + "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/fixtures/corpus.js", + "https://github.com/opsle/context-firewall/pull/2", + "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/context-firewall-value-receipt.json", + "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", + "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", + "https://github.com/opsle/research/pull/11", + "https://github.com/opsle/research/pull/13" + ], + "completion_criteria": [ + "Publish reducer policy and executable conformance validator.", + "Run correctness-gated measured context-reduction experiments.", + "Replicate the safe frontier and publish omission failure modes." + ], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "ACTIVE", + "consolidated_into": null, + "workload_eligible": true, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "decision-evidence-protocol", @@ -293,444 +505,710 @@ "implementation_status": "dependency-free generic value-receipt validator plus independent Context Firewall packet/source/value-receipt validator, Decision Evidence validation receipts, canonical CLI with named operator indicators, and self-contained conformance suite", "implementation_requirement": "As a protocol, an executable validator and conformance suite can satisfy implementation; a full runtime is not required.", "specification_status": "versioned normative experimental Context Firewall packet-v1 and value-receipt validation profiles plus generic opsle.value-receipt.v1 semantic validation", - "test_status": "60 of 60 automated tests and 24 of 24 public-safe conformance vectors passed locally at the verified HEAD; 5 of 5 exact-revision interoperability cases, 8 determinism tests, operator separation, source-unverified, and tamper paths passed; PR #2 CI passed", + "test_status": "Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation.", "benchmark_status": "24 protocol conformance vectors, a 5-case exact-revision interoperability proof, one source-backed observational dogfood validation, 36 source-backed validations in the research-owned frozen EXP-001 offline harness at 04234a65bf36192d63f1dd173c440d45a6604d2b, launch preregistration at 31848c3f25ff9371055932657e8e2f8ad54cc8c7, and six coordinator qualification validations at 9ee43197880c18d4e185cf7e29e02a151d22a12e; no correctness experiment or model run", "measured_experiment_status": "none", "reproducibility_status": "validator tests, self-contained vectors, deterministic CLI/sidecar output, value-receipt validation, and exact-revision Context Firewall interoperability are reproducible locally and in CI; other tool classes and independent replication are missing", "documentation_status": "public theory, packet-v1 and value-receipt profiles, enforced invariants, API/CLI, named operator channel, structural-versus-cryptographic verification, sufficiency/escalation semantics, claim limits, interoperability, and limitations documented", "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Only Context Firewall's strict TAP-subset packet-v1 producer is independently validated; receipt-only source claims and caller-owned raw locator existence remain unverified without external evidence, and no result establishes model correctness, comparative benefit, or independent replication."], + "known_limitations": [ + "Only Context Firewall's strict TAP-subset packet-v1 producer is independently validated; receipt-only source claims and caller-owned raw locator existence remain unverified without external evidence, and no result establishes model correctness, comparative benefit, or independent replication." + ], "dependencies": [], - "dependents": ["agent-recovery-policy", "context-firewall", "durable-supervisor", "gearbox", "verifiable-agent-handoff"], - "active_experiment_ids": ["EXP-001"], - "blockers": ["Additional real tool classes, measured subject decision adequacy, EXP-001 exact budgeted authorization and catalogue/pricing preflight, a controlled comparative result, and independent replication are missing."], - "next_task": "Wait for an integrated Durable Supervisor or Opsle Tasks receipt to expose a decision-evidence gap that blocks a defensible decision.", - "evidence": ["https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/context-firewall-value.js", "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/value-receipt-v1.js", "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/tests/value-receipt.test.js", "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/docs/value-receipt-v1.md", "https://github.com/opsle/decision-evidence-protocol/pull/2", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/decision-evidence-validation.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", "https://github.com/opsle/research/pull/11", "https://github.com/opsle/research/pull/13"], - "completion_criteria": ["Publish a versioned protocol and comprehensive executable conformance suite.", "Measure adequacy across multiple real tool classes.", "Replicate interoperability and document loss/escalation failure modes."], + "dependents": [], + "active_experiment_ids": [ + "EXP-001" + ], + "blockers": [ + "EXP-001 is prepared but unconsumed; measured subject decision adequacy, broader workload coverage and independent replication remain absent." + ], + "next_task": "Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository.", + "evidence": [ + "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/README.md", + "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/context-firewall-value.js", + "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/value-receipt-v1.js", + "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/tests/value-receipt.test.js", + "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/docs/value-receipt-v1.md", + "https://github.com/opsle/decision-evidence-protocol/pull/2", + "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/program/evidence/exp-001-offline-freeze/decision-evidence-validation.json", + "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/program/evidence/exp-001-preregistration/verification-report.json", + "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/program/evidence/exp-001-block-coordinator/qualification-report.json", + "https://github.com/opsle/research/pull/11", + "https://github.com/opsle/research/pull/13" + ], + "completion_criteria": [ + "Publish a versioned protocol and comprehensive executable conformance suite.", + "Measure adequacy across multiple real tool classes.", + "Replicate interoperability and document loss/escalation failure modes." + ], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "ACTIVE", + "consolidated_into": null, + "workload_eligible": true, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "agent-state-ledger", "github_url": "https://github.com/opsle/agent-state-ledger", "default_branch": "main", - "last_verified_head_sha": "acab03b1ff7168222552050e21e7553b07d00e7c", + "last_verified_head_sha": "509a2a55066fd30ca67ceca14eaf640772b84c31", "project_type": "concept", "purpose": "Record append-only agent work evidence and derive deterministic current-state projections.", "lifecycle_stage": "THEORY", - "implementation_status": "none; placeholder source directory only", + "implementation_status": "Historical source preserved; concept contracts consolidated into tasks.", "implementation_requirement": "A portable append-only store and projection validator are required.", - "specification_status": "experimental theory contract; schema and contradiction semantics remain incomplete", - "test_status": "placeholder only; no automated tests", - "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", - "measured_experiment_status": "none", - "reproducibility_status": "not available", - "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", - "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Portable schema, contradiction semantics, projection correctness, and reconstruction-size benchmarks are missing."], + "specification_status": "Historical observation before retirement: experimental theory contract; schema and contradiction semantics remain incomplete", + "test_status": "Historical observation before retirement: placeholder only; no automated tests", + "benchmark_status": "Historical observation before retirement: prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "Historical observation before retirement: none", + "reproducibility_status": "Historical observation before retirement: not available", + "documentation_status": "Historical observation before retirement: public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "Historical observation before retirement: GitHub documentation only; no evidence-backed site publication", + "known_limitations": [ + "Portable schema, contradiction semantics, projection correctness, and reconstruction-size benchmarks are missing." + ], "dependencies": [], - "dependents": ["agent-discovery-control", "agent-execution-authorization", "agent-recovery-policy", "agent-scheduler-runtime", "durable-supervisor", "event-driven-agent-wakeup"], + "dependents": [], "active_experiment_ids": [], - "blockers": ["No portable schema, implementation, projection oracle, or replay fixtures."], - "next_task": "Wait for a real Durable Supervisor or Opsle Tasks workload to expose a portable durable-state deficiency before advancing this repository.", - "evidence": ["https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/THEORY.md", "https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/SPEC.md"], - "completion_criteria": ["Implement append-only storage and deterministic projection.", "Prove replay/reconstruction against failure fixtures.", "Measure and replicate reconstruction size and correctness."], + "blockers": [], + "next_task": "None; historical source is retired from workload selection.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/THEORY.md", + "https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/SPEC.md" + ], + "completion_criteria": [ + "No current completion gate: retired source; highest evidenced maturity and original criteria preserved in historical_evidence." + ], "completion_evidence": [], - "completion_status": "INCOMPLETE", - "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "completion_status": "SUPERSEDED", + "program_state": "historical", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "CONSOLIDATED", + "consolidated_into": "tasks", + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "agent-scheduler-runtime", "github_url": "https://github.com/opsle/agent-scheduler-runtime", "default_branch": "main", - "last_verified_head_sha": "d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0", + "last_verified_head_sha": "261adff7792b12257e57b408a84699479d2b992d", "project_type": "concept", "purpose": "Own queues, dependencies, schedules, leases, cancellation, and pause as deterministic runtime mechanics.", "lifecycle_stage": "THEORY", - "implementation_status": "none; placeholder source directory only", + "implementation_status": "Historical source preserved; concept contracts consolidated into tasks.", "implementation_requirement": "A portable deterministic scheduler core and store adapter are required.", - "specification_status": "experimental theory contract; portable store and fairness semantics remain incomplete", - "test_status": "placeholder only; no automated tests", - "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", - "measured_experiment_status": "none", - "reproducibility_status": "not available", - "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", - "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Portable store interface, fairness/latency evidence, reasoning boundary, and external adapters are missing."], - "dependencies": ["agent-resource-claims", "agent-state-ledger"], - "dependents": ["controlled-agent-acceptance", "durable-supervisor"], + "specification_status": "Historical observation before retirement: experimental theory contract; portable store and fairness semantics remain incomplete", + "test_status": "Historical observation before retirement: placeholder only; no automated tests", + "benchmark_status": "Historical observation before retirement: prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "Historical observation before retirement: none", + "reproducibility_status": "Historical observation before retirement: not available", + "documentation_status": "Historical observation before retirement: public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "Historical observation before retirement: GitHub documentation only; no evidence-backed site publication", + "known_limitations": [ + "Portable store interface, fairness/latency evidence, reasoning boundary, and external adapters are missing." + ], + "dependencies": [], + "dependents": [], "active_experiment_ids": [], - "blockers": ["Ledger and claim interfaces are not portable or executable in these repositories."], - "next_task": "Wait for a real workload to expose a scheduler-runtime deficiency not already satisfied by Durable Supervisor's local implementation.", - "evidence": ["https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/THEORY.md", "https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/BENCHMARK.md"], - "completion_criteria": ["Implement the portable deterministic core and store contract.", "Pass restart, lease, pause, cancellation, duplicate, and fairness campaigns.", "Benchmark and replicate latency and correctness."], + "blockers": [], + "next_task": "None; historical source is retired from workload selection.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/THEORY.md", + "https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/BENCHMARK.md" + ], + "completion_criteria": [ + "No current completion gate: retired source; highest evidenced maturity and original criteria preserved in historical_evidence." + ], "completion_evidence": [], - "completion_status": "INCOMPLETE", - "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "completion_status": "SUPERSEDED", + "program_state": "historical", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "CONSOLIDATED", + "consolidated_into": "tasks", + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "verifiable-agent-handoff", "github_url": "https://github.com/opsle/verifiable-agent-handoff", "default_branch": "main", - "last_verified_head_sha": "399e5cfae94345affa3f087f0f6eb9e77669d33c", + "last_verified_head_sha": "b832c770d9965b6060fc21af4f8826812c80b0a9", "project_type": "concept", "purpose": "Seal exact execution results before source destruction so a fresh verifier can reconstruct and validate them.", "lifecycle_stage": "PROTOTYPED", - "implementation_status": "dependency-free JavaScript HMAC manifest prototype", + "implementation_status": "Historical source preserved; concept contracts consolidated into tasks.", "implementation_requirement": "As a protocol, an executable sealing/verifying library and adversarial conformance suite can satisfy implementation.", - "specification_status": "experimental prototype contract", - "test_status": "3 of 3 automated tests passed locally at the verified HEAD", - "benchmark_status": "prose plan only; no artifact reconstruction harness or frozen fixtures", - "measured_experiment_status": "EXP-001 potential support only; no runs", - "reproducibility_status": "prototype tests reproducible locally; cross-environment handoff not reproduced", - "documentation_status": "public theory, specification, architecture, and benchmark plan present", - "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Independent implementations, cryptographic agility, cross-VCS artifacts, byte ceilings, and correlated-error studies are missing."], - "dependencies": ["decision-evidence-protocol"], - "dependents": ["controlled-agent-acceptance", "ephemeral-agent-workers"], - "active_experiment_ids": ["EXP-001"], - "blockers": ["Prototype authenticates a manifest but does not build, transport, or reconstruct artifacts."], - "next_task": "Wait for a real workload to require evidence survival across an isolation or source-destruction boundary before advancing this protocol.", - "evidence": ["https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/src/seal.js", "https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/tests/seal.test.js"], - "completion_criteria": ["Publish a versioned seal/reconstruction contract and executable conformance suite.", "Benchmark full artifact handoff and failure cases.", "Replicate with an independent implementation and environment."], + "specification_status": "Historical observation before retirement: experimental prototype contract", + "test_status": "Historical observation before retirement: 3 of 3 automated tests passed locally at the verified HEAD", + "benchmark_status": "Historical observation before retirement: prose plan only; no artifact reconstruction harness or frozen fixtures", + "measured_experiment_status": "Historical observation before retirement: EXP-001 potential support only; no runs", + "reproducibility_status": "Historical observation before retirement: prototype tests reproducible locally; cross-environment handoff not reproduced", + "documentation_status": "Historical observation before retirement: public theory, specification, architecture, and benchmark plan present", + "site_publication_status": "Historical observation before retirement: GitHub documentation only; no evidence-backed site publication", + "known_limitations": [ + "Independent implementations, cryptographic agility, cross-VCS artifacts, byte ceilings, and correlated-error studies are missing." + ], + "dependencies": [], + "dependents": [], + "active_experiment_ids": [], + "blockers": [], + "next_task": "None; historical source is retired from workload selection.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/src/seal.js", + "https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/tests/seal.test.js" + ], + "completion_criteria": [ + "No current completion gate: retired source; highest evidenced maturity and original criteria preserved in historical_evidence." + ], "completion_evidence": [], - "completion_status": "INCOMPLETE", - "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "completion_status": "SUPERSEDED", + "program_state": "historical", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "CONSOLIDATED", + "consolidated_into": "tasks", + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [ + "EXP-001" + ] }, { "name": "agent-routing-policy", "github_url": "https://github.com/opsle/agent-routing-policy", "default_branch": "main", - "last_verified_head_sha": "43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8", + "last_verified_head_sha": "179fb15add41f338813bc04cc77c308563942a93", "project_type": "concept", "purpose": "Select purpose-bound agent routes under exact capability, access, independence, budget, and availability constraints.", "lifecycle_stage": "THEORY", - "implementation_status": "none; placeholder source directory only", + "implementation_status": "Historical source preserved; concept contracts consolidated into gearbox.", "implementation_requirement": "An executable deterministic policy evaluator and trace validator are sufficient.", - "specification_status": "experimental theory contract; route schema and ordering semantics remain incomplete", - "test_status": "placeholder only; no automated tests", - "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", - "measured_experiment_status": "none", - "reproducibility_status": "not available", - "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", - "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Quality-adjusted price evidence, provider comparability, portability, and stability/fairness analysis are missing."], + "specification_status": "Historical observation before retirement: experimental theory contract; route schema and ordering semantics remain incomplete", + "test_status": "Historical observation before retirement: placeholder only; no automated tests", + "benchmark_status": "Historical observation before retirement: prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "Historical observation before retirement: none", + "reproducibility_status": "Historical observation before retirement: not available", + "documentation_status": "Historical observation before retirement: public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "Historical observation before retirement: GitHub documentation only; no evidence-backed site publication", + "known_limitations": [ + "Quality-adjusted price evidence, provider comparability, portability, and stability/fairness analysis are missing." + ], "dependencies": [], - "dependents": ["agent-recovery-policy", "controlled-agent-acceptance", "gearbox"], + "dependents": [], "active_experiment_ids": [], - "blockers": ["No normative route schema, evaluator, fixtures, or provider-independent quality evidence."], - "next_task": "Wait for measured Durable Supervisor routing deficiencies before advancing a standalone routing policy.", - "evidence": ["https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/THEORY.md", "https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/SPEC.md"], - "completion_criteria": ["Publish route schema and executable policy evaluator.", "Measure strict and adaptive routing against real baselines.", "Replicate quality/cost claims and document unstable routes."], + "blockers": [], + "next_task": "None; historical source is retired from workload selection.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/THEORY.md", + "https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/SPEC.md" + ], + "completion_criteria": [ + "No current completion gate: retired source; highest evidenced maturity and original criteria preserved in historical_evidence." + ], "completion_evidence": [], - "completion_status": "INCOMPLETE", - "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "completion_status": "SUPERSEDED", + "program_state": "historical", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "CONSOLIDATED", + "consolidated_into": "gearbox", + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "agent-resource-claims", "github_url": "https://github.com/opsle/agent-resource-claims", "default_branch": "main", - "last_verified_head_sha": "dfe0fbc90c67ce5ef4256354bb62f1d511b1304c", + "last_verified_head_sha": "f80f82e79ed07915feaea6a3b873d9af4aed144c", "project_type": "concept", "purpose": "Coordinate atomic resource claim sets with leases and fencing so stale actors lose authority.", "lifecycle_stage": "THEORY", - "implementation_status": "none; placeholder source directory only", + "implementation_status": "Historical source preserved; concept contracts consolidated into tasks.", "implementation_requirement": "A reference claim-set evaluator with lease/fence state machine is required.", - "specification_status": "experimental theory contract; resource identity and conflict semantics remain incomplete", - "test_status": "placeholder only; no automated tests", - "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", - "measured_experiment_status": "none", - "reproducibility_status": "not available", - "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", - "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Deadlock analysis, semantic-region adapters, fairness benchmarks, and interoperability are missing."], + "specification_status": "Historical observation before retirement: experimental theory contract; resource identity and conflict semantics remain incomplete", + "test_status": "Historical observation before retirement: placeholder only; no automated tests", + "benchmark_status": "Historical observation before retirement: prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "Historical observation before retirement: none", + "reproducibility_status": "Historical observation before retirement: not available", + "documentation_status": "Historical observation before retirement: public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "Historical observation before retirement: GitHub documentation only; no evidence-backed site publication", + "known_limitations": [ + "Deadlock analysis, semantic-region adapters, fairness benchmarks, and interoperability are missing." + ], "dependencies": [], - "dependents": ["agent-execution-authorization", "agent-scheduler-runtime", "ephemeral-agent-workers", "semantic-edit-protocol"], + "dependents": [], "active_experiment_ids": [], - "blockers": ["No portable resource catalog, claim-set state machine, concurrency tests, or fairness fixtures."], - "next_task": "Wait for a workload-exposed resource-claim or fencing deficiency not already covered by Durable Supervisor's local claims.", - "evidence": ["https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/THEORY.md", "https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/ARCHITECTURE.md"], - "completion_criteria": ["Implement atomic claims, expiry, renewal, takeover, and fencing.", "Pass deadlock, crash, stale-actor, and fairness campaigns.", "Replicate across at least two store/runtime adapters."], + "blockers": [], + "next_task": "None; historical source is retired from workload selection.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/THEORY.md", + "https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/ARCHITECTURE.md" + ], + "completion_criteria": [ + "No current completion gate: retired source; highest evidenced maturity and original criteria preserved in historical_evidence." + ], "completion_evidence": [], - "completion_status": "INCOMPLETE", - "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "completion_status": "SUPERSEDED", + "program_state": "historical", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "CONSOLIDATED", + "consolidated_into": "tasks", + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "agent-discovery-control", "github_url": "https://github.com/opsle/agent-discovery-control", "default_branch": "main", - "last_verified_head_sha": "926547dd9fd713990b1d6f1f2e650aa6c0883564", + "last_verified_head_sha": "09c11f8a8ffd57521110a7696ab4f6a5181c0218", "project_type": "concept", "purpose": "Admit, merge, ignore, or escalate agent-discovered work under provenance and storm budgets.", "lifecycle_stage": "THEORY", - "implementation_status": "none; placeholder source directory only", + "implementation_status": "Historical source preserved; concept contracts consolidated into tasks.", "implementation_requirement": "An executable admission policy and convergence simulator are sufficient.", - "specification_status": "experimental theory contract; proposal similarity and budget semantics remain incomplete", - "test_status": "placeholder only; no automated tests", - "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", - "measured_experiment_status": "none", - "reproducibility_status": "not available", - "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", - "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Semantic duplicate benchmarks, multi-agent convergence, calibrated confidence, and complexity budgets are missing."], - "dependencies": ["agent-state-ledger"], + "specification_status": "Historical observation before retirement: experimental theory contract; proposal similarity and budget semantics remain incomplete", + "test_status": "Historical observation before retirement: placeholder only; no automated tests", + "benchmark_status": "Historical observation before retirement: prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "Historical observation before retirement: none", + "reproducibility_status": "Historical observation before retirement: not available", + "documentation_status": "Historical observation before retirement: public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "Historical observation before retirement: GitHub documentation only; no evidence-backed site publication", + "known_limitations": [ + "Semantic duplicate benchmarks, multi-agent convergence, calibrated confidence, and complexity budgets are missing." + ], + "dependencies": [], "dependents": [], "active_experiment_ids": [], - "blockers": ["No durable proposal dataset, similarity oracle, policy evaluator, or convergence harness."], - "next_task": "Wait for real duplicate-discovery or already-satisfied work evidence before advancing this repository.", - "evidence": ["https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/THEORY.md", "https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/BENCHMARK.md"], - "completion_criteria": ["Implement deterministic proposal admission and convergence.", "Benchmark duplicates, storms, and already-satisfied proofs.", "Replicate calibrated policies across objective sizes."], + "blockers": [], + "next_task": "None; historical source is retired from workload selection.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/THEORY.md", + "https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/BENCHMARK.md" + ], + "completion_criteria": [ + "No current completion gate: retired source; highest evidenced maturity and original criteria preserved in historical_evidence." + ], "completion_evidence": [], - "completion_status": "INCOMPLETE", - "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "completion_status": "SUPERSEDED", + "program_state": "historical", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "CONSOLIDATED", + "consolidated_into": "tasks", + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "agent-execution-authorization", "github_url": "https://github.com/opsle/agent-execution-authorization", "default_branch": "main", - "last_verified_head_sha": "8fb02e943c83fae121c63185d2f0d0dde8c4260a", + "last_verified_head_sha": "6babce3929e0284cdbb530ed5ebfb29a36388bf7", "project_type": "concept", "purpose": "Bind execution authority to exact jobs, leases, targets, resources, evidence, routes, and purposes.", "lifecycle_stage": "THEORY", - "implementation_status": "none; placeholder source directory only", + "implementation_status": "Historical source preserved; concept contracts consolidated into tasks.", "implementation_requirement": "As an authorization abstraction, an executable grant validator and revocation conformance suite are sufficient.", - "specification_status": "experimental theory contract; portable grant and composition semantics remain incomplete", - "test_status": "placeholder only; no automated tests", - "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", - "measured_experiment_status": "none", - "reproducibility_status": "not available", - "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", - "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Portable grant format, revocation latency, cross-orchestrator subjects, and authority composition are missing."], - "dependencies": ["agent-resource-claims", "agent-state-ledger"], - "dependents": ["controlled-agent-acceptance", "ephemeral-agent-workers", "gearbox"], + "specification_status": "Historical observation before retirement: experimental theory contract; portable grant and composition semantics remain incomplete", + "test_status": "Historical observation before retirement: placeholder only; no automated tests", + "benchmark_status": "Historical observation before retirement: prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "Historical observation before retirement: none", + "reproducibility_status": "Historical observation before retirement: not available", + "documentation_status": "Historical observation before retirement: public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "Historical observation before retirement: GitHub documentation only; no evidence-backed site publication", + "known_limitations": [ + "Portable grant format, revocation latency, cross-orchestrator subjects, and authority composition are missing." + ], + "dependencies": [], + "dependents": [], "active_experiment_ids": [], - "blockers": ["No grant schema, current-authority oracle, revocation model, or adversarial fixtures."], - "next_task": "Wait for a real execution-authorization gap that blocks the current Durable Supervisor or Opsle Tasks objective.", - "evidence": ["https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/THEORY.md", "https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/SPEC.md"], - "completion_criteria": ["Publish a portable grant schema and executable validator.", "Pass stale, revoked, drifted, replayed, and composed-authority cases.", "Measure and replicate revocation and interoperability behavior."], + "blockers": [], + "next_task": "None; historical source is retired from workload selection.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/THEORY.md", + "https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/SPEC.md" + ], + "completion_criteria": [ + "No current completion gate: retired source; highest evidenced maturity and original criteria preserved in historical_evidence." + ], "completion_evidence": [], - "completion_status": "INCOMPLETE", - "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "completion_status": "SUPERSEDED", + "program_state": "historical", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "CONSOLIDATED", + "consolidated_into": "tasks", + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "controlled-agent-acceptance", "github_url": "https://github.com/opsle/controlled-agent-acceptance", "default_branch": "main", - "last_verified_head_sha": "2d652adf56e53953327d09b1ba9c4a9c3445f052", + "last_verified_head_sha": "258325416c852be719c545a622e8a9d175447512", "project_type": "concept", "purpose": "Run exact one-shot autonomy acceptance under immutable authority, provider budgets, pause restoration, and evidence retention.", "lifecycle_stage": "THEORY", - "implementation_status": "none; placeholder source directory only", + "implementation_status": "Historical source preserved; concept contracts consolidated into tasks.", "implementation_requirement": "A provider-independent controller simulator and manifest validator must precede any real-provider harness.", - "specification_status": "experimental theory contract; manifest and interruption semantics remain incomplete", - "test_status": "placeholder only; no automated tests", - "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", - "measured_experiment_status": "none in this repository", - "reproducibility_status": "not available", - "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", - "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Independent harness, interruption model, non-reusable grants, and cross-runtime replication are missing."], - "dependencies": ["agent-execution-authorization", "agent-routing-policy", "agent-scheduler-runtime", "verifiable-agent-handoff"], + "specification_status": "Historical observation before retirement: experimental theory contract; manifest and interruption semantics remain incomplete", + "test_status": "Historical observation before retirement: placeholder only; no automated tests", + "benchmark_status": "Historical observation before retirement: prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "Historical observation before retirement: none in this repository", + "reproducibility_status": "Historical observation before retirement: not available", + "documentation_status": "Historical observation before retirement: public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "Historical observation before retirement: GitHub documentation only; no evidence-backed site publication", + "known_limitations": [ + "Independent harness, interruption model, non-reusable grants, and cross-runtime replication are missing." + ], + "dependencies": [], "dependents": [], "active_experiment_ids": [], - "blockers": ["Authorization, routing, scheduling, and handoff contracts are not yet executable together."], - "next_task": "Wait for real acceptance evidence to show a missing portable capability before advancing the standalone concept.", - "evidence": ["https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/THEORY.md", "https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/ARCHITECTURE.md"], - "completion_criteria": ["Implement and verify a one-shot fail-closed acceptance controller.", "Run bounded acceptance experiments only under separately authorized exact manifests.", "Replicate interruption, restoration, and residue reconciliation behavior."], + "blockers": [], + "next_task": "None; historical source is retired from workload selection.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/THEORY.md", + "https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/ARCHITECTURE.md" + ], + "completion_criteria": [ + "No current completion gate: retired source; highest evidenced maturity and original criteria preserved in historical_evidence." + ], "completion_evidence": [], - "completion_status": "INCOMPLETE", - "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "completion_status": "SUPERSEDED", + "program_state": "historical", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "CONSOLIDATED", + "consolidated_into": "tasks", + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "agent-recovery-policy", "github_url": "https://github.com/opsle/agent-recovery-policy", "default_branch": "main", - "last_verified_head_sha": "1b733a111e26e0a409fee3b96f627048531daefe", + "last_verified_head_sha": "a529954c784a7763fec452c23e143e96ccb97a72", "project_type": "concept", "purpose": "Choose bounded retries, consultations, alternate routes, fallbacks, or terminal outcomes from typed failure evidence.", "lifecycle_stage": "THEORY", - "implementation_status": "none; placeholder source directory only", + "implementation_status": "Historical source preserved; concept contracts consolidated into tasks.", "implementation_requirement": "An executable deterministic policy evaluator and failure-fixture suite are sufficient.", - "specification_status": "experimental theory contract; failure taxonomy and information-gain semantics remain incomplete", - "test_status": "placeholder only; no automated tests", - "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", - "measured_experiment_status": "none", - "reproducibility_status": "not available", - "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", - "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Outcome calibration, cross-provider taxonomy, information-gain estimation, and comparative experiments are missing."], - "dependencies": ["agent-routing-policy", "agent-state-ledger", "decision-evidence-protocol"], + "specification_status": "Historical observation before retirement: experimental theory contract; failure taxonomy and information-gain semantics remain incomplete", + "test_status": "Historical observation before retirement: placeholder only; no automated tests", + "benchmark_status": "Historical observation before retirement: prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "Historical observation before retirement: none", + "reproducibility_status": "Historical observation before retirement: not available", + "documentation_status": "Historical observation before retirement: public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "Historical observation before retirement: GitHub documentation only; no evidence-backed site publication", + "known_limitations": [ + "Outcome calibration, cross-provider taxonomy, information-gain estimation, and comparative experiments are missing." + ], + "dependencies": [], "dependents": [], "active_experiment_ids": [], - "blockers": ["No shared failure schema, attempt ledger, route evaluator, or comparative fixture set."], - "next_task": "Wait for repeated real recovery failures to demonstrate a policy deficiency; do not create speculative recovery work.", - "evidence": ["https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md", "https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/BENCHMARK.md"], - "completion_criteria": ["Publish a typed failure schema and executable bounded policy.", "Compare retry, consultation, alternate route, and terminal baselines.", "Replicate outcome and convergence claims across failure classes."], + "blockers": [], + "next_task": "None; historical source is retired from workload selection.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md", + "https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/BENCHMARK.md" + ], + "completion_criteria": [ + "No current completion gate: retired source; highest evidenced maturity and original criteria preserved in historical_evidence." + ], "completion_evidence": [], - "completion_status": "INCOMPLETE", - "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "completion_status": "SUPERSEDED", + "program_state": "historical", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "CONSOLIDATED", + "consolidated_into": "tasks", + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "ephemeral-agent-workers", "github_url": "https://github.com/opsle/ephemeral-agent-workers", "default_branch": "main", - "last_verified_head_sha": "ad96fcfdfac06d340b5e96d369634980cee78ef4", + "last_verified_head_sha": "e758697528e71290bec55391a6b4e6c5bae83422", "project_type": "concept", "purpose": "Execute in disposable bounded workers through a narrow broker, seal results, and prove destruction.", "lifecycle_stage": "THEORY", - "implementation_status": "none; placeholder source directory only", + "implementation_status": "Historical source preserved; concept contracts consolidated into tasks.", "implementation_requirement": "A portable worker/broker reference adapter and containment test harness are required.", - "specification_status": "experimental theory contract; broker and destruction-proof formats remain incomplete", - "test_status": "placeholder only; no automated tests", - "benchmark_status": "prose plan only; no runnable harness or frozen fixtures", - "measured_experiment_status": "none in this repository", - "reproducibility_status": "not available", - "documentation_status": "public theory, preliminary specification, architecture, and benchmark plan present", - "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["Portable isolation adapters, quantitative containment, kernel threat model, and destruction-proof interoperability are missing."], - "dependencies": ["agent-execution-authorization", "agent-resource-claims", "verifiable-agent-handoff"], + "specification_status": "Historical observation before retirement: experimental theory contract; broker and destruction-proof formats remain incomplete", + "test_status": "Historical observation before retirement: placeholder only; no automated tests", + "benchmark_status": "Historical observation before retirement: prose plan only; no runnable harness or frozen fixtures", + "measured_experiment_status": "Historical observation before retirement: none in this repository", + "reproducibility_status": "Historical observation before retirement: not available", + "documentation_status": "Historical observation before retirement: public theory, preliminary specification, architecture, and benchmark plan present", + "site_publication_status": "Historical observation before retirement: GitHub documentation only; no evidence-backed site publication", + "known_limitations": [ + "Portable isolation adapters, quantitative containment, kernel threat model, and destruction-proof interoperability are missing." + ], + "dependencies": [], "dependents": [], "active_experiment_ids": [], - "blockers": ["Portable authority, claim, and handoff contracts are not ready; no safe synthetic containment harness exists."], - "next_task": "Wait for a real workload to require an ephemeral-worker boundary before advancing this repository.", - "evidence": ["https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md", "https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/ARCHITECTURE.md"], - "completion_criteria": ["Implement a portable bounded broker/worker adapter and destruction proof.", "Pass containment, credential, process, network, seal, and cleanup campaigns.", "Replicate across isolation technologies with a published threat model."], + "blockers": [], + "next_task": "None; historical source is retired from workload selection.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md", + "https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/ARCHITECTURE.md" + ], + "completion_criteria": [ + "No current completion gate: retired source; highest evidenced maturity and original criteria preserved in historical_evidence." + ], "completion_evidence": [], - "completion_status": "INCOMPLETE", - "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "completion_status": "SUPERSEDED", + "program_state": "historical", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "CONSOLIDATED", + "consolidated_into": "tasks", + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "gearbox", "github_url": "https://github.com/opsle/gearbox", "default_branch": "main", - "last_verified_head_sha": "f3fab9f292cf4eabd7200615d444f98881f57d55", + "last_verified_head_sha": "12112c3a04d72b768f3252a6d0f6ea455816c9a1", "project_type": "concept", "purpose": "Let one powerful primary developer execute deterministic operations or one bounded cognitive assignment through an authority-, context-, result-, and budget-constrained transmission boundary.", "lifecycle_stage": "PROTOTYPED", - "implementation_status": "provider-free Python reference core with strict authority-policy admission, exact deterministic argv execution, content-addressed staged helper context, injected one-shot helper transport, passive process waiting, compact results, raw-artifact accounting, fail-closed budgets, and Visible Value receipts", + "implementation_status": "Provider-free bounded execution core and deterministic Tasks phase routing; Agent Routing Policy contracts consolidated with source provenance.", "implementation_requirement": "A runnable provider-free core plus deterministic and injected-helper conformance paths is sufficient for the prototype gate; production provider integration and comparative benefit require later evidence.", "specification_status": "versioned request, policy, result, context-selection, output-contract, budget, helper-transport, raw-artifact, cleanup, and Visible Value boundaries are public at the verified revision", - "test_status": "19 of 19 provider-free automated tests passed locally, in PR #1 CI, and in final-main CI; ruff, shellcheck, actionlint, gitleaks, wheel build, receipt validation, and public raw-locator/hash checks passed", + "test_status": "Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation.", "benchmark_status": "one revision-bound deterministic dogfood fixture with public compact result, raw artifacts, and value receipt plus exact-revision Decision Evidence validation and Trajectory Profiler ingestion; benchmark plan only, with no comparative baseline or model/provider subject", "measured_experiment_status": "none; zero model/provider experiment runs", "reproducibility_status": "tools/verify, package build, provider-free dogfood, receipt validation, and artifact hash checks are reproducible from the public revision; cognitive transport and comparative claims remain untested", "documentation_status": "canonical theory, normative draft specification, architecture, security boundary, known limitations, benchmark plan, provenance, usage, and exact release evidence are public", "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["No production helper transport is bundled; transport isolation is external, Context Firewall reduction is not integrated, Python is the only symbol selector, raw deterministic byte ceilings are post-process, distributed locking and source-write application are absent, and no controlled evidence establishes correctness preservation or intelligence, context, token, latency, cost, or provider-session savings."], - "dependencies": ["context-firewall", "decision-evidence-protocol", "agent-trajectory-profiler", "agent-routing-policy", "agent-execution-authorization"], - "dependents": [], + "known_limitations": [ + "No production helper transport is bundled; transport isolation is external, Context Firewall reduction is not integrated, Python is the only symbol selector, raw deterministic byte ceilings are post-process, distributed locking and source-write application are absent, and no controlled evidence establishes correctness preservation or intelligence, context, token, latency, cost, or provider-session savings.", + "Tasks capability gateway integration is operational evidence, not comparative savings or research completion." + ], + "dependencies": [], + "dependents": [ + "tasks" + ], "active_experiment_ids": [], - "blockers": ["A production-quality bounded helper transport, independently verified isolation and termination, full Context Firewall integration, and a frozen comparative benchmark remain missing."], - "next_task": "Support Durable Supervisor receipt visibility and Opsle Tasks workload measurement; advance the standalone Gearbox only if integrated evidence exposes a routing or preflight deficiency.", - "evidence": ["https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py", "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/tests/test_core.py", "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/evidence/release-001/verification.json", "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/PROVENANCE.md", "program/evidence/gearbox-publication/README.md", "https://github.com/opsle/gearbox/pull/1", "https://github.com/opsle/gearbox/actions/runs/33229696388"], - "completion_criteria": ["Publish and verify a production-quality bounded helper transport without absorbing durable orchestration.", "Benchmark deterministic and bounded-cognitive gears against real direct-execution baselines under equal correctness gates.", "Replicate bounded value claims and document transport, isolation, cleanup, escalation, and unsupported-task failure modes."], + "blockers": [ + "General helper transport, independent isolation evidence and comparative benchmark remain separate research work." + ], + "next_task": "Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository.", + "evidence": [ + "https://github.com/opsle/gearbox/blob/12112c3a04d72b768f3252a6d0f6ea455816c9a1/README.md", + "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py", + "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/tests/test_core.py", + "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/evidence/release-001/verification.json", + "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/PROVENANCE.md", + "program/evidence/gearbox-publication/README.md", + "https://github.com/opsle/gearbox/pull/1", + "https://github.com/opsle/gearbox/actions/runs/33229696388" + ], + "completion_criteria": [ + "Publish and verify a production-quality bounded helper transport without absorbing durable orchestration.", + "Benchmark deterministic and bounded-cognitive gears against real direct-execution baselines under equal correctness gates.", + "Replicate bounded value claims and document transport, isolation, cleanup, escalation, and unsupported-task failure modes." + ], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "ACTIVE", + "consolidated_into": null, + "workload_eligible": true, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "affected-verification", "github_url": "https://github.com/opsle/affected-verification", "default_branch": "main", - "last_verified_head_sha": "97f490a67337552fee25757266f3dc034660dca0", + "last_verified_head_sha": "792c4bb7881f6f430b2b1ba238ba05be43f50d38", "project_type": "concept", "purpose": "Select the smallest verification workload whose sufficiency can be defended from available change-impact, dependency, coverage, policy, and risk evidence.", "lifecycle_stage": "VERIFIED", - "implementation_status": "dependency-free Node.js deterministic plan-v2 planner with check-level dependency-completeness states, mechanism and boundary evidence, check-local fail-closed forced selection, explainable skips, bounded deterministic Python boundary inspection, identity-bound SHADOW validation, frozen-oracle repair/replay harnesses, and opsle.value-receipt.v1 safety-cost telemetry", + "implementation_status": "dependency-free Node.js deterministic plan-v2 planner with check-level dependency-completeness states, mechanism and boundary evidence, check-local fail-closed forced selection, explainable skips, bounded deterministic Python boundary inspection, identity-bound SHADOW validation, frozen-oracle repair/replay harnesses, and opsle.value-receipt.v1 safety-cost telemetry; Tasks verification planning and capture contract integrated.", "implementation_requirement": "A runnable deterministic planner, input validator, plan contract, conformance fixtures, and shadow classifier are sufficient for the prototype gate; real adapters and comparative evidence are later gates.", "specification_status": "versioned input-v2 and opsle.affected-verification.plan.v2 contracts define check-level dependency mechanisms, completeness states, opaque-boundary provenance, forced selection, evidence coverage, policy matching, selection and skip reasons, deterministic identities, Visible Value safety additions, and SHADOW observations; plan v1 remains immutable historical evidence", - "test_status": "107 of 107 automated tests, 15 of 15 conformance scenarios, and 10 of 10 determinism checks passed locally, in PR #4 CI, in exact final-main CI, and from a fresh detached worktree at 97f490a67337552fee25757266f3dc034660dca0; the AV-EXP-003 verifier and full deterministic repair reproduction also passed with identical result and regression-matrix identities", + "test_status": "Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation.", "benchmark_status": "AV-EXP-003 preregistered and recorded an opaque-boundary SHADOW repair: the permanently preserved AV-EXP-002 AV2-006 miss is selected in both repaired AV arms for generalized subprocess/child-interpreter completeness evidence; ten adversarial cases have zero misses, 7 added checks, 13 checks still skipped, 10/10 targeted scenarios, and zero FULL escalations; frozen AV-EXP-001 adds 0 test executions and has 0 repaired misses, while AV-EXP-002 adds 6 AV_CORE and 7 AV_WITH_SELECTOR executions and has 0 repaired misses", "measured_experiment_status": "AV-EXP-001, AV-EXP-002, and AV-EXP-003 RECORDED; AV-EXP-002 permanently remains FAIL with one miss in each AV arm; AV-EXP-003 result sha256:03b2f7d6a380c84f6a1749531067cf8b87404c879f42380de8f07cce48251519 and regression matrix sha256:7260c2d3476a6e78323e75d36c54c8409ea4cb18fa3a8f9a76b5533e1df08615 measure repair selection and exact precision cost; zero provider/model runs", "reproducibility_status": "the public AV-EXP-003 harness deterministically replays frozen AV-EXP-001/002 evidence, reruns ten adversarial full-oracle SHADOW cases, inspects the pinned Click source, validates ten Visible Value receipts, and reproduces identical result and regression-matrix identities from a fresh exact-main worktree on this host; no independent qualifying replication exists", "documentation_status": "public canonical definition, normative specification, source-linked prior-art audit, architecture and independence boundaries, verification catalog and policy semantics, trust ramp, benchmark plan, limitations, usage, security, schema, and fixtures are present", "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["AV-EXP-002 permanently remains a FAIL. AV-EXP-003 repairs the known skip in frozen replay but its bounded static inspector does not trace child-process imports, close arbitrary plugin/reflection behavior, solve dynamic Python dependencies, provide a production adapter, add historical real-change evidence, establish independent replication, prove general safety or correctness equivalence, claim causal savings, or authorize production trust."], + "known_limitations": [ + "AV-EXP-002 permanently remains a FAIL. AV-EXP-003 repairs the known skip in frozen replay but its bounded static inspector does not trace child-process imports, close arbitrary plugin/reflection behavior, solve dynamic Python dependencies, provide a production adapter, add historical real-change evidence, establish independent replication, prove general safety or correctness equivalence, claim causal savings, or authorize production trust." + ], "dependencies": [], - "dependents": [], - "active_experiment_ids": ["AV-EXP-001", "AV-EXP-002", "AV-EXP-003"], - "blockers": ["The bounded repair does not close opaque boundaries or constitute dynamic analysis; historical real-change replay, a production-quality evidence adapter, and independent qualifying replication remain missing. AV remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized."], - "next_task": "Remain OBSERVE/SHADOW and wait for Durable Supervisor-driven Opsle Tasks work to expose a concrete verification-selection deficiency before adding another experiment.", - "evidence": ["https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/REPORT.md", "https://github.com/opsle/affected-verification/blob/7aa4d13e42d6a547973d7f2a6b330821145cedc2/benchmark/av-exp-003/preregistration-v1/preregistration.json", "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/summary.json", "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/repair-regression-matrix.json", "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/evidence-manifest.json", "https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/REPORT.md", "https://github.com/opsle/affected-verification/pull/4", "https://github.com/opsle/affected-verification/actions/runs/33412942072"], - "completion_criteria": ["Publish production evidence adapters and independently validate their completeness boundaries.", "Run correctness-first comparisons against full verification and native selectors with durable shadow miss evidence.", "Replicate bounded workload and safety claims on independent repositories or environments."], + "dependents": [ + "tasks" + ], + "active_experiment_ids": [ + "AV-EXP-001", + "AV-EXP-002", + "AV-EXP-003" + ], + "blockers": [ + "Research remains OBSERVE/SHADOW; AV-EXP-002 failure is permanent. Tasks bounded manifest-backed selection requires complete impact, catalog and check-boundary evidence; incomplete evidence broadens or stops. No general TRUSTED_BOUNDED promotion." + ], + "next_task": "Keep OBSERVE/SHADOW research limits; admit a new calibration or repair only from a demonstrated Tasks verification deficiency.", + "evidence": [ + "https://github.com/opsle/affected-verification/blob/792c4bb7881f6f430b2b1ba238ba05be43f50d38/README.md", + "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/REPORT.md", + "https://github.com/opsle/affected-verification/blob/7aa4d13e42d6a547973d7f2a6b330821145cedc2/benchmark/av-exp-003/preregistration-v1/preregistration.json", + "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/summary.json", + "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/repair-regression-matrix.json", + "https://github.com/opsle/affected-verification/blob/97f490a67337552fee25757266f3dc034660dca0/benchmark/av-exp-003/results-v1/evidence-manifest.json", + "https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/REPORT.md", + "https://github.com/opsle/affected-verification/pull/4", + "https://github.com/opsle/affected-verification/actions/runs/33412942072" + ], + "completion_criteria": [ + "Publish production evidence adapters and independently validate their completeness boundaries.", + "Run correctness-first comparisons against full verification and native selectors with durable shadow miss evidence.", + "Replicate bounded workload and safety claims on independent repositories or environments." + ], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "ACTIVE", + "consolidated_into": null, + "workload_eligible": true, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "research", "github_url": "https://github.com/opsle/research", "default_branch": "main", - "last_verified_head_sha": "e8a2c36678ac4d4f72a79a56e03e9f896811be02", + "last_verified_head_sha": "840c79c04199762acdd59e89e807ca9086f3e45e", "project_type": "program infrastructure", "purpose": "Own public research methodology, portfolio reconciliation, experiment records, and the authoritative program ledger.", "lifecycle_stage": "PROTOTYPED", - "implementation_status": "authoritative 21-repository portfolio and priority ledger, machine-readable 18-concept theory registry, five-experiment registry, lifecycle and anti-nitpick controls, normative Visible Value target, deterministic generated status and priority views, provider-free EXP-001 preparation, and recorded AV-EXP-001/002/003 shadow evidence with integrity CI", + "implementation_status": "Authoritative current/historical research ledger, concept mappings, experiment records and deterministic generated views; Tasks owns workload management.", "implementation_requirement": "Program infrastructure requires a validated registry, generated dashboard, experiment ledger, operating rules, and CI.", "specification_status": "canonical lifecycle, theory classifications and dispositions, Gearbox and Context Firewall definitions, Gearbox-versus-Durable boundary, Visible Value receipt and measurement classes, operator/model channels, observational corpus, shadow/replay attachment, NOW/NEXT/THEN/LATER/PARKED lanes, work-item admission, registry, experiment, and deterministic generated-view controls are specified", - "test_status": "90 of 90 repository tests passed locally at the verified default-branch HEAD; the 21-repository and 18-concept registries, five experiment records, generated dashboard, and deterministic experiment evidence checks passed", + "test_status": "Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation.", "benchmark_status": "EXP-001 provider-free components remain frozen and unconsumed; AV-EXP-001/002/003 are recorded correctness-first shadow calibrations, including the preserved AV-EXP-002 miss and bounded AV-EXP-003 repair evidence; no provider/model subject ran", "measured_experiment_status": "AV-EXP-001, AV-EXP-002, and AV-EXP-003 are recorded shadow experiments; EXP-001 remains planned and unconsumed; no provider/model experiment run exists", "reproducibility_status": "program and receipt validation, both generated views, public dogfood artifacts, exact-revision EXP-001 provider-free qualification, and AV-EXP-001/002/003 deterministic shadow evidence are reproducible; full secret-backed coordinator replay still requires the external seed, and no provider/model subject result exists", - "documentation_status": "public canonical theory map, project reconciliation, generated priority and portfolio views, registered Gearbox boundary and prototype scope, Context Firewall adapter scope, consolidation provenance policy, program documentation, Visible Value semantics, machine controls, and evidence are present", + "documentation_status": "Current research ledger, lifecycle policy and generated concept, status and priority views; superseded pre-consolidation analysis explicitly historical.", "site_publication_status": "GitHub research hub only", - "known_limitations": ["The registry cannot self-reference the commit that contains its own SHA, default-branch HEAD verification remains operator-driven, EXP-001 is still unconsumed, AV remains OBSERVE/SHADOW, the site is not registry-derived, and integrated Durable Supervisor/Opsle Tasks token, cost, first-pass, retry, and avoidable-intelligence measurements do not yet exist."], + "known_limitations": [ + "Recorded default-branch revisions are observations, not promises about later changes. EXP-001 remains unconsumed; AV remains OBSERVE/SHADOW. Tasks telemetry cannot establish controlled comparative benefit." + ], "dependencies": [], - "dependents": [".github", "site"], - "active_experiment_ids": ["EXP-001", "AV-EXP-001", "AV-EXP-002", "AV-EXP-003", "LEGACY-001"], - "blockers": ["Durable Supervisor v0.1 measurement and foreign-workload stopping criteria remain open; Opsle Tasks cannot become the primary workload until v0.1 is declared and frozen."], - "next_task": "Keep the authoritative priority and portfolio views current while Durable Supervisor completes its ten bounded v0.1 stopping criteria.", - "evidence": ["program/registry.json", "PROGRAM_STATUS.md", "program/PRIORITY.md", "program/THEORY_MAP.md", "program/theory-registry.json", "program/evidence/gearbox-publication/README.md", "program/evidence/exp-001-offline-freeze/README.md", "program/evidence/exp-001-preregistration/README.md", "program/evidence/exp-001-block-coordinator/README.md", "program/evidence/exp-001-live-preflight/README.md", "https://github.com/opsle/research/pull/19", "https://github.com/opsle/research/pull/20", "https://github.com/opsle/research/pull/21", "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/experiments/exp-001/benchmark.json", "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/experiments/exp-001/preregistration-v1/preregistration.json", "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/experiments/exp-001/coordinator-v1/coordinator.py", "https://github.com/opsle/research/blob/72f9e4a0326d68d3870e2e79ce4e351acb1d8ffa/program/VISIBLE_VALUE_CONTRACT.md"], - "completion_criteria": ["Keep the 21-repository ledger and dashboard mechanically consistent.", "Retain immutable experiment evidence and lifecycle promotion proof.", "Publish program documentation without stale or unsupported claims."], + "dependents": [], + "active_experiment_ids": [ + "EXP-001", + "AV-EXP-001", + "AV-EXP-002", + "AV-EXP-003", + "LEGACY-001" + ], + "blockers": [ + "Research completion requires controlled evidence and replication; integration completion is insufficient." + ], + "next_task": "Select the next authorized task from Opsle Tasks; use its project scope and workload evidence to identify the smallest justified change. This ledger does not create a parallel queue.", + "evidence": [ + "https://github.com/opsle/research/blob/840c79c04199762acdd59e89e807ca9086f3e45e/README.md", + "program/registry.json", + "PROGRAM_STATUS.md", + "program/PRIORITY.md", + "program/THEORY_MAP.md", + "program/theory-registry.json", + "program/evidence/gearbox-publication/README.md", + "program/evidence/exp-001-offline-freeze/README.md", + "program/evidence/exp-001-preregistration/README.md", + "program/evidence/exp-001-block-coordinator/README.md", + "program/evidence/exp-001-live-preflight/README.md", + "https://github.com/opsle/research/pull/19", + "https://github.com/opsle/research/pull/20", + "https://github.com/opsle/research/pull/21", + "https://github.com/opsle/research/blob/04234a65bf36192d63f1dd173c440d45a6604d2b/experiments/exp-001/benchmark.json", + "https://github.com/opsle/research/blob/31848c3f25ff9371055932657e8e2f8ad54cc8c7/experiments/exp-001/preregistration-v1/preregistration.json", + "https://github.com/opsle/research/blob/9ee43197880c18d4e185cf7e29e02a151d22a12e/experiments/exp-001/coordinator-v1/coordinator.py", + "https://github.com/opsle/research/blob/72f9e4a0326d68d3870e2e79ce4e351acb1d8ffa/program/VISIBLE_VALUE_CONTRACT.md" + ], + "completion_criteria": [ + "Satisfy applicable canonical lifecycle and Visible Value gates with revision-linked evidence; integration completion alone is insufficient." + ], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "active", - "last_verified_at": "2026-09-05T15:28:58Z" + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "ACTIVE", + "consolidated_into": null, + "workload_eligible": true, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": "site", "github_url": "https://github.com/opsle/site", "default_branch": "main", - "last_verified_head_sha": "28ad65be4750dc849976fbf5c9eae9501c6bbb25", + "last_verified_head_sha": "29c94e6440ec91f5e09aded959fe7144187af6d7", "project_type": "program infrastructure", "purpose": "Provide source for the future public research, benchmark, failure, documentation, and product-gateway site.", "lifecycle_stage": "PROTOTYPED", "implementation_status": "React/Vinext source implementation with content routes", "implementation_requirement": "Program infrastructure requires tested source, registry-derived content, accessibility, and separately authorized publication proof.", "specification_status": "README and route structure define the current source scope", - "test_status": "automated build/render tests present; not rerun because this reconciliation kept other repositories read-only", + "test_status": "Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation.", "benchmark_status": "not applicable to the source shell; evidence-backed benchmark content is absent", "measured_experiment_status": "none", "reproducibility_status": "build instructions present; current test result unverified in this run", "documentation_status": "public source README and content routes present", "site_publication_status": "source only; explicitly not deployed", - "known_limitations": ["Content is not registry-generated, evidence-backed research results are absent, and deployment is intentionally unauthorized."], - "dependencies": ["research"], + "known_limitations": [ + "Content is not registry-generated, evidence-backed research results are absent, and deployment is intentionally unauthorized." + ], + "dependencies": [], "dependents": [], "active_experiment_ids": [], - "blockers": ["Wait for validated registry data and measured research; deployment requires separate authorization."], - "next_task": "Remain later until measured research and separate site-release authorization justify registry-derived public content.", - "evidence": ["https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md", "https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/tests/rendered-html.test.mjs"], - "completion_criteria": ["Render registry-derived public state without unsupported claims.", "Pass build, render, accessibility, and content-consistency checks.", "Publish only under a separate authorized release with exact revision evidence."], + "blockers": [ + "Wait for validated registry data and measured research; deployment requires separate authorization." + ], + "next_task": "Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository.", + "evidence": [ + "https://github.com/opsle/site/blob/29c94e6440ec91f5e09aded959fe7144187af6d7/README.md", + "https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md", + "https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/tests/rendered-html.test.mjs" + ], + "completion_criteria": [ + "Render registry-derived public state without unsupported claims.", + "Pass build, render, accessibility, and content-consistency checks.", + "Publish only under a separate authorized release with exact revision evidence." + ], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "ACTIVE", + "consolidated_into": null, + "workload_eligible": true, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] }, { "name": ".github", @@ -743,24 +1221,230 @@ "implementation_status": "documentation-only organization profile", "implementation_requirement": "Program infrastructure may remain documentation-only but must be mechanically checked against the authoritative registry before completion.", "specification_status": "profile purpose is documented; no synchronization contract", - "test_status": "not applicable to current single Markdown profile; consistency is unverified", + "test_status": "Historical verification remains pinned in historical_evidence; current default-branch documentation and source inspected. No new test execution is claimed by this reconciliation.", "benchmark_status": "not applicable", "measured_experiment_status": "none", "reproducibility_status": "not applicable", "documentation_status": "public organization profile present and lists all 16 concepts", "site_publication_status": "published as the GitHub organization profile", - "known_limitations": ["Repository list is manually maintained and can drift from the authoritative registry."], - "dependencies": ["research"], + "known_limitations": [ + "Repository list is manually maintained and can drift from the authoritative registry." + ], + "dependencies": [], "dependents": [], "active_experiment_ids": [], - "blockers": ["No mechanical registry consistency check exists in this repository."], - "next_task": "Remain parked until a broken organization-profile link or external release condition creates a concrete need.", - "evidence": ["https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md"], - "completion_criteria": ["Keep organization identity and repository links consistent with the authoritative registry.", "Automate or verify consistency at exact revisions.", "Document ownership and update workflow."], + "blockers": [ + "No mechanical registry consistency check exists in this repository." + ], + "next_task": "Use Opsle Tasks to admit work only for a demonstrated deficiency in this repository.", + "evidence": [ + "https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/README.md", + "https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md" + ], + "completion_criteria": [ + "Keep organization identity and repository links consistent with the authoritative registry.", + "Automate or verify consistency at exact revisions.", + "Document ownership and update workflow." + ], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "waiting", - "last_verified_at": "2026-09-05T15:28:58Z" + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "ACTIVE", + "consolidated_into": null, + "workload_eligible": false, + "historical_evidence": "program/history/pre-consolidation/registry.json", + "historical_experiment_ids": [] + }, + { + "name": "tasks", + "github_url": "https://github.com/opsle/tasks", + "default_branch": "main", + "last_verified_head_sha": "e1207c5264c59e14efe9838bba3a33ba504665d2", + "project_type": "program infrastructure", + "purpose": "Current workload lifecycle, task management and remote execution substrate with plug-and-play capability contract.", + "lifecycle_stage": "PROTOTYPED", + "implementation_status": "Current workload lifecycle, task management and remote execution substrate with plug-and-play capability contract.", + "implementation_requirement": "Executable implementation and meaningful boundary tests.", + "specification_status": "Versioned public contract and implementation at recorded revision.", + "test_status": "Implementation and boundary-test sources inspected at recorded HEAD; upstream proof receipts are observational evidence. Final research verification belongs to the pipeline.", + "benchmark_status": "No controlled comparative benefit established.", + "measured_experiment_status": "Integration observations only; no research completion inferred.", + "reproducibility_status": "Pinned source and documented commands; independent research replication not established.", + "documentation_status": "Current README and versioned contract inspected.", + "site_publication_status": "Separate publication requirements apply.", + "known_limitations": [ + "Integration completion does not establish causal savings, comparative correctness or research completion." + ], + "dependencies": [ + "gearbox", + "context-firewall", + "affected-verification", + "visible-value" + ], + "dependents": [], + "active_experiment_ids": [], + "blockers": [ + "Controlled comparative evidence and independent replication remain absent." + ], + "next_task": "Use Tasks workload evidence to admit only the smallest justified deficiency repair.", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/README.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/CAPABILITIES.md" + ], + "completion_criteria": [ + "Satisfy applicable canonical lifecycle and Visible Value gates with revision-linked evidence; integration completion alone is insufficient." + ], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "active", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "ACTIVE", + "consolidated_into": null, + "workload_eligible": true, + "historical_evidence": null, + "historical_experiment_ids": [] + }, + { + "name": "visible-value", + "github_url": "https://github.com/opsle/visible-value", + "default_branch": "main", + "last_verified_head_sha": "f011395cae86a7a736b424d658bbe3afe65ba870", + "project_type": "concept", + "purpose": "Independent receipt validation and operator summaries preserving evidence classes and missing measurements.", + "lifecycle_stage": "PROTOTYPED", + "implementation_status": "Independent receipt validation and operator summaries preserving evidence classes and missing measurements.", + "implementation_requirement": "Executable implementation and meaningful boundary tests.", + "specification_status": "Versioned public contract and implementation at recorded revision.", + "test_status": "Implementation and boundary-test sources inspected at recorded HEAD; upstream proof receipts are observational evidence. Final research verification belongs to the pipeline.", + "benchmark_status": "No controlled comparative benefit established.", + "measured_experiment_status": "Integration observations only; no research completion inferred.", + "reproducibility_status": "Pinned source and documented commands; independent research replication not established.", + "documentation_status": "Current README and versioned contract inspected.", + "site_publication_status": "Separate publication requirements apply.", + "known_limitations": [ + "Integration completion does not establish causal savings, comparative correctness or research completion.", + "Validates producer-supplied claims; does not authenticate external artifacts. Only compatible EXACT/OBSERVED numeric SUM values aggregate; missing values remain absent." + ], + "dependencies": [], + "dependents": [ + "tasks" + ], + "active_experiment_ids": [], + "blockers": [ + "Controlled comparative evidence and independent replication remain absent." + ], + "next_task": "Use Tasks workload evidence to admit only the smallest justified deficiency repair.", + "evidence": [ + "https://github.com/opsle/visible-value/blob/f011395cae86a7a736b424d658bbe3afe65ba870/README.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/CAPABILITIES.md" + ], + "completion_criteria": [ + "Satisfy applicable canonical lifecycle and Visible Value gates with revision-linked evidence; integration completion alone is insufficient." + ], + "completion_evidence": [], + "completion_status": "INCOMPLETE", + "program_state": "waiting", + "last_verified_at": "2026-09-09T00:00:00Z", + "repository_disposition": "ACTIVE", + "consolidated_into": null, + "workload_eligible": true, + "historical_evidence": null, + "historical_experiment_ids": [] } - ] + ], + "membership": { + "current": [ + "agent-trajectory-profiler", + "semantic-edit-protocol", + "context-firewall", + "decision-evidence-protocol", + "gearbox", + "affected-verification", + "research", + "site", + ".github", + "tasks", + "visible-value" + ], + "historical": [ + "agent-discovery-control", + "agent-execution-authorization", + "agent-recovery-policy", + "agent-resource-claims", + "agent-routing-policy", + "agent-scheduler-runtime", + "agent-state-ledger", + "controlled-agent-acceptance", + "ephemeral-agent-workers", + "event-driven-agent-wakeup", + "verifiable-agent-handoff", + "durable-supervisor" + ] + }, + "repository_counts": { + "current": 11, + "historical": 12, + "workload_eligible": 10 + }, + "inventory_evidence": "program/evidence/post-consolidation/inventory.json", + "historical_state": "program/history/pre-consolidation/registry.json", + "completed_work": [ + { + "id": "taslos-tasks-retirement", + "status": "COMPLETED", + "evidence": [ + "https://github.com/opsle/tasks/commit/9915dd4a7f0a4bf73c62b2310523b1429f46d9b1" + ], + "scope": "Retired execution target; historical extraction provenance retained." + }, + { + "id": "paperclip-retirement", + "status": "COMPLETED", + "evidence": [ + "https://github.com/opsle/tasks/commit/2d06527f583e49dbe75adb84051806812712e920" + ], + "scope": "Retired orchestration references removed; no current dependency." + }, + { + "id": "repository-consolidation", + "status": "COMPLETED", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.md" + ], + "scope": "Final 22:42 UTC receipt supersedes partial observations, including Durable Supervisor host move. Source licenses, bundles, history and project records preserved." + }, + { + "id": "remote-incus-execution", + "status": "COMPLETED", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/REMOTE_EXECUTION_PROOF_2026-09-07.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/remote-execution/no-local-fallback.json" + ], + "scope": "Recorded normal Tasks Codex and Claude SSH executions to Incus targets completed; unreachable target did not fall back locally. Integration observations, not controlled savings evidence." + }, + { + "id": "task-15-capability-contract", + "status": "COMPLETED", + "evidence": [ + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/TASK_15.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/CAPABILITIES.md" + ], + "scope": "Versioned manifest discovery, immutable selection, operator grants, lifecycle hooks, receipts and scoped model gateway implemented. Historical agent-run remains retired. Implementation completion does not imply deployment or research completion." + } + ], + "optional_tools": { + "graphify": { + "role": "OPTIONAL_STANDALONE_CLI", + "tasks_capability": false, + "basis": "Task 21 approved scope fixes the final policy; the historical external-capability acceptance in Tasks commit 00d2bac is superseded. Current bundled manifests contain only Gearbox, Context Firewall, Affected Verification and Visible Value." + } + }, + "relationship_scope": "Current repository dependencies record integrated public-interface consumers; historical theoretical dependency edges remain in historical_state. Capability installation and project grants are separate operator state.", + "inventory_evidence_sha256": "sha256:35da0023c7e232f55ec21fbb6951031c0891421015d960e828ff8f08c554c63d", + "consolidation_evidence": { + "path": "program/evidence/post-consolidation/consolidation.json", + "sha256": "sha256:01fb77a6227e7b6f1aee8ee56ff15672a7d0657611371d0f0f63597268f945ce" + } } diff --git a/program/theory-registry.json b/program/theory-registry.json index dd58e53..b0cd4e2 100644 --- a/program/theory-registry.json +++ b/program/theory-registry.json @@ -1,9 +1,9 @@ { - "schema_version": 1, - "registry_id": "opsle.theory-registry.v1", - "verified_at": "2026-08-31T02:59:01Z", - "source_repository_count": 21, - "current_concept_repository_count": 18, + "schema_version": 2, + "registry_id": "opsle.theory-registry.v2", + "verified_at": "2026-09-09T00:00:00Z", + "source_repository_count": 23, + "current_concept_repository_count": 8, "canonical_definitions": { "gearbox": "Agent Gearbox lets a powerful primary developer delegate routine operations and bounded work to deterministic software or less expensive models, then receive only the compact result needed to continue.", "context_firewall": "Context Firewall is a deterministic boundary that keeps operational noise out of an AI agent's context while preserving the compact evidence, provenance, and escalation path the agent needs to make correct decisions.", @@ -42,8 +42,8 @@ ], "consumers": [], "current_implementation_fidelity": { - "status": "NARROW_PROTOTYPE", - "assessment": "The public provider-free core implements strict request and policy admission, exact deterministic execution, content-addressed staged helper context, one injected helper transport, blocking wait, literal budgets, raw-artifact locators, compact results, termination checks, and Visible Value receipts. No production provider transport, Context Firewall integration, comparative benchmark, or autonomous orchestration exists." + "status": "PROTOTYPED", + "assessment": "Provider-free bounded core, deterministic Tasks phase routing and consolidated Routing Policy contract. Tasks uses an execution-scoped capability model gateway; no comparative savings established." }, "drift_status": "RESTORED_PUBLIC_HOME", "current_name_accuracy": "Canonical name and repository boundary now match the restored primary-developer theory.", @@ -54,14 +54,19 @@ "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py", "https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/evidence/release-001/verification.json", "program/evidence/gearbox-publication/README.md", - "https://github.com/opsle/gearbox/pull/1" + "https://github.com/opsle/gearbox/pull/1", + "https://github.com/opsle/gearbox/blob/12112c3a04d72b768f3252a6d0f6ea455816c9a1/README.md" ], "confidence": "HIGH", "unresolved_questions": [ "Which production helper transport can independently prove filesystem, credential, network, provider, and termination bounds?", "Which external routing and authorization profiles should receive formal Gearbox adapters without moving their ownership into this repository?", "What comparative provider-session, correctness, and primary-context effects survive a controlled benchmark without inventing counterfactual savings?" - ] + ], + "source_repository": "gearbox", + "highest_evidenced_stage": "PROTOTYPED", + "concept_disposition": "ACTIVE", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "agent-trajectory-profiler", @@ -93,14 +98,19 @@ "evidence_references": [ "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/README.md", "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/metrics.js", - "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/value-summary.js" + "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/src/value-summary.js", + "https://github.com/opsle/agent-trajectory-profiler/blob/0a89661640721d6a39f127514b993d29bd728d47/README.md" ], "confidence": "HIGH", "unresolved_questions": [ - "Should generic Visible Value receipt validation remain in this repository?", + "How should independent pinned profiler validation stay conformant with the reusable Visible Value contract and reporting layer?", "Which metrics predict useful agent outcomes across models and tasks?", "How should semantic-region ground truth be established?" - ] + ], + "source_repository": "agent-trajectory-profiler", + "highest_evidenced_stage": "VERIFIED", + "concept_disposition": "ACTIVE", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "semantic-edit-protocol", @@ -131,52 +141,55 @@ "current_name_accuracy": "Accurate.", "evidence_references": [ "https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/README.md", - "https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/SPEC.md" + "https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/SPEC.md", + "https://github.com/opsle/semantic-edit-protocol/blob/29caad5c03827cde17aabd71c38bc25899413a33/README.md" ], "confidence": "HIGH", "unresolved_questions": [ "What is the first concrete operation and language adapter?", "Can anchors, transactions, and rollback be made portable?", "Does the protocol outperform ordinary patching under equal correctness gates?" - ] + ], + "source_repository": "semantic-edit-protocol", + "highest_evidenced_stage": "THEORY", + "concept_disposition": "ACTIVE", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "durable-supervisor", "canonical_concept_name": "Durable Supervisor", - "one_sentence_definition": "Durable Supervisor owns autonomous objective progress, reconstructs durable state, and resumes decision-making across activations while a non-reasoning runner handles lifecycle mechanics.", + "one_sentence_definition": "Historical autonomous objective-progress hypothesis; its runtime and repository are retired and impose no current authority or prerequisite.", "original_problem": "Long-horizon autonomous work loses state across interruptions and wastes cognition on waiting, lifecycle mechanics, and reconstruction.", "primary_classification": "DURABLE_ORCHESTRATION", - "current_repository": "durable-supervisor", - "recommended_disposition": "KEEP_AS_RESEARCH", - "disposition_rationale": "The autonomous cross-activation hypothesis is distinct from Gearbox, but the installable boundary and family packaging are not yet proven.", - "disposition_gain": "The separate research home prevents durable objective ownership and reconstruction from being conflated with bounded primary-developer delegation.", - "disposition_risk": "Remaining an umbrella research repository may defer decisions about ledger, scheduler, and wakeup module ownership.", - "provenance_concerns": "Retain the extraction commit and explicitly attribute any future consolidation of ledger, scheduler, wakeup, discovery, or recovery research.", - "relationship_to_gearbox": "Formally separate: it may operate across activations without a continuously active primary developer; Gearbox remains subordinate to one primary developer's bounded call.", - "relationship_to_context_firewall": "May consume compact evidence packets from child or runner activity, but Context Firewall provides no persistence, scheduling, or objective ownership.", - "dependencies": [ - "agent-state-ledger", - "agent-scheduler-runtime", - "event-driven-agent-wakeup", - "decision-evidence-protocol" - ], + "current_repository": null, + "recommended_disposition": "RETIRED", + "disposition_rationale": "Final consolidation receipt executed this disposition; prior recommendations remain historical.", + "disposition_gain": "Current home and source history are explicit; no retired source is eligible for new workload.", + "disposition_risk": "Consolidation or retirement does not establish general correctness, independent reuse or comparative benefit.", + "provenance_concerns": "Final migration preserves source revisions, licenses, attribution, archived repository history and Tasks project history. Original analysis remains in historical_analysis.", + "relationship_to_gearbox": "Routing belongs to Gearbox; task lifecycle belongs to Tasks. Historical family analysis is preserved separately.", + "relationship_to_context_firewall": "Evidence reduction belongs to Context Firewall; historical contracts retain their provenance.", + "dependencies": [], "consumers": [], "current_implementation_fidelity": { - "status": "THEORY_ONLY", - "assessment": "The theory is coherent, but no ledger, runner, reconstruction envelope, fixture, or automated test exists. Shared fresh-child and passive-wait language caused external conflation with Gearbox." + "status": "HISTORICAL_EVIDENCE_PRESERVED", + "assessment": "Source maturity preserved independently of current home. Consolidation does not promote the concept." }, - "drift_status": "BOUNDARY_CONFLATION", - "current_name_accuracy": "Accurate when durable objective ownership and cross-activation reconstruction remain explicit.", + "drift_status": "Final retirement/consolidation receipt supersedes prior proposed disposition.", + "current_name_accuracy": "Historical concept identity retained; current repository home is recorded independently.", "evidence_references": [ "https://github.com/opsle/durable-supervisor/blob/555ebedb992ac74236bb7da8230b4d6b0489830b/README.md", - "https://github.com/opsle/durable-supervisor/blob/555ebedb992ac74236bb7da8230b4d6b0489830b/ARCHITECTURE.md" + "https://github.com/opsle/durable-supervisor/blob/555ebedb992ac74236bb7da8230b4d6b0489830b/ARCHITECTURE.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json" ], "confidence": "HIGH", "unresolved_questions": [ - "What is the minimum durable reconstruction envelope?", - "Which family mechanisms should share one implementation repository?", - "Can correctness match a continuously active supervisor under restart and duplicate-event faults?" - ] + "Highest evidenced research maturity is preserved; this historical source has no active work or activation requirement." + ], + "source_repository": "durable-supervisor", + "highest_evidenced_stage": "VERIFIED", + "concept_disposition": "RETIRED", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "event-driven-agent-wakeup", @@ -184,38 +197,39 @@ "one_sentence_definition": "Event-Driven Agent Wakeup durably registers a wait, suspends model activity, and reactivates an orchestrator exactly once on a decision-relevant event.", "original_problem": "Polling and waiting were implemented as repeated model turns even though they are runtime concerns.", "primary_classification": "DURABLE_ORCHESTRATION", - "current_repository": "event-driven-agent-wakeup", - "recommended_disposition": "CONSOLIDATE_WITH_OTHER", - "disposition_rationale": "Durable wait registration, event delivery, scheduling, and reconstruction share one state boundary with the Durable Supervisor family.", - "disposition_gain": "One durable event and scheduling contract avoids duplicate persistence and wakeup semantics.", - "disposition_risk": "Independent polling-versus-wakeup benchmarks and reusable adapters could become less visible.", - "provenance_concerns": "Preserve the repository commit, two prototype tests, lifecycle record, citations, and a future redirect before consolidation.", - "relationship_to_gearbox": "Not Gearbox's passive wait: Gearbox blocks synchronously at the OS or transport layer inside one bounded call; this concept persists a wait and reactivates an orchestrator.", - "relationship_to_context_firewall": "May use adapters for event/process chatter, but wake identity, timeout, failure, approval, and loss signals are unsafe to hide.", + "current_repository": "tasks", + "recommended_disposition": "CONSOLIDATED", + "disposition_rationale": "Final consolidation receipt executed this disposition; prior recommendations remain historical.", + "disposition_gain": "Current home and source history are explicit; no retired source is eligible for new workload.", + "disposition_risk": "Consolidation or retirement does not establish general correctness, independent reuse or comparative benefit.", + "provenance_concerns": "Final migration preserves source revisions, licenses, attribution, archived repository history and Tasks project history. Original analysis remains in historical_analysis.", + "relationship_to_gearbox": "Routing belongs to Gearbox; task lifecycle belongs to Tasks. Historical family analysis is preserved separately.", + "relationship_to_context_firewall": "Evidence reduction belongs to Context Firewall; historical contracts retain their provenance.", "dependencies": [ "agent-state-ledger", "agent-scheduler-runtime" ], - "consumers": [ - "durable-supervisor" - ], + "consumers": [], "current_implementation_fidelity": { - "status": "NARROW_PROTOTYPE", - "assessment": "The eight-line in-memory transition function exercises duplicate and wrong-wait behavior only; it does not implement durable registration, restart recovery, event delivery, or timeouts." + "status": "HISTORICAL_EVIDENCE_PRESERVED", + "assessment": "Contracts and source provenance now reside in tasks; migration does not establish new research maturity." }, - "drift_status": "PROTOTYPE_SCOPE_OVERSTATED", - "current_name_accuracy": "Accurate for the theory; the current implementation is not durable.", + "drift_status": "Final retirement/consolidation receipt supersedes prior proposed disposition.", + "current_name_accuracy": "Historical concept identity retained; current repository home is recorded independently.", "evidence_references": [ "https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/README.md", "https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/src/runtime.js", - "https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/tests/runtime.test.js" + "https://github.com/opsle/event-driven-agent-wakeup/blob/a6209860c2151450cc28ed648bc8c2631c8db7ef/tests/runtime.test.js", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json" ], "confidence": "MEDIUM_HIGH", "unresolved_questions": [ - "Who owns timeout production and late-event retention?", - "What proves exactly-once decision activation across restart?", - "Do reusable wait adapters justify independent packaging?" - ] + "Highest evidenced research maturity is preserved; this historical source has no active work or activation requirement." + ], + "source_repository": "event-driven-agent-wakeup", + "highest_evidenced_stage": "PROTOTYPED", + "concept_disposition": "CONSOLIDATED", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "context-firewall", @@ -235,8 +249,7 @@ "consumers": [ "gearbox", "decision-evidence-protocol", - "agent-trajectory-profiler", - "durable-supervisor" + "agent-trajectory-profiler" ], "intended_adapter_families": [ "tests", @@ -248,22 +261,27 @@ "helper_agent_result" ], "current_implementation_fidelity": { - "status": "ADAPTER_PROTOTYPED", - "assessment": "The strict TAP-compatible reducer faithfully implements one deterministic adapter and packet profile. Other adapter families and the safe correctness frontier do not exist." + "status": "PROTOTYPED", + "assessment": "Flat TAP and Node spec/dot reporter reducer with semantic-only projection; Tasks command-evidence adapter integrated. Actual downstream delivery is separate from canonical packet telemetry." }, "drift_status": "PARTIAL_ADAPTER_SCOPE", "current_name_accuracy": "Accurate for the concept; implementation maturity must remain adapter-scoped.", "evidence_references": [ "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/README.md", "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/SPEC.md", - "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/src/reducer.js" + "https://github.com/opsle/context-firewall/blob/953c48f1cfd154d6b7ed10b51b87fe54e4df45f2/src/reducer.js", + "https://github.com/opsle/context-firewall/blob/6dd6e5fdf21f28dc5ebfa07954aaa9bed2dbcc32/README.md" ], "confidence": "HIGH", "unresolved_questions": [ "What common adapter interface preserves tool-specific safety?", "How are raw artifact availability and hash validity independently proven?", "Where is the experimentally safe reduction frontier by task and adapter family?" - ] + ], + "source_repository": "context-firewall", + "highest_evidenced_stage": "PROTOTYPED", + "concept_disposition": "ACTIVE", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "decision-evidence-protocol", @@ -285,7 +303,6 @@ "context-firewall", "agent-trajectory-profiler", "verifiable-agent-handoff", - "durable-supervisor", "agent-recovery-policy" ], "current_implementation_fidelity": { @@ -297,14 +314,19 @@ "evidence_references": [ "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/README.md", "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/context-firewall-v1.js", - "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/validate.js" + "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/src/validate.js", + "https://github.com/opsle/decision-evidence-protocol/blob/b17ae3b41cea7cb0b9e0befe43e885b5aa0e4a09/README.md" ], "confidence": "HIGH", "unresolved_questions": [ "What is the canonical generic envelope independent of one producer profile?", - "Should Decision Evidence be the sole normative Visible Value validator?", + "How should independent Decision Evidence conformance checks stay compatible with the reusable Visible Value validator without disturbing pinned experiment evidence?", "Which tool-class profiles demonstrate adequate decision coverage?" - ] + ], + "source_repository": "decision-evidence-protocol", + "highest_evidenced_stage": "VERIFIED", + "concept_disposition": "ACTIVE", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "agent-state-ledger", @@ -312,40 +334,42 @@ "one_sentence_definition": "Agent State Ledger records immutable orchestration facts and derives a deterministic stale-aware projection of current work, evidence, authority, contradictions, and uncertainty.", "original_problem": "Conversation history could not reliably reconstruct what autonomous work was done, verified, failed, authorized, uncertain, or next after interruption.", "primary_classification": "DURABLE_ORCHESTRATION", - "current_repository": "agent-state-ledger", - "recommended_disposition": "CONSOLIDATE_WITH_OTHER", - "disposition_rationale": "The ledger is the authoritative history and projection module of the Durable Supervisor family, not a second independent state owner.", - "disposition_gain": "One durable state contract prevents supervisor, scheduler, discovery, and recovery from maintaining competing histories.", - "disposition_risk": "Potential reuse as a generic audit/projection library may be obscured.", - "provenance_concerns": "Preserve the current repository, extraction commit, benchmark theory, and future negative evidence through an attributable module and redirect.", - "relationship_to_gearbox": "Normally unnecessary for one bounded Gearbox call; absorbing it would pull Gearbox toward persistent supervision.", - "relationship_to_context_firewall": "May record packet hashes, raw locators, and escalation facts, but compact packets are not a replacement for authoritative history.", + "current_repository": "tasks", + "recommended_disposition": "CONSOLIDATED", + "disposition_rationale": "Final consolidation receipt executed this disposition; prior recommendations remain historical.", + "disposition_gain": "Current home and source history are explicit; no retired source is eligible for new workload.", + "disposition_risk": "Consolidation or retirement does not establish general correctness, independent reuse or comparative benefit.", + "provenance_concerns": "Final migration preserves source revisions, licenses, attribution, archived repository history and Tasks project history. Original analysis remains in historical_analysis.", + "relationship_to_gearbox": "Routing belongs to Gearbox; task lifecycle belongs to Tasks. Historical family analysis is preserved separately.", + "relationship_to_context_firewall": "Evidence reduction belongs to Context Firewall; historical contracts retain their provenance.", "dependencies": [ "decision-evidence-protocol" ], "consumers": [ - "durable-supervisor", "event-driven-agent-wakeup", "agent-scheduler-runtime", "agent-discovery-control", "agent-recovery-policy" ], "current_implementation_fidelity": { - "status": "THEORY_ONLY", - "assessment": "The theory is coherent, but the generic specification defines no event schema, projection contract, contradiction rules, fixture, implementation, or test." + "status": "HISTORICAL_EVIDENCE_PRESERVED", + "assessment": "Contracts and source provenance now reside in tasks; migration does not establish new research maturity." }, - "drift_status": "OVER_EXTRACTED_REPOSITORY_BOUNDARY", - "current_name_accuracy": "Accurate, though the mechanism is not intrinsically agent-specific.", + "drift_status": "Final retirement/consolidation receipt supersedes prior proposed disposition.", + "current_name_accuracy": "Historical concept identity retained; current repository home is recorded independently.", "evidence_references": [ "https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/README.md", - "https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/SPEC.md" + "https://github.com/opsle/agent-state-ledger/blob/acab03b1ff7168222552050e21e7553b07d00e7c/SPEC.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json" ], "confidence": "MEDIUM_HIGH", "unresolved_questions": [ - "Do non-orchestration consumers justify independent packaging?", - "What event and projection schemas preserve contradictions and uncertainty?", - "What compaction is safe without losing reconstructability?" - ] + "Highest evidenced research maturity is preserved; this historical source has no active work or activation requirement." + ], + "source_repository": "agent-state-ledger", + "highest_evidenced_stage": "THEORY", + "concept_disposition": "CONSOLIDATED", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "agent-scheduler-runtime", @@ -353,39 +377,41 @@ "one_sentence_definition": "Agent Scheduler Runtime performs deterministic readiness, queue, time, cancellation, pause, and bounded-concurrency transitions over durable execution identities.", "original_problem": "Queues, dependency readiness, retry timing, timeouts, and schedules were mechanical decisions left inside model loops.", "primary_classification": "DURABLE_ORCHESTRATION", - "current_repository": "agent-scheduler-runtime", - "recommended_disposition": "CONSOLIDATE_WITH_OTHER", - "disposition_rationale": "Scheduling is the deterministic runtime module of Durable Supervisor and shares persistence and event boundaries with the ledger and wakeup mechanism.", - "disposition_gain": "One orchestration family can own readiness, time, pause, durable event, and reconstruction semantics without duplicating state.", - "disposition_risk": "Independent scheduler reuse may be lost and the Durable Supervisor family could become too monolithic.", - "provenance_concerns": "Retain the existing theory and benchmark plan as a named module with attributable history if folded.", - "relationship_to_gearbox": "Not Gearbox core: a bounded Gearbox call does not require queues, schedules, Global Pause, autonomous retries, or restart recovery.", - "relationship_to_context_firewall": "May expose scheduler and process evidence through adapters, but Context Firewall cannot decide readiness or mutate schedules.", + "current_repository": "tasks", + "recommended_disposition": "CONSOLIDATED", + "disposition_rationale": "Final consolidation receipt executed this disposition; prior recommendations remain historical.", + "disposition_gain": "Current home and source history are explicit; no retired source is eligible for new workload.", + "disposition_risk": "Consolidation or retirement does not establish general correctness, independent reuse or comparative benefit.", + "provenance_concerns": "Final migration preserves source revisions, licenses, attribution, archived repository history and Tasks project history. Original analysis remains in historical_analysis.", + "relationship_to_gearbox": "Routing belongs to Gearbox; task lifecycle belongs to Tasks. Historical family analysis is preserved separately.", + "relationship_to_context_firewall": "Evidence reduction belongs to Context Firewall; historical contracts retain their provenance.", "dependencies": [ "agent-state-ledger", "agent-resource-claims" ], "consumers": [ - "durable-supervisor", "event-driven-agent-wakeup", "controlled-agent-acceptance" ], "current_implementation_fidelity": { - "status": "THEORY_ONLY", - "assessment": "No queue state machine, store, fake clock, pause contract, fixture, implementation, or automated test exists; the initial theory also overlaps claims, recovery, and routing." + "status": "HISTORICAL_EVIDENCE_PRESERVED", + "assessment": "Contracts and source provenance now reside in tasks; migration does not establish new research maturity." }, - "drift_status": "OVER_BROAD_INITIAL_SCOPE", - "current_name_accuracy": "Accurate if claims, routing, and recovery remain external.", + "drift_status": "Final retirement/consolidation receipt supersedes prior proposed disposition.", + "current_name_accuracy": "Historical concept identity retained; current repository home is recorded independently.", "evidence_references": [ "https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/README.md", - "https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/THEORY.md" + "https://github.com/opsle/agent-scheduler-runtime/blob/d97cd3c218b20e0b2b0e09873f6a3d15c396b3a0/THEORY.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json" ], "confidence": "HIGH", "unresolved_questions": [ - "What exact transitions belong to scheduler versus claims and recovery?", - "Who owns timeout and provider-availability events?", - "Can the scheduler remain a reusable module inside a consolidated repository?" - ] + "Highest evidenced research maturity is preserved; this historical source has no active work or activation requirement." + ], + "source_repository": "agent-scheduler-runtime", + "highest_evidenced_stage": "THEORY", + "concept_disposition": "CONSOLIDATED", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "verifiable-agent-handoff", @@ -393,14 +419,14 @@ "one_sentence_definition": "Verifiable Agent Handoff seals exact source, result, artifact, and authority identities before source cleanup so a fresh verifier can reconstruct and validate without fallback.", "original_problem": "A verifier could not establish an executor's exact result when evidence remained only in the executor's mutable or disposable environment.", "primary_classification": "CROSS_CUTTING_PROTOCOL", - "current_repository": "verifiable-agent-handoff", - "recommended_disposition": "KEEP_AS_PROTOCOL", - "disposition_rationale": "The durable result-transfer and independent-verification boundary is reusable across Gearbox, durable orchestration, CI, and isolation systems.", - "disposition_gain": "Artifact identity, publication-before-destruction, and reconstruction semantics remain independently conformable.", - "disposition_risk": "Another evidence envelope may duplicate Decision Evidence unless layering is explicit.", - "provenance_concerns": "Preserve the result-loss failure narrative and scope the current HMAC prototype to manifest authentication rather than claiming full handoff.", - "relationship_to_gearbox": "Optional external protocol when a bounded helper's source environment will be destroyed; not gear selection, admission, or lifecycle ownership.", - "relationship_to_context_firewall": "Handoff preserves exact durable artifacts; Context Firewall decides which verified facts enter model context and must escalate rather than hide required handoff evidence.", + "current_repository": "tasks", + "recommended_disposition": "CONSOLIDATED", + "disposition_rationale": "Final consolidation receipt executed this disposition; prior recommendations remain historical.", + "disposition_gain": "Current home and source history are explicit; no retired source is eligible for new workload.", + "disposition_risk": "Consolidation or retirement does not establish general correctness, independent reuse or comparative benefit.", + "provenance_concerns": "Final migration preserves source revisions, licenses, attribution, archived repository history and Tasks project history. Original analysis remains in historical_analysis.", + "relationship_to_gearbox": "Routing belongs to Gearbox; task lifecycle belongs to Tasks. Historical family analysis is preserved separately.", + "relationship_to_context_firewall": "Evidence reduction belongs to Context Firewall; historical contracts retain their provenance.", "dependencies": [ "decision-evidence-protocol" ], @@ -410,21 +436,24 @@ "controlled-agent-acceptance" ], "current_implementation_fidelity": { - "status": "SUBCOMPONENT_PROTOTYPED", - "assessment": "The HMAC manifest prototype validates required bindings and a caller-supplied destruction assertion; it does not publish, transport, prove destruction, reconstruct, or independently verify artifacts." + "status": "HISTORICAL_EVIDENCE_PRESERVED", + "assessment": "Contracts and source provenance now reside in tasks; migration does not establish new research maturity." }, - "drift_status": "PROTOTYPE_SCOPE_OVERSTATED", - "current_name_accuracy": "Accurate for the theory; current code is a seal subcomponent.", + "drift_status": "Final retirement/consolidation receipt supersedes prior proposed disposition.", + "current_name_accuracy": "Historical concept identity retained; current repository home is recorded independently.", "evidence_references": [ "https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/README.md", - "https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/src/seal.js" + "https://github.com/opsle/verifiable-agent-handoff/blob/399e5cfae94345affa3f087f0f6eb9e77669d33c/src/seal.js", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json" ], "confidence": "HIGH", "unresolved_questions": [ - "How does the handoff profile layer over Decision Evidence?", - "What independently proves source destruction and publication durability?", - "Which non-Git artifact formats are supported?" - ] + "Highest evidenced research maturity is preserved; this historical source has no active work or activation requirement." + ], + "source_repository": "verifiable-agent-handoff", + "highest_evidenced_stage": "PROTOTYPED", + "concept_disposition": "CONSOLIDATED", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "agent-routing-policy", @@ -432,14 +461,14 @@ "one_sentence_definition": "Agent Routing Policy filters and selects a purpose-bound model, provider, profile, or executor route under exact capability, access, availability, independence, and budget constraints with a durable reason.", "original_problem": "Opaque routing violated provider budgets, access boundaries, exact model or profile constraints, and reviewer independence.", "primary_classification": "GEARBOX_SUPPORTING_POLICY", - "current_repository": "agent-routing-policy", - "recommended_disposition": "FUTURE_GEARBOX_POLICY", - "disposition_rationale": "Gearbox needs bounded model, effort, and provider route selection after deterministic-versus-cognitive admission, while broader review and recovery routing must remain explicit extensions.", - "disposition_gain": "A versioned Gearbox-facing selection policy would replace ad hoc model and cost choice.", - "disposition_risk": "Absorption could erase reviewer-independence or autonomous routing research and accidentally authorize fallback retries.", - "provenance_concerns": "Migrate only the Gearbox-facing route profile; retain the original general routing theory, commit, and later experiments with attribution.", - "relationship_to_gearbox": "Supporting policy: Gearbox first admits deterministic versus cognitive work; routing then selects one allowed cognitive route and never autonomously retries or falls back.", - "relationship_to_context_firewall": "May consume compact availability or performance evidence, but incomplete or uncertain packets must make selection fail closed.", + "current_repository": "gearbox", + "recommended_disposition": "CONSOLIDATED", + "disposition_rationale": "Final consolidation receipt executed this disposition; prior recommendations remain historical.", + "disposition_gain": "Current home and source history are explicit; no retired source is eligible for new workload.", + "disposition_risk": "Consolidation or retirement does not establish general correctness, independent reuse or comparative benefit.", + "provenance_concerns": "Final migration preserves source revisions, licenses, attribution, archived repository history and Tasks project history. Original analysis remains in historical_analysis.", + "relationship_to_gearbox": "Routing belongs to Gearbox; task lifecycle belongs to Tasks. Historical family analysis is preserved separately.", + "relationship_to_context_firewall": "Evidence reduction belongs to Context Firewall; historical contracts retain their provenance.", "dependencies": [ "decision-evidence-protocol" ], @@ -449,21 +478,24 @@ "controlled-agent-acceptance" ], "current_implementation_fidelity": { - "status": "THEORY_ONLY", - "assessment": "The initial theory mixes gear/model selection, provider routing, authorization inputs, reviewer independence, and fallback eligibility; no evaluator or normative precedence exists." + "status": "HISTORICAL_EVIDENCE_PRESERVED", + "assessment": "Contracts and source provenance now reside in gearbox; migration does not establish new research maturity." }, - "drift_status": "MIXED_DECISION_OWNERSHIP", - "current_name_accuracy": "Accurate for broad routing; not a substitute name for Gearbox.", + "drift_status": "Final retirement/consolidation receipt supersedes prior proposed disposition.", + "current_name_accuracy": "Historical concept identity retained; current repository home is recorded independently.", "evidence_references": [ "https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/README.md", - "https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/THEORY.md" + "https://github.com/opsle/agent-routing-policy/blob/43fc2a72d2c8494b2dcdca7b5a209de61d8fe2d8/THEORY.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json" ], "confidence": "MEDIUM_HIGH", "unresolved_questions": [ - "Where exactly does gear selection stop and route selection begin?", - "Who owns fallback eligibility and reviewer routing?", - "What deterministic precedence applies to price, quality, access, and availability?" - ] + "Highest evidenced research maturity is preserved; this historical source has no active work or activation requirement." + ], + "source_repository": "agent-routing-policy", + "highest_evidenced_stage": "THEORY", + "concept_disposition": "CONSOLIDATED", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "agent-resource-claims", @@ -471,14 +503,14 @@ "one_sentence_definition": "Agent Resource Claims establishes current concurrent ownership through canonical resource identities, atomic shared or exclusive claim sets, expiring leases, and fencing tokens.", "original_problem": "A process could remain alive after losing authority, while independent locks could deadlock or admit partially conflicting work.", "primary_classification": "GEARBOX_SUPPORTING_POLICY", - "current_repository": "agent-resource-claims", - "recommended_disposition": "KEEP_AS_RESEARCH", - "disposition_rationale": "The policy can support Gearbox writes, durable orchestration, semantic editing, and isolation, but a portable state machine and independent packaging value are not yet proven.", - "disposition_gain": "Research remains broadly reusable and avoids prematurely coupling lease and fence semantics to Gearbox.", - "disposition_risk": "Repository granularity and overlap with scheduler and authorization remain unresolved.", - "provenance_concerns": "Preserve the fenced-claim failure narrative and original extraction before any later module migration.", - "relationship_to_gearbox": "Optional policy for bounded shared-resource or write authority; unnecessary when a Gearbox task touches no shared mutable resource.", - "relationship_to_context_firewall": "Claim observations may be compacted, but fence mismatch, expiry, and takeover evidence are unsafe to hide and reduction never grants authority.", + "current_repository": "tasks", + "recommended_disposition": "CONSOLIDATED", + "disposition_rationale": "Final consolidation receipt executed this disposition; prior recommendations remain historical.", + "disposition_gain": "Current home and source history are explicit; no retired source is eligible for new workload.", + "disposition_risk": "Consolidation or retirement does not establish general correctness, independent reuse or comparative benefit.", + "provenance_concerns": "Final migration preserves source revisions, licenses, attribution, archived repository history and Tasks project history. Original analysis remains in historical_analysis.", + "relationship_to_gearbox": "Routing belongs to Gearbox; task lifecycle belongs to Tasks. Historical family analysis is preserved separately.", + "relationship_to_context_firewall": "Evidence reduction belongs to Context Firewall; historical contracts retain their provenance.", "dependencies": [], "consumers": [ "semantic-edit-protocol", @@ -487,21 +519,24 @@ "ephemeral-agent-workers" ], "current_implementation_fidelity": { - "status": "THEORY_ONLY", - "assessment": "The useful invariants are present, but no catalog schema, conflict lattice, lease state machine, store adapter, concurrency fixture, or automated test exists." + "status": "HISTORICAL_EVIDENCE_PRESERVED", + "assessment": "Contracts and source provenance now reside in tasks; migration does not establish new research maturity." }, - "drift_status": "BOUNDARY_OVERLAP", - "current_name_accuracy": "Substantially accurate; the mechanism is generic concurrency authority rather than intrinsically agent-specific.", + "drift_status": "Final retirement/consolidation receipt supersedes prior proposed disposition.", + "current_name_accuracy": "Historical concept identity retained; current repository home is recorded independently.", "evidence_references": [ "https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/README.md", - "https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/THEORY.md" + "https://github.com/opsle/agent-resource-claims/blob/dfe0fbc90c67ce5ef4256354bb62f1d511b1304c/THEORY.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json" ], "confidence": "MEDIUM", "unresolved_questions": [ - "Does a portable resource identity and conflict lattice justify a standalone package?", - "Which time, fairness, and takeover semantics are required?", - "Where does resource authority stop and execution authorization begin?" - ] + "Highest evidenced research maturity is preserved; this historical source has no active work or activation requirement." + ], + "source_repository": "agent-resource-claims", + "highest_evidenced_stage": "THEORY", + "concept_disposition": "CONSOLIDATED", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "agent-discovery-control", @@ -509,37 +544,38 @@ "one_sentence_definition": "Agent Discovery Control admits, merges, ignores, or escalates provenance-bearing discovered-work proposals under similarity, depth, existing-work, and storm-budget constraints.", "original_problem": "Autonomous discovery created duplicate proposals, recursive work trees, and churn for already-satisfied work.", "primary_classification": "DURABLE_ORCHESTRATION", - "current_repository": "agent-discovery-control", - "recommended_disposition": "CONSOLIDATE_WITH_OTHER", - "disposition_rationale": "Autonomous discovered-work admission depends on the Durable Supervisor family's objective graph, ledger, and shared budgets.", - "disposition_gain": "One durable proposal and convergence model avoids a standalone shell with duplicate state.", - "disposition_risk": "Independent admission-policy research and non-agent issue-triage reuse could become less visible.", - "provenance_concerns": "Preserve the repository commit, benchmark plan, extraction citation, and future fixtures through an attributable module and redirect.", - "relationship_to_gearbox": "Outside Gearbox: canonical Gearbox starts from an explicit primary-developer task and must not create recursive autonomous work.", - "relationship_to_context_firewall": "May receive compact proposal evidence, but Context Firewall cannot decide semantic duplication or replace proposal provenance.", + "current_repository": "tasks", + "recommended_disposition": "CONSOLIDATED", + "disposition_rationale": "Final consolidation receipt executed this disposition; prior recommendations remain historical.", + "disposition_gain": "Current home and source history are explicit; no retired source is eligible for new workload.", + "disposition_risk": "Consolidation or retirement does not establish general correctness, independent reuse or comparative benefit.", + "provenance_concerns": "Final migration preserves source revisions, licenses, attribution, archived repository history and Tasks project history. Original analysis remains in historical_analysis.", + "relationship_to_gearbox": "Routing belongs to Gearbox; task lifecycle belongs to Tasks. Historical family analysis is preserved separately.", + "relationship_to_context_firewall": "Evidence reduction belongs to Context Firewall; historical contracts retain their provenance.", "dependencies": [ "agent-state-ledger", "decision-evidence-protocol" ], - "consumers": [ - "durable-supervisor" - ], + "consumers": [], "current_implementation_fidelity": { - "status": "THEORY_ONLY", - "assessment": "The theory is coherent, but the generic spec does not define proposal dispositions, similarity evidence, budgets, convergence, or already-satisfied proofs; no implementation exists." + "status": "HISTORICAL_EVIDENCE_PRESERVED", + "assessment": "Contracts and source provenance now reside in tasks; migration does not establish new research maturity." }, - "drift_status": "OVER_EXTRACTED_REPOSITORY_BOUNDARY", - "current_name_accuracy": "Accurate for an orchestration admission policy, not a standalone agent product.", + "drift_status": "Final retirement/consolidation receipt supersedes prior proposed disposition.", + "current_name_accuracy": "Historical concept identity retained; current repository home is recorded independently.", "evidence_references": [ "https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/README.md", - "https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/THEORY.md" + "https://github.com/opsle/agent-discovery-control/blob/926547dd9fd713990b1d6f1f2e650aa6c0883564/THEORY.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json" ], "confidence": "HIGH", "unresolved_questions": [ - "What semantic duplicate oracle is safe?", - "How do storm budgets scale with objective complexity?", - "Does a reusable admission library exist independently of a supervisor?" - ] + "Highest evidenced research maturity is preserved; this historical source has no active work or activation requirement." + ], + "source_repository": "agent-discovery-control", + "highest_evidenced_stage": "THEORY", + "concept_disposition": "CONSOLIDATED", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "agent-execution-authorization", @@ -547,14 +583,14 @@ "one_sentence_definition": "Agent Execution Authorization validates capability-style grants that bind immutable source authority and current execution authority to one exact task, target, purpose, resource, route, and expiry.", "original_problem": "Machine actors received broad or caller-asserted authority instead of server-derived grants bound to exact work and current state.", "primary_classification": "GEARBOX_SUPPORTING_POLICY", - "current_repository": "agent-execution-authorization", - "recommended_disposition": "FUTURE_GEARBOX_POLICY", - "disposition_rationale": "Gearbox must fail closed before deterministic or cognitive execution, while the grant contract should remain reusable by durable orchestration and isolation infrastructure.", - "disposition_gain": "Gearbox obtains an exact authority envelope instead of ad hoc caller assertions.", - "disposition_risk": "Narrow absorption could discard cross-orchestrator composition or duplicate claim, routing, and handoff state.", - "provenance_concerns": "Preserve a separately versioned grant schema and conformance history, plus the immutable-source/current-continuation failure narrative.", - "relationship_to_gearbox": "Supporting policy that supplies and validates authority; Gearbox owns applying the result to its requested gear, context, model, budget, and task.", - "relationship_to_context_firewall": "Independent: reduced or redacted evidence is never authorization; a firewall may compact an authorization receipt only after authoritative validation.", + "current_repository": "tasks", + "recommended_disposition": "CONSOLIDATED", + "disposition_rationale": "Final consolidation receipt executed this disposition; prior recommendations remain historical.", + "disposition_gain": "Current home and source history are explicit; no retired source is eligible for new workload.", + "disposition_risk": "Consolidation or retirement does not establish general correctness, independent reuse or comparative benefit.", + "provenance_concerns": "Final migration preserves source revisions, licenses, attribution, archived repository history and Tasks project history. Original analysis remains in historical_analysis.", + "relationship_to_gearbox": "Routing belongs to Gearbox; task lifecycle belongs to Tasks. Historical family analysis is preserved separately.", + "relationship_to_context_firewall": "Evidence reduction belongs to Context Firewall; historical contracts retain their provenance.", "dependencies": [ "agent-resource-claims", "decision-evidence-protocol", @@ -566,21 +602,24 @@ "controlled-agent-acceptance" ], "current_implementation_fidelity": { - "status": "THEORY_ONLY", - "assessment": "The source/current authority distinction is preserved, but no portable grant, issuance, revocation, replay, composition, validator, or conformance suite exists." + "status": "HISTORICAL_EVIDENCE_PRESERVED", + "assessment": "Contracts and source provenance now reside in tasks; migration does not establish new research maturity." }, - "drift_status": "BOUNDARY_OVERLAP", - "current_name_accuracy": "Accurate.", + "drift_status": "Final retirement/consolidation receipt supersedes prior proposed disposition.", + "current_name_accuracy": "Historical concept identity retained; current repository home is recorded independently.", "evidence_references": [ "https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/README.md", - "https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/THEORY.md" + "https://github.com/opsle/agent-execution-authorization/blob/8fb02e943c83fae121c63185d2f0d0dde8c4260a/THEORY.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json" ], "confidence": "HIGH", "unresolved_questions": [ - "What is the portable grant schema and revocation oracle?", - "Are claims and routes embedded or content-addressed references?", - "How are source and continuation authority composed without replay?" - ] + "Highest evidenced research maturity is preserved; this historical source has no active work or activation requirement." + ], + "source_repository": "agent-execution-authorization", + "highest_evidenced_stage": "THEORY", + "concept_disposition": "CONSOLIDATED", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "controlled-agent-acceptance", @@ -588,14 +627,14 @@ "one_sentence_definition": "Controlled Agent Acceptance is an experiment-control method for one-shot real-agent runs under an immutable manifest, literal budgets, fail-closed activation, restoration, evidence retention, and residue reconciliation.", "original_problem": "Real-agent acceptance could exceed provider budgets, target the wrong revision, conflict with active work, leave reusable authority, fail to restore pause state, or leave residue.", "primary_classification": "RESEARCH_ONLY_HYPOTHESIS", - "current_repository": "controlled-agent-acceptance", - "recommended_disposition": "KEEP_AS_RESEARCH", - "disposition_rationale": "This is a provider-independent acceptance methodology for testing autonomous systems, not a production runtime primitive or Gearbox component.", - "disposition_gain": "Harness and system-under-test verdicts, first-failure behavior, budgets, and restoration remain independently testable.", - "disposition_risk": "A separate repository may be excessive if the eventual controller belongs beside each system under test.", - "provenance_concerns": "Retain predecessor run citations and future immutable manifests without promoting private operational evidence into generalized proof.", - "relationship_to_gearbox": "May later test one-shot Gearbox budget and cleanup behavior, but does not define or implement Gearbox.", - "relationship_to_context_firewall": "Can use adapters for captured logs; preflight defects, first failures, budget use, restoration failures, residue, and dual verdicts are unsafe to hide.", + "current_repository": "tasks", + "recommended_disposition": "CONSOLIDATED", + "disposition_rationale": "Final consolidation receipt executed this disposition; prior recommendations remain historical.", + "disposition_gain": "Current home and source history are explicit; no retired source is eligible for new workload.", + "disposition_risk": "Consolidation or retirement does not establish general correctness, independent reuse or comparative benefit.", + "provenance_concerns": "Final migration preserves source revisions, licenses, attribution, archived repository history and Tasks project history. Original analysis remains in historical_analysis.", + "relationship_to_gearbox": "Routing belongs to Gearbox; task lifecycle belongs to Tasks. Historical family analysis is preserved separately.", + "relationship_to_context_firewall": "Evidence reduction belongs to Context Firewall; historical contracts retain their provenance.", "dependencies": [ "agent-execution-authorization", "agent-routing-policy", @@ -604,21 +643,24 @@ ], "consumers": [], "current_implementation_fidelity": { - "status": "THEORY_ONLY", - "assessment": "The theory is coherent, but no manifest schema, state machine, CAS semantics, restoration protocol, residue model, controller, fixture, or test exists." + "status": "HISTORICAL_EVIDENCE_PRESERVED", + "assessment": "Contracts and source provenance now reside in tasks; migration does not establish new research maturity." }, - "drift_status": "REPOSITORY_BOUNDARY_OVERSTATED", - "current_name_accuracy": "Accurate for a methodology or harness.", + "drift_status": "Final retirement/consolidation receipt supersedes prior proposed disposition.", + "current_name_accuracy": "Historical concept identity retained; current repository home is recorded independently.", "evidence_references": [ "https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/README.md", - "https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/ARCHITECTURE.md" + "https://github.com/opsle/controlled-agent-acceptance/blob/2d652adf56e53953327d09b1ba9c4a9c3445f052/ARCHITECTURE.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json" ], "confidence": "MEDIUM_HIGH", "unresolved_questions": [ - "Which manifest and interruption semantics generalize beyond the predecessor?", - "Where should the eventual harness implementation live?", - "How are non-reusable grants and residue proved across runtimes?" - ] + "Highest evidenced research maturity is preserved; this historical source has no active work or activation requirement." + ], + "source_repository": "controlled-agent-acceptance", + "highest_evidenced_stage": "THEORY", + "concept_disposition": "CONSOLIDATED", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "agent-recovery-policy", @@ -626,38 +668,39 @@ "one_sentence_definition": "Agent Recovery Policy permits bounded recovery only when typed failure evidence shows that another authorized attempt can add information or change conditions, otherwise converging to a terminal or human-input state.", "original_problem": "Retries, consultations, fallbacks, and verification failures bypassed shared budgets or repeated the same failure indefinitely.", "primary_classification": "DURABLE_ORCHESTRATION", - "current_repository": "agent-recovery-policy", - "recommended_disposition": "CONSOLIDATE_WITH_OTHER", - "disposition_rationale": "Recovery permission depends on the Durable Supervisor family's attempt ledger, global budgets, scheduling, and terminal-state model.", - "disposition_gain": "One authoritative attempt and budget model prevents recovery paths from bypassing orchestration limits.", - "disposition_risk": "Independent deterministic failure-policy testing and reuse could become less visible.", - "provenance_concerns": "Retain extraction history, failure narratives, benchmark plan, and future policy fixtures as an attributable module.", - "relationship_to_gearbox": "Outside canonical Gearbox: a Gearbox run returns one failed or indeterminate result and does not autonomously retry, consult, route, or fall back.", - "relationship_to_context_firewall": "May consume typed failure packets, but incomplete, truncated, hash-invalid, or unavailable raw evidence must terminate or escalate rather than authorize recovery.", + "current_repository": "tasks", + "recommended_disposition": "CONSOLIDATED", + "disposition_rationale": "Final consolidation receipt executed this disposition; prior recommendations remain historical.", + "disposition_gain": "Current home and source history are explicit; no retired source is eligible for new workload.", + "disposition_risk": "Consolidation or retirement does not establish general correctness, independent reuse or comparative benefit.", + "provenance_concerns": "Final migration preserves source revisions, licenses, attribution, archived repository history and Tasks project history. Original analysis remains in historical_analysis.", + "relationship_to_gearbox": "Routing belongs to Gearbox; task lifecycle belongs to Tasks. Historical family analysis is preserved separately.", + "relationship_to_context_firewall": "Evidence reduction belongs to Context Firewall; historical contracts retain their provenance.", "dependencies": [ "agent-state-ledger", "agent-routing-policy", "decision-evidence-protocol" ], - "consumers": [ - "durable-supervisor" - ], + "consumers": [], "current_implementation_fidelity": { - "status": "THEORY_ONLY", - "assessment": "No failure taxonomy, attempt ledger, information-gain rule, budget transition, evaluator, fixture, or automated test exists; alternate-route ownership overlaps routing." + "status": "HISTORICAL_EVIDENCE_PRESERVED", + "assessment": "Contracts and source provenance now reside in tasks; migration does not establish new research maturity." }, - "drift_status": "BOUNDARY_OVERLAP", - "current_name_accuracy": "Accurate, with routing and scheduling boundaries made explicit.", + "drift_status": "Final retirement/consolidation receipt supersedes prior proposed disposition.", + "current_name_accuracy": "Historical concept identity retained; current repository home is recorded independently.", "evidence_references": [ "https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/README.md", - "https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md" + "https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json" ], "confidence": "HIGH", "unresolved_questions": [ - "What portable failure taxonomy and information-gain rule are adequate?", - "Does recovery choose an alternate-route class while routing chooses the exact route?", - "Which retry timing decisions belong to the scheduler?" - ] + "Highest evidenced research maturity is preserved; this historical source has no active work or activation requirement." + ], + "source_repository": "agent-recovery-policy", + "highest_evidenced_stage": "THEORY", + "concept_disposition": "CONSOLIDATED", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "ephemeral-agent-workers", @@ -665,39 +708,41 @@ "one_sentence_definition": "Ephemeral Agent Workers provides restricted-broker execution in disposable bounded environments with scoped resources, sealed result capture, termination, destruction, and independent destruction verification.", "original_problem": "Autonomous execution exposed host mounts, production secrets, broad network access, lingering processes, and ambiguous cleanup.", "primary_classification": "EXECUTION_ISOLATION_INFRASTRUCTURE", - "current_repository": "ephemeral-agent-workers", - "recommended_disposition": "KEEP_STANDALONE", - "disposition_rationale": "Isolation, containment, worker lifecycle, and destruction proof have independent use for Gearbox, durable orchestration, CI, and other untrusted workloads.", - "disposition_gain": "A separate threat model, broker contract, portability layer, and containment harness can mature independently.", - "disposition_risk": "The implementation may duplicate authorization, claims, or handoff schemas unless those remain explicit external contracts.", - "provenance_concerns": "Preserve the Incus-derived predecessor evidence while keeping Incus an optional adapter and not copying private infrastructure.", - "relationship_to_gearbox": "Optional external executor for bounded cognitive helpers; deterministic local gears need not use it, and it never selects a gear.", - "relationship_to_context_firewall": "May expose helper logs and results through adapters, but Context Firewall proves neither containment, result sealing, nor destruction.", + "current_repository": "tasks", + "recommended_disposition": "CONSOLIDATED", + "disposition_rationale": "Final consolidation receipt executed this disposition; prior recommendations remain historical.", + "disposition_gain": "Current home and source history are explicit; no retired source is eligible for new workload.", + "disposition_risk": "Consolidation or retirement does not establish general correctness, independent reuse or comparative benefit.", + "provenance_concerns": "Final migration preserves source revisions, licenses, attribution, archived repository history and Tasks project history. Original analysis remains in historical_analysis.", + "relationship_to_gearbox": "Routing belongs to Gearbox; task lifecycle belongs to Tasks. Historical family analysis is preserved separately.", + "relationship_to_context_firewall": "Evidence reduction belongs to Context Firewall; historical contracts retain their provenance.", "dependencies": [ "agent-execution-authorization", "agent-resource-claims", "verifiable-agent-handoff" ], "consumers": [ - "gearbox", - "durable-supervisor" + "gearbox" ], "current_implementation_fidelity": { - "status": "THEORY_ONLY", - "assessment": "No portable broker or worker adapter, resource profile, threat model, containment harness, destruction receipt, fixture, or test exists." + "status": "HISTORICAL_EVIDENCE_PRESERVED", + "assessment": "Contracts and source provenance now reside in tasks; migration does not establish new research maturity." }, - "drift_status": "ALIGNED_THEORY_ONLY", - "current_name_accuracy": "Reasonably accurate, though the isolation primitive is not inherently agent-specific.", + "drift_status": "Final retirement/consolidation receipt supersedes prior proposed disposition.", + "current_name_accuracy": "Historical concept identity retained; current repository home is recorded independently.", "evidence_references": [ "https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/README.md", - "https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/ARCHITECTURE.md" + "https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/ARCHITECTURE.md", + "https://github.com/opsle/tasks/blob/e1207c5264c59e14efe9838bba3a33ba504665d2/docs/migrations/20260907-consolidation.json" ], "confidence": "HIGH", "unresolved_questions": [ - "What portable isolation and destruction-proof contracts work across technologies?", - "What kernel and broker threat model is in scope?", - "How much result sealing belongs solely to Verifiable Handoff?" - ] + "Highest evidenced research maturity is preserved; this historical source has no active work or activation requirement." + ], + "source_repository": "ephemeral-agent-workers", + "highest_evidenced_stage": "THEORY", + "concept_disposition": "CONSOLIDATED", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" }, { "id": "affected-verification", @@ -720,10 +765,10 @@ "agent-trajectory-profiler" ], "current_implementation_fidelity": { - "status": "VERIFIED_NARROW_PROTOTYPE", - "assessment": "The public dependency-free Node.js core validates normalized evidence/catalog/policy input, computes reverse impact, selects and explains checks, fails closed on uncertainty, emits canonical plans and Visible Value receipts, and validates identity-bound SHADOW results. AV-EXP-001 adds one benchmark-only Zustand/Vitest adapter and a frozen full-catalog oracle: both AV arms selected all 8 relevant checks observed across ten scenarios, while the uncertainty case forced full verification. It has no production adapter, independent qualifying replication, or trusted selective-verification class." + "status": "VERIFIED", + "assessment": "Tasks verification contract integrated. Research remains OBSERVE/SHADOW; permanent AV-EXP-002 failure and bounded AV-EXP-003 repair remain intact. Manifest-backed selection does not establish general trust." }, - "drift_status": "NEW_ALIGNED_PUBLIC_HOME", + "drift_status": "TASKS_CONTRACT_INTEGRATED_RESEARCH_LIMITS_RETAINED", "current_name_accuracy": "Accurate for the implemented verification-planning boundary and explicitly broader than affected-test selection.", "evidence_references": [ "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/SPEC.md", @@ -731,14 +776,60 @@ "https://github.com/opsle/affected-verification/blob/0544362d7659093b7f0b4f89ee8f68023fd269c3/benchmark/av-exp-001/preregistration-v1/preregistration.json", "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/REPORT.md", "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/summary.json", - "https://github.com/opsle/affected-verification/pull/2" + "https://github.com/opsle/affected-verification/pull/2", + "https://github.com/opsle/affected-verification/blob/792c4bb7881f6f430b2b1ba238ba05be43f50d38/README.md" + ], + "confidence": "HIGH", + "unresolved_questions": [ + "Which independent evidence could justify a bounded trust promotion beyond OBSERVE/SHADOW without erasing the permanent AV-EXP-002 failure?", + "Can opaque-boundary verification regain precision while preserving fail-closed selection and complete check evidence?", + "How generalizable is the bounded AV-EXP-003 repair beyond its frozen replay and adversarial corpora?" + ], + "source_repository": "affected-verification", + "highest_evidenced_stage": "VERIFIED", + "concept_disposition": "ACTIVE", + "historical_analysis": "program/history/pre-consolidation/theory-registry.json" + }, + { + "id": "visible-value", + "canonical_concept_name": "Visible Value", + "one_sentence_definition": "Validate mechanism receipts and summarize operator evidence without inventing benefit.", + "original_problem": "Mechanism activity needs inspectable evidence and honest measurement classes.", + "primary_classification": "INDEPENDENT_OPSLE_TOOL", + "current_repository": "visible-value", + "recommended_disposition": "KEEP_STANDALONE", + "disposition_rationale": "Independent reusable receipt validator and reporting CLI.", + "disposition_gain": "Shared validation and honest compatible totals.", + "disposition_risk": "Producer claims are not independent attestations.", + "provenance_concerns": "Preserve research Visible Value Contract attribution and Apache-2.0 license.", + "relationship_to_gearbox": "Consumes receipts without choosing routes.", + "relationship_to_context_firewall": "Consumes byte receipts without inferring downstream delivery or savings.", + "dependencies": [], + "consumers": [], + "current_implementation_fidelity": { + "status": "IMPLEMENTED", + "assessment": "Installable library and CLI with Tasks receipt observer integration." + }, + "drift_status": "Reconciled to current standalone implementation.", + "current_name_accuracy": "Accurate receipt-validation and presentation role.", + "evidence_references": [ + "https://github.com/opsle/visible-value/blob/f011395cae86a7a736b424d658bbe3afe65ba870/README.md" ], "confidence": "HIGH", "unresolved_questions": [ - "Does the zero-observed-miss result persist in a second public repository and different ecosystem?", - "Which benchmark-only evidence adapter is valuable enough to harden into a production-quality adapter?", - "Which evidence-based promotion criteria justify TRUSTED_BOUNDED authority for a narrowly defined change class?" - ] + "Controlled comparative benefit and independent artifact authentication are not established." + ], + "source_repository": "visible-value", + "highest_evidenced_stage": "PROTOTYPED", + "concept_disposition": "ACTIVE", + "historical_analysis": null } - ] + ], + "historical_state": "program/history/pre-consolidation/theory-registry.json", + "concept_counts": { + "ACTIVE": 7, + "CONSOLIDATED": 11, + "RETIRED": 1 + }, + "concept_count": 19 } diff --git a/tests/test_gearbox_publication_evidence.py b/tests/test_gearbox_publication_evidence.py index 65ef9c9..1ab4699 100644 --- a/tests/test_gearbox_publication_evidence.py +++ b/tests/test_gearbox_publication_evidence.py @@ -29,7 +29,7 @@ def test_consumers_bind_one_public_receipt(self): self.assertEqual(self.decision["input"], self.trajectory["input"]) self.assertRegex(self.decision["input"]["sha256"], r"^[0-9a-f]{64}$") self.assertIn( - self.repositories["gearbox"]["last_verified_head_sha"], + self.registry["gearbox_publication"]["final_main_sha"], self.decision["input"]["locator"], ) diff --git a/tests/test_post_consolidation.py b/tests/test_post_consolidation.py new file mode 100644 index 0000000..699f1de --- /dev/null +++ b/tests/test_post_consolidation.py @@ -0,0 +1,125 @@ +"""Reject topology regressions without reactivating archived research sources.""" +from __future__ import annotations + +import copy +import sys +import unittest +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[1] +sys.path.insert(0, str(ROOT / 'tools')) +from validate_program import DEFAULT_REGISTRY, DEFAULT_EXPERIMENTS, DEFAULT_THEORY_REGISTRY, load_json, validate, validate_theory +from render_program_status import render, render_priority, render_theory_map + + +class PostConsolidationTests(unittest.TestCase): + def setUp(self): + self.registry = load_json(DEFAULT_REGISTRY) + self.experiments = load_json(DEFAULT_EXPERIMENTS) + self.theory = load_json(DEFAULT_THEORY_REGISTRY) + + def reject(self, fragment): + self.assertTrue(any(fragment in e for e in validate(self.registry, self.experiments))) + + def test_retired_system_cannot_be_current_authority(self): + self.registry['program_control']['workload_authority'] = 'paperclip' + self.reject('current workload authority') + + def test_retired_system_cannot_be_next_action_or_gate(self): + for field in ('exact_next_execution', 'operating_question'): + with self.subTest(field=field): + registry = copy.deepcopy(self.registry) + registry['program_control'][field] = 'Wait for Durable Supervisor v0.1.' + self.assertTrue(any('retired system' in e for e in validate(registry, self.experiments))) + + def test_current_repo_cannot_depend_on_retired_source(self): + repo = next(r for r in self.registry['repositories'] if r['name'] == 'research') + repo['dependencies'] = ['agent-state-ledger'] + self.reject('current relationship targets historical') + + def test_retired_repo_cannot_gain_work(self): + repo = next(r for r in self.registry['repositories'] if r['name'] == 'durable-supervisor') + repo['next_task'] = 'Implement a new measurement.' + self.reject('retired repository has active work') + + def test_dangling_consolidation_destination_fails(self): + repo = next(r for r in self.registry['repositories'] if r['name'] == 'agent-state-ledger') + repo['consolidated_into'] = 'missing-home' + self.reject('dangling consolidation destination') + + def test_repository_count_drift_fails(self): + self.registry['repository_counts']['current'] += 1 + self.reject('repository count drift') + + def test_retired_membership_cannot_be_deleted_with_counts(self): + self.registry['membership']['historical'].remove('durable-supervisor') + self.registry['repository_counts']['historical'] -= 1 + self.reject('historical membership differs') + + def test_github_housekeeping_is_not_workload_eligibility(self): + repo = next(r for r in self.registry['repositories'] if r['name'] == '.github') + self.assertEqual(repo['repository_disposition'], 'ACTIVE') + self.assertFalse(repo['workload_eligible']) + repo['workload_eligible'] = True + self.reject('workload eligibility drift') + + def test_graphify_capability_claim_fails(self): + self.registry['optional_tools']['graphify']['tasks_capability'] = True + self.reject('Graphify') + + def test_graphify_cannot_enter_tasks_capability_inventory(self): + self.registry['program_control']['opsle_tasks']['capability_ids'].append('opsle.graphify') + self.reject('Graphify is standalone only') + + def test_retirement_cannot_promote_repository_maturity(self): + repo = next(r for r in self.registry['repositories'] if r['name'] == 'agent-state-ledger') + repo['lifecycle_stage'] = 'PROTOTYPED' + self.reject('retirement must preserve highest evidenced maturity') + + def test_missing_inventory_object_reports_error(self): + self.registry['membership'] = None + self.reject('registry.membership must be an object') + + def test_concept_count_drift_fails(self): + self.theory['concept_counts']['CONSOLIDATED'] -= 1 + errors = validate_theory(self.theory, self.registry, render_theory_map(self.theory, self.registry)) + self.assertTrue(any('disposition count drift' in e for e in errors)) + + def test_retired_concept_cannot_gain_a_current_home(self): + concept = next(c for c in self.theory['concepts'] if c['id'] == 'durable-supervisor') + concept['current_repository'] = 'tasks' + errors = validate_theory(self.theory, self.registry, render_theory_map(self.theory, self.registry)) + self.assertTrue(any('current home drift' in e for e in errors)) + + def test_historical_experiment_participation_survives(self): + repo = next(r for r in self.registry['repositories'] if r['name'] == 'verifiable-agent-handoff') + self.assertEqual(repo['active_experiment_ids'], []) + self.assertIn('EXP-001', repo['historical_experiment_ids']) + repo['historical_experiment_ids'] = [] + self.reject('does not reciprocally list') + + def test_experiment_readiness_and_failed_verdict_are_preserved(self): + exp = next(e for e in self.experiments['experiments'] if e['id'] == 'EXP-001') + self.assertEqual(exp['status'], 'PLANNED') + self.assertEqual(exp['run_identities'], []) + failed = next(e for e in self.experiments['experiments'] if e['id'] == 'AV-EXP-002') + self.assertIn('FAIL', failed['verdict']) + + def test_all_generated_views_are_deterministic(self): + views = ( + (render, (self.registry, self.experiments), ROOT / 'PROGRAM_STATUS.md'), + (render_priority, (self.registry, self.experiments), ROOT / 'program/PRIORITY.md'), + (render_theory_map, (self.theory, self.registry), ROOT / 'program/THEORY_MAP.md'), + ) + for renderer, args, path in views: + with self.subTest(path=path.name): + self.assertEqual(renderer(*args), renderer(*copy.deepcopy(args))) + self.assertEqual(renderer(*args), path.read_text()) + priority = render_priority(self.registry, self.experiments) + self.assertNotIn('Durable Supervisor', priority) + self.assertNotIn('stopping criteria', priority) + self.assertIn('current workload and task-management authority', priority) + + +if __name__ == '__main__': + unittest.main() diff --git a/tests/test_theory_registry.py b/tests/test_theory_registry.py index de61d85..bb9d02f 100644 --- a/tests/test_theory_registry.py +++ b/tests/test_theory_registry.py @@ -34,12 +34,15 @@ def errors_for(self, theory=None, registry=None, theory_map=None): def test_authoritative_theory_registry_is_valid(self): self.assertEqual(self.errors_for(), []) - def test_every_current_concept_repository_is_mapped_once(self): + def test_consolidated_concepts_share_a_home(self): + concepts = [c for c in self.theory["concepts"] if c["current_repository"] == "tasks"] + self.assertGreater(len(concepts), 1) + self.assertEqual(self.errors_for(), []) + + def test_wrong_current_home_fails(self): theory = copy.deepcopy(self.theory) theory["concepts"][-1]["current_repository"] = "context-firewall" - errors = self.errors_for(theory=theory) - self.assertTrue(any("duplicate current concept repository mappings" in error for error in errors)) - self.assertTrue(any("missing current concept repository mappings" in error for error in errors)) + self.assertTrue(any("current home drift" in e for e in self.errors_for(theory=theory))) def test_gearbox_must_claim_its_current_repository(self): theory = copy.deepcopy(self.theory) @@ -77,39 +80,24 @@ def test_required_semantic_text_cannot_be_blank(self): errors = self.errors_for(theory=theory) self.assertTrue(any("original_problem must be a nonempty string" in error for error in errors)) - def test_registry_reconciliation_distinction_is_required(self): - registry = copy.deepcopy(self.registry) - del registry["theory_reconciliation"]["gearbox_vs_durable_supervisor"] - errors = self.errors_for(registry=registry) - self.assertTrue( - any("Gearbox versus Durable Supervisor distinction" in error for error in errors) - ) - - def test_registry_gearbox_repository_status_is_required(self): + def test_post_consolidation_reconciliation_is_required(self): registry = copy.deepcopy(self.registry) - del registry["theory_reconciliation"]["gearbox_repository_status"] - errors = self.errors_for(registry=registry) - self.assertTrue( - any("created and prototyped" in error for error in errors) - ) + registry["theory_reconciliation"]["status"] = "IMPLEMENTED_HOME_REGISTERED" + self.assertTrue(any("post-consolidation state" in e for e in self.errors_for(registry=registry))) def test_gearbox_publication_head_must_match_registry(self): registry = copy.deepcopy(self.registry) registry["gearbox_publication"]["final_main_sha"] = "0" * 40 errors = self.errors_for(registry=registry) self.assertTrue( - any("publication SHA must match repository HEAD" in error for error in errors) + any("publication SHA must match historical publication" in error for error in errors) ) - def test_existing_repository_lifecycle_changes_are_forbidden(self): - registry = copy.deepcopy(self.registry) - registry["theory_reconciliation"][ - "existing_repository_lifecycle_changes_executed" - ] = True - errors = self.errors_for(registry=registry) - self.assertTrue( - any("no existing-repository lifecycle changes" in error for error in errors) - ) + def test_consolidation_does_not_promote_concept_maturity(self): + theory = copy.deepcopy(self.theory) + concept = next(c for c in theory["concepts"] if c["id"] == "agent-state-ledger") + concept["highest_evidenced_stage"] = "COMPLETE" + self.assertTrue(any("research maturity drift" in e for e in self.errors_for(theory=theory))) def test_theory_map_classification_drift_fails(self): drifted = self.theory_map.replace( diff --git a/tests/test_validate_program.py b/tests/test_validate_program.py index 6a6d980..a727da3 100644 --- a/tests/test_validate_program.py +++ b/tests/test_validate_program.py @@ -29,14 +29,14 @@ def errors_for(self, registry=None, experiments=None): def test_authoritative_registries_are_valid(self): self.assertEqual(self.errors_for(), []) - def test_affected_verification_is_the_twenty_first_repository(self): - self.assertEqual(len(self.registry["repositories"]), 21) + def test_current_inventory_preserves_affected_verification_experiments(self): + self.assertEqual(len(self.registry["repositories"]), self.registry["authoritative_repository_count"]) gearbox = next( item for item in self.registry["repositories"] if item["name"] == "gearbox" ) self.assertEqual(gearbox["lifecycle_stage"], "PROTOTYPED") - self.assertEqual( + self.assertNotEqual( gearbox["last_verified_head_sha"], self.registry["gearbox_publication"]["final_main_sha"], ) @@ -47,7 +47,7 @@ def test_affected_verification_is_the_twenty_first_repository(self): self.assertEqual(affected["lifecycle_stage"], "VERIFIED") self.assertEqual( affected["last_verified_head_sha"], - "97f490a67337552fee25757266f3dc034660dca0", + "792c4bb7881f6f430b2b1ba238ba05be43f50d38", ) self.assertEqual( affected["active_experiment_ids"], @@ -130,13 +130,13 @@ def test_duplicate_repository_fails(self): def test_priority_lanes_cover_each_repository_exactly_once(self): lanes = self.registry["program_control"]["lanes"] repositories = [name for lane in lanes for name in lane["repositories"]] - self.assertEqual(len(repositories), 21) - self.assertEqual(len(set(repositories)), 21) + self.assertEqual(len(repositories), len(self.registry["membership"]["current"])) + self.assertEqual(set(repositories), set(self.registry["membership"]["current"])) def test_duplicate_priority_repository_fails(self): registry = copy.deepcopy(self.registry) registry["program_control"]["lanes"][1]["repositories"].append( - "durable-supervisor" + "tasks" ) errors = self.errors_for(registry=registry) self.assertTrue( @@ -161,13 +161,10 @@ def test_active_repositories_match_now_lane(self): errors = self.errors_for(registry=registry) self.assertTrue(any("current priority lane" in error for error in errors)) - def test_durable_supervisor_stopping_criteria_are_fenced(self): + def test_retired_supervisor_controls_cannot_return(self): registry = copy.deepcopy(self.registry) - registry["program_control"]["durable_supervisor_v0_1"][ - "stopping_criteria" - ].pop() - errors = self.errors_for(registry=registry) - self.assertTrue(any("ten ordered stopping criteria" in error for error in errors)) + registry["program_control"]["durable_supervisor_v0_1"] = {"status": "IN_PROGRESS"} + self.assertTrue(any("retired system" in e for e in self.errors_for(registry=registry))) def test_anti_nitpick_reasons_are_fenced(self): registry = copy.deepcopy(self.registry) diff --git a/tools/render_program_status.py b/tools/render_program_status.py index 79ce90b..92326c4 100644 --- a/tools/render_program_status.py +++ b/tools/render_program_status.py @@ -13,7 +13,7 @@ DEFAULT_REGISTRY, DEFAULT_THEORY_MAP, DEFAULT_THEORY_REGISTRY, - EXPECTED_REPOSITORIES, + _canonical_json_sha256, LIFECYCLE_STAGES, load_json, validate, @@ -32,7 +32,7 @@ def _cell(value: str) -> str: def _validate_for_render(registry: dict, experiments: dict) -> dict: errors = validate(registry, experiments) theory = load_json(DEFAULT_THEORY_REGISTRY) - theory_map_text = DEFAULT_THEORY_MAP.read_text(encoding="utf-8") + theory_map_text = render_theory_map(theory, registry) errors.extend(validate_theory(theory, registry, theory_map_text)) if errors: raise ValueError("cannot render invalid registry: " + "; ".join(errors)) @@ -57,7 +57,6 @@ def render_priority(registry: dict, experiments: dict) -> str: _validate_for_render(registry, experiments) control = registry["program_control"] admission = control["work_item_admission"] - ds = control["durable_supervisor_v0_1"] tasks = control["opsle_tasks"] visible = registry["visible_value"] lines = [ @@ -65,7 +64,7 @@ def render_priority(registry: dict, experiments: dict) -> str: "", "", "", - "`program/registry.json` is the sole priority authority. Edit the registry, not this file.", + "`program/registry.json` records research priorities. Opsle Tasks is the current workload and task-management authority. Edit the registry, not this generated view.", "", "## Operating question", "", @@ -95,28 +94,12 @@ def render_priority(registry: dict, experiments: dict) -> str: ] ) - lines.extend( - [ - "", - "## Durable Supervisor v0.1 stopping criteria", - "", - f"Status: **{ds['status']}**. Foundation: `{ds['durability_foundation']}`.", - "", - f"Verified main: `{ds['verified_main_sha']}`. Runtime: `{ds['verified_runtime_state']}` " - f"as durably recorded at `{ds['runtime_state_recorded_at']}`.", - "", - ] - ) - lines.extend( - f"{index}. **{item['status']}** — {item['criterion']}" - for index, item in enumerate(ds["stopping_criteria"], 1) - ) lines.extend( [ "", "## Opsle Tasks boundary", "", - f"{tasks['future_name']} is the {tasks['role']}. Its current repository remains " + f"{tasks['name']} is the {tasks['role']}. Its repository is " f"`{tasks['current_repository']}`.", "", "Measure: " + ", ".join(tasks["measurements"]) + ".", @@ -152,8 +135,8 @@ def render_priority(registry: dict, experiments: dict) -> str: "", "Per-child receipt: " + "; ".join(visible["per_child_receipt_fields"]) + ".", "", - "Supervisor/run summary: " - + "; ".join(visible["supervisor_summary_fields"]) + "Task/run summary: " + + "; ".join(visible["run_summary_fields"]) + ".", "", "## Later", @@ -179,7 +162,6 @@ def render(registry: dict, experiments: dict) -> str: theory = _validate_for_render(registry, experiments) repositories = registry["repositories"] control = registry["program_control"] - ds = control["durable_supervisor_v0_1"] stages = Counter(repo["lifecycle_stage"] for repo in repositories) states = Counter(repo["program_state"] for repo in repositories) experiment_by_id = {item["id"]: item for item in experiments["experiments"]} @@ -188,9 +170,6 @@ def render(registry: dict, experiments: dict) -> str: for concept in theory["concepts"] if concept["current_repository"] is not None ] - satisfied = sum( - item["status"] == "SATISFIED" for item in ds["stopping_criteria"] - ) lines = [ "# Opsle program status", "", @@ -198,7 +177,7 @@ def render(registry: dict, experiments: dict) -> str: "program/experiments.json, program/theory-registry.json, and " "program/THEORY_MAP.md. -->", "", - f"**Coverage: {len(repositories)}/{len(EXPECTED_REPOSITORIES)} expected repositories; duplicates: 0.**", + f"**Coverage: {len(repositories)} recorded repositories; {registry['repository_counts']['current']} current, {registry['repository_counts']['historical']} historical; {registry['repository_counts']['workload_eligible']} workload eligible.**", "", f"Last verified: `{registry['last_verified_at']}`. HEADs are the verified default-branch revisions, not an assumption about later changes.", "", @@ -206,8 +185,7 @@ def render(registry: dict, experiments: dict) -> str: "", f"> {control['operating_question']}", "", - f"Current lane: **{control['current_lane']}**. Durable Supervisor v0.1: " - f"**{ds['status']}** ({satisfied}/{len(ds['stopping_criteria'])} stopping criteria satisfied).", + f"Current lane: **{control['current_lane']}**. Workload and task-management authority: **Opsle Tasks**.", "", "Full generated priority view: `program/PRIORITY.md`.", "", @@ -222,7 +200,7 @@ def render(registry: dict, experiments: dict) -> str: "", control["exact_next_execution"], "", - "## Portfolio totals", + "## Highest evidenced research maturity (including historical sources)", "", "| Lifecycle stage | Count |", "|---|---:|", @@ -233,11 +211,11 @@ def render(registry: dict, experiments: dict) -> str: lines.extend( [ "", - f"Program state totals: active {states.get('active', 0)}; waiting {states.get('waiting', 0)}; complete {states.get('complete', 0)}.", + f"Program state totals: active {states.get('active', 0)}; waiting {states.get('waiting', 0)}; complete {states.get('complete', 0)}; historical {states.get('historical', 0)}.", "", "## Repository inventory", "", - "| # | Repository | Type | HEAD | Stage | Evidence | Blocker | Next task | Dependencies | State |", + "| # | Repository | Type | HEAD | Stage | Evidence | Blocker | Next task | Dependencies | State / disposition |", "|---:|---|---|---|---|---|---|---|---|---|", ] ) @@ -262,19 +240,23 @@ def render(registry: dict, experiments: dict) -> str: _cell(blocker), _cell(repo["next_task"]), dependencies, - repo["program_state"], + repo["program_state"] + " / " + repo["repository_disposition"] + (" → " + repo["consolidated_into"] if repo["consolidated_into"] else ""), ] ) + " |" ) + lines.extend(["", "## Completed integration and retirement work", ""]) + for item in registry["completed_work"]: + lines.append(f"- **{item['id']} — {item['status']}**: {item['scope']} " + ", ".join(f"[evidence {i}]({ref})" for i, ref in enumerate(item["evidence"], 1))) + lines.extend(["", "Graphify is an optional standalone CLI, not an Opsle Tasks capability. Historical capability acceptance is superseded.", "", "Historical source evidence and superseded controls: `program/history/pre-consolidation/`. Retirement is not research completion."]) exp001 = experiment_by_id["EXP-001"] lines.extend( [ "", "## Theory reconciliation", "", - f"Concept coverage: {len(theory['concepts'])} canonical concepts; {len(current_concepts)} current concept repositories mapped exactly once; Agent Gearbox maps to opsle/gearbox.", + f"Concept coverage: {len(theory['concepts'])} canonical concepts; {len(current_concepts)} live concepts sharing {theory['current_concept_repository_count']} current homes; Agent Gearbox maps to opsle/gearbox.", "", "Canonical map: `program/THEORY_MAP.md`. Machine registry: `program/theory-registry.json`.", "", @@ -286,7 +268,8 @@ def render(registry: dict, experiments: dict) -> str: "", "## Mechanical source of truth", "", - "- Portfolio and priority authority: `program/registry.json`", + "- Research portfolio and priority evidence: `program/registry.json`", + "- Workload and task-management authority: Opsle Tasks (`opsle/tasks`)", "- Generated priority view: `program/PRIORITY.md`", "- Experiments: `program/experiments.json`", "- Theory registry: `program/theory-registry.json`", @@ -297,6 +280,37 @@ def render(registry: dict, experiments: dict) -> str: "", ] ) + lines.extend(["", "### All recorded experiments", "", "| Experiment | Status | Verdict |", "|---|---|---|"]) + for experiment in experiments["experiments"]: + lines.append(f"| `{experiment['id']}` | {experiment['status']} | {_cell(experiment['verdict'])} |") + lines.append("") + return "\n".join(lines) + + +def render_theory_map(theory: dict, registry: dict) -> str: + lines = [ + "# Current theory map", "", + "", "", + "Opsle Tasks owns current workload and task management. The program ledger records research state; it does not create a parallel task queue.", "", + "Gearbox determines where bounded work executes; Context Firewall determines what evidence returns. Visible Value validates receipts and presents operator measurements. Affected Verification plans checks under explicit evidence boundaries.", "", + "Task 15 completed the versioned capability contract: trusted manifests, operator grants, immutable repository selection, lifecycle hooks, receipts and an execution-scoped model gateway. Graphify is an optional standalone CLI, not a Tasks capability.", "", + "The final consolidation receipt supersedes earlier partial observations. Durable Supervisor is retired, with no current home or activation gate. Other consolidated concepts share Tasks or Gearbox; their source maturity does not increase through migration. `.github` remains a repository but is ineligible for workload execution.", "", + "Historical definitions, extraction analysis, initial SHAs, licenses, negative evidence and superseded recommendations are preserved in [the historical map](history/pre-consolidation/THEORY_MAP.md) and [historical registry](history/pre-consolidation/theory-registry.json). They authorize no work.", "", + f"Coverage: {theory['concept_count']} concepts; " + "; ".join(f"{key.lower()} {value}" for key, value in theory["concept_counts"].items()) + f"; {theory['current_concept_repository_count']} distinct current homes.", "", + "Theory registry canonical SHA-256:", f"`{_canonical_json_sha256(theory)}`.", "", + "| Concept identity | Classification | Executed / retained disposition | Confidence | Current home and maturity |", + "|---|---|---|---|---|", + ] + for concept in theory["concepts"]: + home = concept["current_repository"] or "none (retired)" + lines.append(f"| `{concept['id']}` | `{concept['primary_classification']}` | `{concept['recommended_disposition']}` | `{concept['confidence']}` | {home}; {concept['highest_evidenced_stage']} |") + lines.extend(["", "## Evidence limits", "", + "Context Firewall remains PROTOTYPED: flat TAP plus Node spec/dot reporters and a semantic-only projection. Canonical packet metrics are audit measurements; downstream delivery is measured separately. This is not an arbitrary-log parser or production security boundary.", "", + "Gearbox remains PROTOTYPED with provider-free bounded execution and deterministic phase routing. Routing Policy is consolidated into its documented contract. Capability integration does not prove comparative savings.", "", + "Visible Value is PROTOTYPED: an independent validator and reporting CLI. It does not authenticate external producer evidence. Missing measurements remain absent and incompatible totals are excluded.", "", + "Affected Verification remains VERIFIED, with research limited to OBSERVE/SHADOW. AV-EXP-002 remains FAIL; AV-EXP-003 repaired its known skip without establishing general selector completeness. Tasks may use bounded manifest-backed selection only with complete impact, catalog and check-boundary evidence; unknown evidence broadens or stops. This is not a general TRUSTED_BOUNDED promotion.", "", + "EXP-001 remains PLANNED and unconsumed. Remote Incus execution proofs are operational observations; Task 15 implementation and repository consolidation do not complete research or establish token, cost, correctness or causal savings.", "", + "Revision-linked completed work and verified inventory: [PROGRAM_STATUS](../PROGRAM_STATUS.md), [evidence](evidence/post-consolidation/README.md).", ""]) return "\n".join(lines) @@ -310,6 +324,7 @@ def main(argv: list[str] | None = None) -> int: args = parser.parse_args(argv) registry = load_json(DEFAULT_REGISTRY) experiments = load_json(DEFAULT_EXPERIMENTS) + theory_content = render_theory_map(load_json(DEFAULT_THEORY_REGISTRY), registry) try: content = render(registry, experiments) priority_content = render_priority(registry, experiments) @@ -318,6 +333,8 @@ def main(argv: list[str] | None = None) -> int: return 1 if args.check: stale = [] + if not DEFAULT_THEORY_MAP.exists() or DEFAULT_THEORY_MAP.read_text(encoding="utf-8") != theory_content: + stale.append(str(DEFAULT_THEORY_MAP)) if not args.output.exists() or args.output.read_text(encoding="utf-8") != content: stale.append(str(args.output)) if ( @@ -334,12 +351,13 @@ def main(argv: list[str] | None = None) -> int: ) return 1 print( - f"PASS: {args.output.name} and {args.priority_output.name} match the registry" + f"PASS: {args.output.name}, {args.priority_output.name}, and {DEFAULT_THEORY_MAP.name} match the registry" ) return 0 + DEFAULT_THEORY_MAP.write_text(theory_content, encoding="utf-8") args.output.write_text(content, encoding="utf-8") args.priority_output.write_text(priority_content, encoding="utf-8") - print(f"Rendered {args.output} and {args.priority_output}") + print(f"Rendered {args.output}, {args.priority_output}, and {DEFAULT_THEORY_MAP}") return 0 diff --git a/tools/validate_program.py b/tools/validate_program.py index 9f7df6e..5d6b771 100644 --- a/tools/validate_program.py +++ b/tools/validate_program.py @@ -19,29 +19,21 @@ DEFAULT_THEORY_REGISTRY = ROOT / "program" / "theory-registry.json" DEFAULT_THEORY_MAP = ROOT / "program" / "THEORY_MAP.md" -EXPECTED_REPOSITORIES = ( - "agent-trajectory-profiler", - "semantic-edit-protocol", - "durable-supervisor", - "event-driven-agent-wakeup", - "context-firewall", - "decision-evidence-protocol", - "agent-state-ledger", - "agent-scheduler-runtime", - "verifiable-agent-handoff", - "agent-routing-policy", - "agent-resource-claims", - "agent-discovery-control", - "agent-execution-authorization", - "controlled-agent-acceptance", - "agent-recovery-policy", - "ephemeral-agent-workers", - "gearbox", - "affected-verification", - "research", - "site", - ".github", -) +DEFAULT_INVENTORY = ROOT / "program/evidence/post-consolidation/inventory.json" +REPOSITORY_DISPOSITIONS = {"ACTIVE", "CONSOLIDATED", "RETIRED"} +RETIRED_SYSTEMS = ("durable-supervisor", "durable supervisor", "taslos tasks", "paperclip", "agent-run") + + +def retired_reference(value: Any) -> bool: + text = json.dumps(value, ensure_ascii=False).lower() + return any(name in text for name in RETIRED_SYSTEMS) + + +def membership(registry: dict, kind: str) -> set[str]: + value = registry.get("membership", {}) + values = value.get(kind, []) if isinstance(value, dict) else [] + return {v for v in values if isinstance(v, str)} if isinstance(values, list) else set() + LIFECYCLE_STAGES = ( "THEORY", @@ -91,7 +83,7 @@ } ) -SUPERVISOR_VALUE_FIELDS = frozenset( +RUN_VALUE_FIELDS = frozenset( { "total children", "model/effort distribution", @@ -110,8 +102,7 @@ { "Gearbox", "Context Firewall", - "Decision Evidence Protocol", - "Agent Trajectory Profiler", + "Visible Value", "Affected Verification", } ) @@ -157,12 +148,17 @@ "CONSOLIDATE_WITH_OTHER", "RENAME", "DEPRECATE_AFTER_PROVENANCE_PRESERVATION", + "CONSOLIDATED", + "RETIRED", } ) THEORY_CONFIDENCE_LEVELS = frozenset({"HIGH", "MEDIUM_HIGH", "MEDIUM", "LOW"}) REQUIRED_THEORY_CONCEPT_FIELDS = ( + "source_repository", + "concept_disposition", + "highest_evidenced_stage", "id", "canonical_concept_name", "one_sentence_definition", @@ -198,12 +194,6 @@ "and escalation path the agent needs to make correct decisions." ) -GEARBOX_VS_DURABLE_SUPERVISOR = ( - "Gearbox enhances a primary developer; durable orchestration owns autonomous " - "objective progress across durable state and multiple activations without " - "requiring that developer to remain continuously active." -) - EXP001_CONTEXT_FIREWALL_SCOPE = ( "The deterministic multi-adapter evidence boundary is the concept; the current " "TAP-subset reducer is one adapter prototype." @@ -269,6 +259,10 @@ ) REQUIRED_REPOSITORY_FIELDS = ( + "repository_disposition", + "consolidated_into", + "workload_eligible", + "historical_experiment_ids", "name", "github_url", "default_branch", @@ -405,7 +399,59 @@ def validate( if not isinstance(experiment_entries, list): return ["experiments.experiments must be an array"] - expected = set(EXPECTED_REPOSITORIES) + if registry.get("schema_version") != 2: + errors.append("registry.schema_version must be 2") + if registry.get("inventory_evidence") != "program/evidence/post-consolidation/inventory.json": + errors.append("verified inventory path drift") + if registry.get("inventory_evidence_sha256") != "sha256:" + hashlib.sha256(DEFAULT_INVENTORY.read_bytes()).hexdigest(): + errors.append("verified inventory hash drift") + consolidation_path = ROOT / "program/evidence/post-consolidation/consolidation.json" + if registry.get("consolidation_evidence") != { + "path": "program/evidence/post-consolidation/consolidation.json", + "sha256": "sha256:" + hashlib.sha256(consolidation_path.read_bytes()).hexdigest(), + }: + errors.append("final consolidation receipt identity drift") + receipt = load_json(consolidation_path) + if receipt.get("overall") != "complete" or not receipt.get("completedAt"): + errors.append("final consolidation receipt must be complete") + historical_snapshot = load_json(ROOT / "program/history/pre-consolidation/registry.json") + historical_repos = {r["name"]: r for r in historical_snapshot["repositories"]} + milestones = registry.get("completed_work", []) + required_milestones = {"taslos-tasks-retirement", "paperclip-retirement", "repository-consolidation", "remote-incus-execution", "task-15-capability-contract"} + if (not isinstance(milestones, list) or + {m.get("id") for m in milestones if isinstance(m, dict) and isinstance(m.get("id"), str)} != required_milestones + or len(milestones) != len(required_milestones)): + errors.append("completed integration and retirement milestones drifted") + elif any(not isinstance(m, dict) or m.get("status") != "COMPLETED" + or not _string_list(m.get("evidence"), allow_empty=False) + or not _nonempty_string(m.get("scope")) for m in milestones): + errors.append("completed milestones require status, scope and evidence") + current = membership(registry, "current") + historical = membership(registry, "historical") + expected = current | historical + inventory = load_json(DEFAULT_INVENTORY) + if inventory["consolidation_destinations"] != {item["source"]: item["destination"] for item in receipt["sources"]}: + errors.append("inventory consolidation destinations differ from final receipt") + membership_record = registry.get("membership") + if not isinstance(membership_record, dict): + errors.append("registry.membership must be an object") + membership_record = {} + for kind, recorded in (("current", current), ("historical", historical)): + pinned = inventory["active_repositories" if kind == "current" else "historical_repositories"] + raw = membership_record.get(kind, []) + if not _string_list(raw, allow_empty=False) or len(raw) != len(recorded) or recorded != set(pinned): + errors.append(f"{kind} membership differs from verified inventory") + if current & historical: + errors.append("current and historical membership must be disjoint") + if registry.get("repository_counts") != { + "current": len(current), "historical": len(historical), + "workload_eligible": len(current - set(inventory["workload_ineligible"])), + }: + errors.append("repository count drift") + optional_tools = registry.get("optional_tools") + if not isinstance(optional_tools, dict) or optional_tools.get("graphify") != inventory["graphify"]: + errors.append("Graphify must remain an optional standalone CLI, never a Tasks capability") + names: list[str] = [] for index, repo in enumerate(repositories): if not isinstance(repo, dict): @@ -420,13 +466,13 @@ def validate( missing = sorted(expected - set(names)) unexpected = sorted(set(names) - expected, key=str) - if len(repositories) != len(EXPECTED_REPOSITORIES): + if len(repositories) != len(expected): errors.append( - f"expected {len(EXPECTED_REPOSITORIES)} repositories, found {len(repositories)}" + f"expected {len(expected)} repositories, found {len(repositories)}" ) - if registry.get("authoritative_repository_count") != len(EXPECTED_REPOSITORIES): + if registry.get("authoritative_repository_count") != len(expected): errors.append( - f"authoritative_repository_count must be {len(EXPECTED_REPOSITORIES)}" + f"authoritative_repository_count must be {len(expected)}" ) if duplicates: errors.append(f"duplicate repositories: {', '.join(duplicates)}") @@ -477,8 +523,8 @@ def validate( if not isinstance(control, dict): errors.append("registry.program_control must be an object") else: - if control.get("schema_version") != 1: - errors.append("program_control.schema_version must be 1") + if control.get("schema_version") != 2: + errors.append("program_control.schema_version must be 2") if control.get("priority_order") != list(PRIORITY_LANES): errors.append("program_control.priority_order must be NOW, NEXT, THEN, LATER, PARKED") if control.get("current_lane") != "NOW": @@ -535,8 +581,8 @@ def validate( "program_control duplicate priority repositories: " + ", ".join(duplicate_priority_repositories) ) - missing_priority_repositories = sorted(expected - set(priority_repositories)) - unexpected_priority_repositories = sorted(set(priority_repositories) - expected) + missing_priority_repositories = sorted(current - set(priority_repositories)) + unexpected_priority_repositories = sorted(set(priority_repositories) - current) if missing_priority_repositories: errors.append( "program_control missing priority repositories: " @@ -553,36 +599,10 @@ def validate( if active_repositories != sorted(current_lane_repositories): errors.append("active repositories must exactly match the current priority lane") - ds = control.get("durable_supervisor_v0_1") - if not isinstance(ds, dict): - errors.append("program_control.durable_supervisor_v0_1 must be an object") - else: - if ds.get("status") not in {"IN_PROGRESS", "DECLARED_FROZEN"}: - errors.append("durable_supervisor_v0_1 has invalid status") - if ds.get("verified_runtime_state") != "PAUSED_NO_ACTIVE_TASK_OR_ATTEMPT": - errors.append("durable_supervisor_v0_1 runtime state drifted") - if not _valid_timestamp(ds.get("runtime_state_recorded_at")): - errors.append("durable_supervisor_v0_1 runtime state needs a recorded timestamp") - if not _nonempty_string(ds.get("runtime_state_source")): - errors.append("durable_supervisor_v0_1 runtime observation needs a source") - durable_repository = repo_by_name.get("durable-supervisor", {}) - if ds.get("verified_main_sha") != durable_repository.get("last_verified_head_sha"): - errors.append("durable_supervisor_v0_1 SHA must match the repository ledger") - criteria = ds.get("stopping_criteria") - if not isinstance(criteria, list): - errors.append("durable_supervisor_v0_1.stopping_criteria must be an array") - criteria = [] - expected_ids = [f"DS-V0.1-{index:02d}" for index in range(1, 11)] - if [item.get("id") for item in criteria if isinstance(item, dict)] != expected_ids: - errors.append("durable_supervisor_v0_1 must retain ten ordered stopping criteria") - for item in criteria: - if not isinstance(item, dict): - errors.append("durable_supervisor_v0_1 criteria must be objects") - continue - if item.get("status") not in {"OPEN", "SATISFIED"}: - errors.append(f"{item.get('id')}: invalid stopping-criterion status") - if not _nonempty_string(item.get("criterion")): - errors.append(f"{item.get('id')}: criterion must be nonempty") + if "durable_supervisor_v0_1" in control or retired_reference(control): + errors.append("retired system in current program control") + if control.get("workload_authority") != "opsle/tasks": + errors.append("Opsle Tasks must be the current workload authority") opsle_tasks = control.get("opsle_tasks") if not isinstance(opsle_tasks, dict): @@ -590,8 +610,14 @@ def validate( else: if opsle_tasks.get("current_repository") != "opsle/tasks": errors.append("Opsle Tasks current repository identity drifted") - if "NEXT primary real-world workload" not in str(opsle_tasks.get("role")): - errors.append("Opsle Tasks must remain the NEXT primary real-world workload") + if opsle_tasks.get("role") != "current workload and task-management authority": + errors.append("Opsle Tasks must be the current workload and task-management authority") + expected_capabilities = {"opsle.gearbox", "opsle.context-firewall", "opsle.affected-verification", "opsle.visible-value"} + capability_ids = opsle_tasks.get("capability_ids") + if (not _string_list(capability_ids, allow_empty=False) + or set(capability_ids) != expected_capabilities + or len(capability_ids) != len(expected_capabilities)): + errors.append("Tasks capability inventory drift; Graphify is standalone only") measurements = opsle_tasks.get("measurements") if ( not _string_list(measurements, allow_empty=False) @@ -621,7 +647,7 @@ def validate( errors.append( f"concept_activation index {index} repositories must be nonempty strings" ) - elif not set(activation_repositories).issubset(expected): + elif not set(activation_repositories).issubset(current): errors.append( f"concept_activation index {index} references an unknown repository" ) @@ -631,17 +657,6 @@ def validate( or set(later_items) != LATER_ITEMS ): errors.append("program_control later items drifted") - parked_items = control.get("parked_items") - if _string_list(parked_items, allow_empty=False): - required_parked_fragments = ( - "src/cli.js", - "Background projection", - "Historical pre-fix", - "architectural polishing", - ) - for fragment in required_parked_fragments: - if not any(fragment in item for item in parked_items): - errors.append(f"program_control parked item missing {fragment}") visible_value = registry.get("visible_value") if not isinstance(visible_value, dict): @@ -668,12 +683,12 @@ def validate( or set(child_fields) != PER_CHILD_VALUE_FIELDS ): errors.append("visible_value per-child receipt fields drifted") - summary_fields = visible_value.get("supervisor_summary_fields") + summary_fields = visible_value.get("run_summary_fields") if ( not _string_list(summary_fields, allow_empty=False) - or set(summary_fields) != SUPERVISOR_VALUE_FIELDS + or set(summary_fields) != RUN_VALUE_FIELDS ): - errors.append("visible_value supervisor summary fields drifted") + errors.append("visible_value run summary fields drifted") for index, repo in enumerate(repositories): label = ( @@ -690,7 +705,7 @@ def validate( errors.append(f"{label}: invalid lifecycle stage {repo.get('lifecycle_stage')!r}") if repo.get("project_type") not in {"concept", "program infrastructure"}: errors.append(f"{label}: invalid project_type {repo.get('project_type')!r}") - if repo.get("program_state") not in {"active", "waiting", "complete"}: + if repo.get("program_state") not in {"active", "waiting", "complete", "historical"}: errors.append(f"{label}: invalid program_state {repo.get('program_state')!r}") if repo.get("completion_status") not in { "INCOMPLETE", @@ -711,6 +726,7 @@ def validate( "dependencies", "dependents", "active_experiment_ids", + "historical_experiment_ids", "blockers", "evidence", "completion_criteria", @@ -754,6 +770,46 @@ def validate( if experiment_id not in experiment_id_set: errors.append(f"{label}: references nonexistent experiment {experiment_id}") + disposition = repo.get("repository_disposition") + if disposition not in REPOSITORY_DISPOSITIONS: + errors.append(f"{label}: invalid repository disposition") + is_current = label in current + if (disposition == "ACTIVE") != is_current: + errors.append(f"{label}: disposition and membership disagree") + if (repo.get("program_state") == "historical") != (label in historical): + errors.append(f"{label}: historical state and membership disagree") + eligible = is_current and label not in inventory["workload_ineligible"] + if repo.get("workload_eligible") is not eligible: + errors.append(f"{label}: workload eligibility drift") + destination = repo.get("consolidated_into") + if destination != inventory["consolidation_destinations"].get(label): + errors.append(f"{label}: consolidation destination differs from final receipt") + if disposition == "CONSOLIDATED" and (not isinstance(destination, str) or destination not in current): + errors.append(f"{label}: dangling consolidation destination") + if disposition != "CONSOLIDATED" and destination is not None: + errors.append(f"{label}: unexpected consolidation destination") + if is_current: + if repo.get("last_verified_head_sha") != inventory["heads"].get(label): + errors.append(f"{label}: current HEAD differs from verified inventory") + if retired_reference({k: repo.get(k) for k in ("next_task", "blockers", "dependencies", "dependents")}): + errors.append(f"{label}: retired active-work reference") + if any(n in historical for n in dependencies + dependents if isinstance(n, str)): + errors.append(f"{label}: current relationship targets historical repository") + elif (dependencies or dependents or repo.get("active_experiment_ids") + or repo.get("next_task") != "None; historical source is retired from workload selection."): + errors.append(f"{label}: retired repository has active work") + + if label in historical: + if repo.get("completion_status") != "SUPERSEDED": + errors.append(f"{label}: retired source completion must be SUPERSEDED") + origin = historical_repos.get(label, {}) + if repo.get("lifecycle_stage") != origin.get("lifecycle_stage"): + errors.append(f"{label}: retirement must preserve highest evidenced maturity") + if repo.get("historical_experiment_ids") != origin.get("active_experiment_ids"): + errors.append(f"{label}: historical experiment membership drift") + elif repo.get("historical_experiment_ids") != []: + errors.append(f"{label}: current experiment links belong in active_experiment_ids") + complete = repo.get("completion_status") == "COMPLETE" complete_stage = repo.get("lifecycle_stage") == "COMPLETE" if complete != complete_stage: @@ -816,7 +872,7 @@ def validate( continue if project not in expected: errors.append(f"experiment {label}: references nonexistent project {project}") - elif label not in repo_by_name.get(project, {}).get("active_experiment_ids", []): + elif label not in repo_by_name.get(project, {}).get("historical_experiment_ids" if project in historical else "active_experiment_ids", []): errors.append( f"experiment {label}: participating repository {project} " "does not reciprocally list the experiment" @@ -881,9 +937,7 @@ def validate( ): errors.append("EXP-001 must record the completed independent Gearbox publication") gearbox_repository = repo_by_name.get("gearbox", {}) - if reconciliation.get("gearbox_repository_head_sha") != gearbox_repository.get( - "last_verified_head_sha" - ): + if reconciliation.get("gearbox_repository_head_sha") != registry.get("gearbox_publication", {}).get("final_main_sha"): errors.append("EXP-001 Gearbox publication SHA must match the registry") freeze = exp001.get("offline_benchmark_freeze") if not isinstance(freeze, dict): @@ -1156,16 +1210,19 @@ def validate_theory( if not isinstance(repositories, list): return ["registry.repositories must be an array before theory validation"] - if theory.get("registry_id") != "opsle.theory-registry.v1": - errors.append("theory.registry_id must be opsle.theory-registry.v1") - if theory.get("schema_version") != 1: - errors.append("theory.schema_version must be 1") - if theory.get("source_repository_count") != len(EXPECTED_REPOSITORIES): - errors.append( - f"theory.source_repository_count must be {len(EXPECTED_REPOSITORIES)}" - ) - if theory.get("current_concept_repository_count") != 18: - errors.append("theory.current_concept_repository_count must be 18") + if theory.get("registry_id") != "opsle.theory-registry.v2": + errors.append("theory.registry_id must be opsle.theory-registry.v2") + if theory.get("schema_version") != 2: + errors.append("theory.schema_version must be 2") + expected = membership(registry, "current") | membership(registry, "historical") + current_repositories = membership(registry, "current") + if theory.get("source_repository_count") != len(expected): + errors.append("theory source repository count drift") + if theory.get("concept_count") != len(concepts): + errors.append("theory concept count drift") + actual_counts = Counter(x.get("concept_disposition") for x in concepts if isinstance(x, dict)) + if theory.get("concept_counts") != dict(actual_counts): + errors.append("theory disposition count drift") if not _valid_timestamp(theory.get("verified_at")): errors.append("theory.verified_at must be an ISO-8601 UTC timestamp") @@ -1190,10 +1247,14 @@ def validate_theory( duplicate_ids = sorted(item for item, count in Counter(ids).items() if count > 1) if duplicate_ids: errors.append(f"duplicate theory concept IDs: {', '.join(duplicate_ids)}") - if len(concepts) != 18: - errors.append(f"expected 18 theory concepts, found {len(concepts)}") - - current_concept_repositories = set(EXPECTED_REPOSITORIES[:18]) + # Concepts can share a consolidated home. Identity coverage uses source concepts. + current_concept_repositories = { + r.get("name") for r in repositories + if isinstance(r, dict) and r.get("project_type") == "concept" + and isinstance(r.get("name"), str) + } + if set(ids) != current_concept_repositories: + errors.append("theory concept identity coverage differs from repository concepts") current_mappings: list[str] = [] concept_ids = set(ids) concept_by_id = { @@ -1244,7 +1305,7 @@ def validate_theory( current_mappings.append(current_repository) if ( isinstance(current_repository, str) - and current_repository not in current_concept_repositories + and current_repository not in current_repositories ): errors.append( f"theory concept {label}: invalid current repository {current_repository!r}" @@ -1276,15 +1337,31 @@ def validate_theory( f"theory concept {label}: current_implementation_fidelity requires status and assessment" ) - mapping_counts = Counter(current_mappings) - duplicate_mappings = sorted(name for name, count in mapping_counts.items() if count > 1) - missing_mappings = sorted(current_concept_repositories - set(current_mappings)) - if duplicate_mappings: - errors.append(f"duplicate current concept repository mappings: {', '.join(duplicate_mappings)}") - if missing_mappings: - errors.append(f"missing current concept repository mappings: {', '.join(missing_mappings)}") - if len(current_mappings) != 18: - errors.append(f"expected 18 current concept repository mappings, found {len(current_mappings)}") + if theory.get("current_concept_repository_count") != len(set(current_mappings)): + errors.append("theory current concept repository count drift") + repo_by_name = {r.get("name"): r for r in repositories if isinstance(r, dict) and isinstance(r.get("name"), str)} + for concept in concepts: + if not isinstance(concept, dict) or not isinstance(concept.get("id"), str): + continue + label = concept["id"] + repo = repo_by_name.get(label, {}) + expected_disposition = {"ACTIVE": "ACTIVE", "CONSOLIDATED": "CONSOLIDATED", "RETIRED": "RETIRED"}.get(repo.get("repository_disposition")) + expected_home = repo.get("consolidated_into") if expected_disposition == "CONSOLIDATED" else (label if expected_disposition == "ACTIVE" else None) + if concept.get("concept_disposition") != expected_disposition or concept.get("current_repository") != expected_home: + errors.append(f"{label}: concept disposition or current home drift") + if concept.get("source_repository") != label: + errors.append(f"{label}: concept source provenance drift") + if concept.get("highest_evidenced_stage") != repo.get("lifecycle_stage"): + errors.append(f"{label}: concept research maturity drift") + if expected_disposition != "ACTIVE" and concept.get("recommended_disposition") != expected_disposition: + errors.append(f"{label}: executed disposition must supersede recommendation") + for field in ("dependencies", "consumers"): + relations = concept.get(field, []) + if not isinstance(relations, list): + continue + for target in relations: + if isinstance(target, str) and concept_by_id.get(target, {}).get("concept_disposition") == "RETIRED": + errors.append(f"{label}: current concept relationship targets retired concept") gearbox = concept_by_id.get("gearbox") if gearbox is None: @@ -1336,26 +1413,12 @@ def validate_theory( errors.append("registry Gearbox reconciliation definition drifted") if reconciliation.get("canonical_context_firewall_definition") != CANONICAL_CONTEXT_FIREWALL_DEFINITION: errors.append("registry Context Firewall reconciliation definition drifted") - if reconciliation.get("status") != "IMPLEMENTED_HOME_REGISTERED": - errors.append( - "registry theory reconciliation status must record the implemented home" - ) + if reconciliation.get("status") != "POST_CONSOLIDATION_RECONCILED": + errors.append("registry theory reconciliation must record post-consolidation state") if not _valid_timestamp(reconciliation.get("verified_at")): - errors.append("registry theory reconciliation verified_at must be an ISO-8601 UTC timestamp") - if reconciliation.get("gearbox_vs_durable_supervisor") != GEARBOX_VS_DURABLE_SUPERVISOR: - errors.append("registry Gearbox versus Durable Supervisor distinction drifted") - if reconciliation.get("gearbox_repository_status") != "CREATED_PROTOTYPED": - errors.append("registry must record Gearbox as created and prototyped") - if reconciliation.get("repository_topology_operations_executed") is not True: - errors.append("registry must record the authorized Gearbox repository creation") - if reconciliation.get("lifecycle_changes_executed") is not True: - errors.append("registry must record the evidence-backed Gearbox lifecycle assignment") - if reconciliation.get("existing_repository_lifecycle_changes_executed") is not False: - errors.append("registry must record no existing-repository lifecycle changes") - if reconciliation.get("existing_repository_dispositions_executed") is not False: - errors.append("registry must record no executed existing-repository dispositions") + errors.append("registry theory reconciliation needs a verified timestamp") if reconciliation.get("model_provider_runs_added") != 0: - errors.append("registry must record zero model/provider runs for Gearbox publication") + errors.append("ledger reconciliation must add zero provider runs") publication = registry.get("gearbox_publication") if not isinstance(publication, dict): errors.append("registry.gearbox_publication must be an object") @@ -1380,10 +1443,9 @@ def validate_theory( errors.append("registry Gearbox publication CI must be successful") if not _nonempty_string(publication.get("ci_run")): errors.append("registry Gearbox publication CI run is required") - if publication.get("final_main_sha") != gearbox_repository.get( - "last_verified_head_sha" - ): - errors.append("registry Gearbox publication SHA must match repository HEAD") + historical = load_json(ROOT / "program/history/pre-consolidation/registry.json") + if publication.get("final_main_sha") != historical["gearbox_publication"]["final_main_sha"]: + errors.append("registry Gearbox publication SHA must match historical publication") if gearbox_repository.get("lifecycle_stage") != "PROTOTYPED": errors.append("registered Gearbox lifecycle stage must be PROTOTYPED") for field in ("final_main_sha", "implementation_revision"): @@ -1416,7 +1478,7 @@ def validate_theory( item for item in concepts if isinstance(item, dict) - and item.get("current_repository") == repository + and item.get("id") == repository ), None, ) @@ -1461,7 +1523,7 @@ def main(argv: list[str] | None = None) -> int: print(f"- {error}", file=sys.stderr) return 1 print( - f"PASS: {len(EXPECTED_REPOSITORIES)} repositories, " + f"PASS: {len(registry['repositories'])} repositories, " f"{len(theory['concepts'])} concepts, and " f"{len(experiments['experiments'])} experiments validated" )