Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions PROGRAM_STATUS.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,16 +4,16 @@

**Coverage: 21/21 expected repositories; duplicates: 0.**

Last verified: `2026-08-31T01:29:54Z`. HEADs are the verified default-branch revisions, not an assumption about later changes.
Last verified: `2026-08-31T02:59:01Z`. HEADs are the verified default-branch revisions, not an assumption about later changes.

## Portfolio totals

| Lifecycle stage | Count |
|---|---:|
| `THEORY` | 12 |
| `SPECIFIED` | 0 |
| `PROTOTYPED` | 7 |
| `VERIFIED` | 2 |
| `PROTOTYPED` | 6 |
| `VERIFIED` | 3 |
| `BENCHMARK_READY` | 0 |
| `EXPERIMENTED` | 0 |
| `REPRODUCED` | 0 |
Expand Down Expand Up @@ -43,7 +43,7 @@ Program state totals: active 5; waiting 16; complete 0.
| 15 | [agent-recovery-policy](https://github.com/opsle/agent-recovery-policy) | concept | `1b733a111e26` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md); placeholder only; no automated tests | No shared failure schema, attempt ledger, route evaluator, or comparative fixture set. | After decision evidence and route schemas stabilize, define same-failure convergence on synthetic failures. | `agent-routing-policy`, `agent-state-ledger`, `decision-evidence-protocol` | waiting |
| 16 | [ephemeral-agent-workers](https://github.com/opsle/ephemeral-agent-workers) | concept | `ad96fcfdfac0` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md); placeholder only; no automated tests | Portable authority, claim, and handoff contracts are not ready; no safe synthetic containment harness exists. | Wait for prerequisite contracts, then define a fake worker adapter and destruction receipt without infrastructure changes. | `agent-execution-authorization`, `agent-resource-claims`, `verifiable-agent-handoff` | waiting |
| 17 | [gearbox](https://github.com/opsle/gearbox) | concept | `f3fab9f292cf` | `PROTOTYPED` | [provider-free Python reference core with strict authority-policy admission, exact deterministic argv execution, content-addressed staged helper context, injected one-shot helper transport, passive process waiting, compact results, raw-artifact accounting, fail-closed budgets, and Visible Value receipts](https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py); 19 of 19 provider-free automated tests passed locally, in PR #1 CI, and in final-main CI; ruff, shellcheck, actionlint, gitleaks, wheel build, receipt validation, and public raw-locator/hash checks passed | A production-quality bounded helper transport, independently verified isolation and termination, full Context Firewall integration, and a frozen comparative benchmark remain missing. | Freeze a provider-free deterministic-versus-direct baseline and helper-transport conformance corpus before considering any live model/provider run. | `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `agent-routing-policy`, `agent-execution-authorization` | waiting |
| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `12076522c9b8` | `PROTOTYPED` | [dependency-free Node.js 20 deterministic planner with normalized change/impact evidence, reverse-dependency closure, explicit verification catalog and policy rules, selected and skipped check arguments, fail-closed escalation, canonical SHA-256 provenance, opsle.value-receipt.v1 output, and fixture-level shadow miss classification](https://github.com/opsle/affected-verification/blob/12076522c9b82501794d816f1fcc0b7775fad6e1/SPEC.md); 54 of 54 automated tests, 15 of 15 positive/boundary/negative/tampered conformance scenarios, and 5 of 5 determinism checks passed locally and in PR #1 CI; exact main CI, current Opsle value-receipt validation, actionlint, gitleaks, and clean local/remote parity passed | A frozen real public-repository full-verification baseline, complete catalog, production evidence adapter, relevance oracle, native-selector comparison, shadow observations, and independent review are missing. | Freeze one real public repository's full-verification baseline, catalog, adapter outputs, and relevance oracle, then run an OBSERVE/SHADOW comparison without replacing CI. | — | active |
| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `641aee9d29a8` | `VERIFIED` | [dependency-free Node.js deterministic planner plus an AV-EXP-001 benchmark-only Git/catalog/source-graph/Vitest adapter, identity-bound SHADOW result validator, complete frozen-oracle harness, explainable skip records, fail-closed uncertainty handling, and opsle.value-receipt.v1 telemetry](https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/REPORT.md); 70 of 70 automated tests, 15 of 15 conformance scenarios, and 6 of 6 determinism checks passed locally, in PR #2 CI, and from a fresh detached worktree at exact main; invalid-state coverage includes target, patch, catalog, selector, adapter, baseline, oracle, scenario, skip-reason, policy, trust-state, and result tampering | A second ecosystem, historical real-change replay, production-quality evidence adapter, and independent qualifying replication remain missing; AV therefore remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized. | Preregister and run a second public-repository shadow calibration in a different ecosystem with a meaningful native selector and the same full-catalog oracle discipline. | — | active |
| 19 | [research](https://github.com/opsle/research) | program infrastructure | `9ee43197880c` | `PROTOTYPED` | [authoritative 21-repository ledger, machine-readable 18-concept theory registry including Affected Verification, canonical theory map, normative Visible Value controls, and provider-free EXP-001 benchmark, launch, one-block coordinator, external four-label LIVE_PROVIDER_RUN authorization, and current catalogue/pricing preflight artifacts with six content-addressed tasks, deterministic oracle, four arm contracts, sealed blinded allocation, exact subject configuration and adapter, exact authorization admission, private boundaries, receipts, mutation tests, and integrity CI](program/THEORY_MAP.md); 88 of 88 repository tests pass locally after deliberate migration to the 21-repository, 18-concept anti-forgetting set, including 13 authorization validations and two byte-identical replays; generated status and registry validation pass | The exact live authorization set remains unconsumed and unreleased, account-specific API entitlement is unverified under the zero-provider-call policy, no immutable dated model snapshot is documented, and the program has no canonical measured concept experiment. | Independently review and release the provider-free live-authorization and catalogue/pricing preflight; do not consume authorization or launch a provider/model subject. | — | active |
| 20 | [site](https://github.com/opsle/site) | program infrastructure | `28ad65be4750` | `PROTOTYPED` | [React/Vinext source implementation with content routes](https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md); automated build/render tests present; not rerun because this reconciliation kept other repositories read-only | Wait for validated registry data and measured research; deployment requires separate authorization. | After registry merge, add a read-only registry ingestion design without deploying the site. | `research` | waiting |
| 21 | [.github](https://github.com/opsle/.github) | program infrastructure | `01c38e726db7` | `THEORY` | [documentation-only organization profile](https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md); not applicable to current single Markdown profile; consistency is unverified | No mechanical registry consistency check exists in this repository. | After registry merge, design a read-only consistency check for organization-profile repository links. | `research` | waiting |
Expand Down
14 changes: 8 additions & 6 deletions program/PRIORITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,12 +82,14 @@ conversation-local.

## Workstream 6: affected verification

Affected Verification is independently prototyped for normalized synthetic
evidence and remains outside Gearbox. Its next evidence gate is not another
selector implementation: freeze one real public repository's full-verification
baseline, complete catalog, native-selector adapter output, and relevance oracle,
then run only in `OBSERVE`/`SHADOW` without replacing authoritative CI. Correctness
and relevant selection misses precede workload reduction.
Affected Verification is independently verified for its narrow planner and the
AV-EXP-001 Zustand/Vitest shadow calibration, and remains outside Gearbox. Both
AV arms selected 8/8 oracle-relevant checks in the frozen ten-scenario corpus;
the native tests-only arm selected 6/8 and omitted relevant lint and typecheck.
This is one-ecosystem calibration evidence, not general safety or bounded trust.
Its exact next execution is a preregistered second public-repository shadow
calibration in a different ecosystem with a meaningful native selector and the
same full-catalog oracle discipline.

## Exact next execution

Expand Down
14 changes: 8 additions & 6 deletions program/THEORY_MAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ Consolidation and disposition operations remain recommendations only.
Machine source: [`theory-registry.json`](theory-registry.json).

Theory registry canonical SHA-256:
`d3a5f6e03f28d0ae8911763e6cb15c1a3f7306559f1801a2647219563159cb79`.
`7cdc866ada723f24fd16af005bb54b9a60873e0623effb70df691a7a58fcdc47`.

This map corrects an extraction-boundary error. The original 16 concept
repositories were useful hypotheses isolated from one production system, but
Expand Down Expand Up @@ -408,11 +408,13 @@ repositories.
- Verifiable Agent Handoff is prototyped only for HMAC manifest binding and
caller-supplied destruction state; publication, transport, destruction proof,
reconstruction, and independent verification are not implemented.
- Affected Verification is prototyped for normalized synthetic evidence,
deterministic catalog/policy planning, fail-closed escalation, skip reasons,
Visible Value receipts, and fixture-level shadow classification. No production
adapter, check execution, real-repository benchmark, or trusted selective
verification class exists.
- Affected Verification is verified for its narrow deterministic planner and
the revision-bound AV-EXP-001 Zustand/Vitest shadow calibration. Its frozen
corpus observed 8/8 relevant checks selected by both AV arms, including
conservative full broadening under incomplete impact evidence. The result is
still one repository, one ecosystem, and mostly synthetic change/fault
shapes; no production adapter, independent qualifying replication, or trusted
selective-verification class exists.
- The remaining theory repositories contain coherent falsifiable narratives,
but their `SPEC.md` files are generic templates and their source/tests are
placeholders. Their current lifecycle remains `THEORY`.
Expand Down
96 changes: 95 additions & 1 deletion program/experiments.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"schema_version": 1,
"last_verified_at": "2026-08-30T18:49:10Z",
"last_verified_at": "2026-08-31T02:59:01Z",
"experiments": [
{
"id": "EXP-001",
Expand Down Expand Up @@ -249,6 +249,100 @@
},
"next_task": "Independently review and release the provider-free live-authorization preflight branch; do not consume authorization or launch a provider/model subject."
},
{
"id": "AV-EXP-001",
"title": "Minimum Defensible Verification — Real Repository Shadow Calibration",
"status": "RECORDED",
"hypothesis": "On one pinned real public repository, Affected Verification can propose less than the full verification workload while selecting every check that the frozen full-catalog oracle demonstrates was relevant.",
"participating_repositories": [
"affected-verification",
"research"
],
"roles": {
"primary": "affected-verification",
"expected_support": [
"research"
],
"potential_support": []
},
"baseline": "The complete frozen 17-check Zustand verification catalog at b57db4f86ef179285da216eeb291266da82c361c, executed for every scenario and authoritative over all selector predictions.",
"experimental_arms": [
"FULL frozen verification catalog",
"NATIVE Vitest 4.1.10 related selector",
"AV_CORE with normalized Git, catalog, source-graph, and policy evidence",
"AV_WITH_NATIVE_EVIDENCE with native output as one normalized evidence source"
],
"primary_metric": "Selection misses reported individually against checks whose full-catalog outcome changed because of a frozen scenario.",
"secondary_metrics": [
"relevant-check recall and scenario-level misses",
"exact selected and skipped test files, test executions, and non-test checks as separate units",
"uncertainty broadening and full-verification escalation",
"observed wall-clock telemetry without a causal time-saved claim",
"native-versus-AV selection differences"
],
"correctness_gate": "Every frozen catalog check runs for every scenario; selector predictions never accept a scenario, and a miss is any omitted oracle-relevant failing check.",
"failure_classifications": [
"selection miss",
"selected relevant failure",
"irrelevant full-catalog failure",
"incomplete or indeterminate full run",
"baseline instability",
"harness defect requiring a versioned amendment",
"evidence or identity drift",
"conservative broadening",
"insufficient evidence"
],
"dataset_fixture_identity": "Zustand b57db4f86ef179285da216eeb291266da82c361c; preregistration commit 0544362d7659093b7f0b4f89ee8f68023fd269c3; catalog sha256:8c5b224deaa7077690341248a18a2155310e2b16072e1607a9b4cd546e3a0914; ten scenario identities and patches frozen in the preregistration.",
"model_provider_configuration": "NONE: no model/provider benchmark subject or external provider workload was used.",
"run_identities": [
"sha256:68b8582a9ce7b86bfa5431d89d2dea07f8c34b88d1d0350bab25c99fa5b236df"
],
"result_artifacts": [
"https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/REPORT.md",
"https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/summary.json",
"https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/analysis.json",
"https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/evidence-manifest.json"
],
"target": {
"repository": "https://github.com/pmndrs/zustand.git",
"sha": "b57db4f86ef179285da216eeb291266da82c361c",
"license": "MIT"
},
"preregistration": {
"commit_sha": "0544362d7659093b7f0b4f89ee8f68023fd269c3",
"amendment_count": 3,
"comparative_outcomes_observed_before_commit": false
},
"benchmark_result": {
"affected_verification_main_sha": "641aee9d29a89e2a8819f00817ccee8e5d234dcb",
"results_commit_sha": "97f0301c28e9840e18c5aa35a6cbf95b95f7c6cf",
"summary_identity": "sha256:68b8582a9ce7b86bfa5431d89d2dea07f8c34b88d1d0350bab25c99fa5b236df",
"analysis_identity": "sha256:c431d8849edce79d6121a290f49288ed2710e406600d8b66d3588e6b82c73a1d",
"evidence_bundle_identity": "sha256:1e176b7a40b5f16451797d87784f560f932b686f2fe261731526709331ff1172",
"native_selector_identity": "sha256:f9f8f41244923a6daa6f86b0818889c955b6b26362432b76ae4ee361f10171d6",
"scenario_count": 10,
"synthetic_fault_count": 6,
"synthetic_benign_change_shape_count": 4,
"relevant_check_count": 8,
"native_selected_relevant_check_count": 6,
"native_missed_relevant_check_count": 2,
"av_core_selected_relevant_check_count": 8,
"av_core_missed_relevant_check_count": 0,
"av_with_native_selected_relevant_check_count": 8,
"av_with_native_missed_relevant_check_count": 0,
"av_full_escalation_count": 3
},
"replication_status": "SAME_HOST_DETERMINISTIC_REPLAY_ONLY",
"verdict": "PASS for completing the preregistered shadow calibration: no AV selection miss was observed in the frozen corpus. This does not establish general safety, correctness equivalence, causal savings, or bounded trust.",
"blockers": [
"Only one repository and one TypeScript/Vitest ecosystem were calibrated.",
"The corpus uses six synthetic faults and four synthetic benign change shapes rather than a historical real-change replay.",
"The adapter is benchmark-only and no independent qualifying replication exists.",
"Affected Verification remains OBSERVE/SHADOW; no TRUSTED_BOUNDED class is authorized."
],
"lifecycle_impact": "PROMOTE_TO_VERIFIED_ONLY: the narrow scoped correctness and safety claims pass meaningful automated and revision-bound benchmark checks, but this run is explicitly capped below BENCHMARK_READY and does not authorize trusted execution.",
"next_task": "Preregister and run a second public-repository shadow calibration in a different ecosystem with a meaningful native selector and the same full-catalog oracle discipline."
},
{
"id": "LEGACY-001",
"title": "Graphify plus Antigravity semantic adapter integration observation",
Expand Down
Loading