diff --git a/PROGRAM_STATUS.md b/PROGRAM_STATUS.md index 90ca65c..4efbc99 100644 --- a/PROGRAM_STATUS.md +++ b/PROGRAM_STATUS.md @@ -4,7 +4,7 @@ **Coverage: 21/21 expected repositories; duplicates: 0.** -Last verified: `2026-08-31T01:29:54Z`. HEADs are the verified default-branch revisions, not an assumption about later changes. +Last verified: `2026-08-31T02:59:01Z`. HEADs are the verified default-branch revisions, not an assumption about later changes. ## Portfolio totals @@ -12,8 +12,8 @@ Last verified: `2026-08-31T01:29:54Z`. HEADs are the verified default-branch rev |---|---:| | `THEORY` | 12 | | `SPECIFIED` | 0 | -| `PROTOTYPED` | 7 | -| `VERIFIED` | 2 | +| `PROTOTYPED` | 6 | +| `VERIFIED` | 3 | | `BENCHMARK_READY` | 0 | | `EXPERIMENTED` | 0 | | `REPRODUCED` | 0 | @@ -43,7 +43,7 @@ Program state totals: active 5; waiting 16; complete 0. | 15 | [agent-recovery-policy](https://github.com/opsle/agent-recovery-policy) | concept | `1b733a111e26` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md); placeholder only; no automated tests | No shared failure schema, attempt ledger, route evaluator, or comparative fixture set. | After decision evidence and route schemas stabilize, define same-failure convergence on synthetic failures. | `agent-routing-policy`, `agent-state-ledger`, `decision-evidence-protocol` | waiting | | 16 | [ephemeral-agent-workers](https://github.com/opsle/ephemeral-agent-workers) | concept | `ad96fcfdfac0` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md); placeholder only; no automated tests | Portable authority, claim, and handoff contracts are not ready; no safe synthetic containment harness exists. | Wait for prerequisite contracts, then define a fake worker adapter and destruction receipt without infrastructure changes. | `agent-execution-authorization`, `agent-resource-claims`, `verifiable-agent-handoff` | waiting | | 17 | [gearbox](https://github.com/opsle/gearbox) | concept | `f3fab9f292cf` | `PROTOTYPED` | [provider-free Python reference core with strict authority-policy admission, exact deterministic argv execution, content-addressed staged helper context, injected one-shot helper transport, passive process waiting, compact results, raw-artifact accounting, fail-closed budgets, and Visible Value receipts](https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py); 19 of 19 provider-free automated tests passed locally, in PR #1 CI, and in final-main CI; ruff, shellcheck, actionlint, gitleaks, wheel build, receipt validation, and public raw-locator/hash checks passed | A production-quality bounded helper transport, independently verified isolation and termination, full Context Firewall integration, and a frozen comparative benchmark remain missing. | Freeze a provider-free deterministic-versus-direct baseline and helper-transport conformance corpus before considering any live model/provider run. | `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `agent-routing-policy`, `agent-execution-authorization` | waiting | -| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `12076522c9b8` | `PROTOTYPED` | [dependency-free Node.js 20 deterministic planner with normalized change/impact evidence, reverse-dependency closure, explicit verification catalog and policy rules, selected and skipped check arguments, fail-closed escalation, canonical SHA-256 provenance, opsle.value-receipt.v1 output, and fixture-level shadow miss classification](https://github.com/opsle/affected-verification/blob/12076522c9b82501794d816f1fcc0b7775fad6e1/SPEC.md); 54 of 54 automated tests, 15 of 15 positive/boundary/negative/tampered conformance scenarios, and 5 of 5 determinism checks passed locally and in PR #1 CI; exact main CI, current Opsle value-receipt validation, actionlint, gitleaks, and clean local/remote parity passed | A frozen real public-repository full-verification baseline, complete catalog, production evidence adapter, relevance oracle, native-selector comparison, shadow observations, and independent review are missing. | Freeze one real public repository's full-verification baseline, catalog, adapter outputs, and relevance oracle, then run an OBSERVE/SHADOW comparison without replacing CI. | — | active | +| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `641aee9d29a8` | `VERIFIED` | [dependency-free Node.js deterministic planner plus an AV-EXP-001 benchmark-only Git/catalog/source-graph/Vitest adapter, identity-bound SHADOW result validator, complete frozen-oracle harness, explainable skip records, fail-closed uncertainty handling, and opsle.value-receipt.v1 telemetry](https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/REPORT.md); 70 of 70 automated tests, 15 of 15 conformance scenarios, and 6 of 6 determinism checks passed locally, in PR #2 CI, and from a fresh detached worktree at exact main; invalid-state coverage includes target, patch, catalog, selector, adapter, baseline, oracle, scenario, skip-reason, policy, trust-state, and result tampering | A second ecosystem, historical real-change replay, production-quality evidence adapter, and independent qualifying replication remain missing; AV therefore remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized. | Preregister and run a second public-repository shadow calibration in a different ecosystem with a meaningful native selector and the same full-catalog oracle discipline. | — | active | | 19 | [research](https://github.com/opsle/research) | program infrastructure | `9ee43197880c` | `PROTOTYPED` | [authoritative 21-repository ledger, machine-readable 18-concept theory registry including Affected Verification, canonical theory map, normative Visible Value controls, and provider-free EXP-001 benchmark, launch, one-block coordinator, external four-label LIVE_PROVIDER_RUN authorization, and current catalogue/pricing preflight artifacts with six content-addressed tasks, deterministic oracle, four arm contracts, sealed blinded allocation, exact subject configuration and adapter, exact authorization admission, private boundaries, receipts, mutation tests, and integrity CI](program/THEORY_MAP.md); 88 of 88 repository tests pass locally after deliberate migration to the 21-repository, 18-concept anti-forgetting set, including 13 authorization validations and two byte-identical replays; generated status and registry validation pass | The exact live authorization set remains unconsumed and unreleased, account-specific API entitlement is unverified under the zero-provider-call policy, no immutable dated model snapshot is documented, and the program has no canonical measured concept experiment. | Independently review and release the provider-free live-authorization and catalogue/pricing preflight; do not consume authorization or launch a provider/model subject. | — | active | | 20 | [site](https://github.com/opsle/site) | program infrastructure | `28ad65be4750` | `PROTOTYPED` | [React/Vinext source implementation with content routes](https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md); automated build/render tests present; not rerun because this reconciliation kept other repositories read-only | Wait for validated registry data and measured research; deployment requires separate authorization. | After registry merge, add a read-only registry ingestion design without deploying the site. | `research` | waiting | | 21 | [.github](https://github.com/opsle/.github) | program infrastructure | `01c38e726db7` | `THEORY` | [documentation-only organization profile](https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md); not applicable to current single Markdown profile; consistency is unverified | No mechanical registry consistency check exists in this repository. | After registry merge, design a read-only consistency check for organization-profile repository links. | `research` | waiting | diff --git a/program/PRIORITY.md b/program/PRIORITY.md index 30eab72..3c62628 100644 --- a/program/PRIORITY.md +++ b/program/PRIORITY.md @@ -82,12 +82,14 @@ conversation-local. ## Workstream 6: affected verification -Affected Verification is independently prototyped for normalized synthetic -evidence and remains outside Gearbox. Its next evidence gate is not another -selector implementation: freeze one real public repository's full-verification -baseline, complete catalog, native-selector adapter output, and relevance oracle, -then run only in `OBSERVE`/`SHADOW` without replacing authoritative CI. Correctness -and relevant selection misses precede workload reduction. +Affected Verification is independently verified for its narrow planner and the +AV-EXP-001 Zustand/Vitest shadow calibration, and remains outside Gearbox. Both +AV arms selected 8/8 oracle-relevant checks in the frozen ten-scenario corpus; +the native tests-only arm selected 6/8 and omitted relevant lint and typecheck. +This is one-ecosystem calibration evidence, not general safety or bounded trust. +Its exact next execution is a preregistered second public-repository shadow +calibration in a different ecosystem with a meaningful native selector and the +same full-catalog oracle discipline. ## Exact next execution diff --git a/program/THEORY_MAP.md b/program/THEORY_MAP.md index 4711ca5..5473e9c 100644 --- a/program/THEORY_MAP.md +++ b/program/THEORY_MAP.md @@ -7,7 +7,7 @@ Consolidation and disposition operations remain recommendations only. Machine source: [`theory-registry.json`](theory-registry.json). Theory registry canonical SHA-256: -`d3a5f6e03f28d0ae8911763e6cb15c1a3f7306559f1801a2647219563159cb79`. +`7cdc866ada723f24fd16af005bb54b9a60873e0623effb70df691a7a58fcdc47`. This map corrects an extraction-boundary error. The original 16 concept repositories were useful hypotheses isolated from one production system, but @@ -408,11 +408,13 @@ repositories. - Verifiable Agent Handoff is prototyped only for HMAC manifest binding and caller-supplied destruction state; publication, transport, destruction proof, reconstruction, and independent verification are not implemented. -- Affected Verification is prototyped for normalized synthetic evidence, - deterministic catalog/policy planning, fail-closed escalation, skip reasons, - Visible Value receipts, and fixture-level shadow classification. No production - adapter, check execution, real-repository benchmark, or trusted selective - verification class exists. +- Affected Verification is verified for its narrow deterministic planner and + the revision-bound AV-EXP-001 Zustand/Vitest shadow calibration. Its frozen + corpus observed 8/8 relevant checks selected by both AV arms, including + conservative full broadening under incomplete impact evidence. The result is + still one repository, one ecosystem, and mostly synthetic change/fault + shapes; no production adapter, independent qualifying replication, or trusted + selective-verification class exists. - The remaining theory repositories contain coherent falsifiable narratives, but their `SPEC.md` files are generic templates and their source/tests are placeholders. Their current lifecycle remains `THEORY`. diff --git a/program/experiments.json b/program/experiments.json index a86ac55..292884c 100644 --- a/program/experiments.json +++ b/program/experiments.json @@ -1,6 +1,6 @@ { "schema_version": 1, - "last_verified_at": "2026-08-30T18:49:10Z", + "last_verified_at": "2026-08-31T02:59:01Z", "experiments": [ { "id": "EXP-001", @@ -249,6 +249,100 @@ }, "next_task": "Independently review and release the provider-free live-authorization preflight branch; do not consume authorization or launch a provider/model subject." }, + { + "id": "AV-EXP-001", + "title": "Minimum Defensible Verification — Real Repository Shadow Calibration", + "status": "RECORDED", + "hypothesis": "On one pinned real public repository, Affected Verification can propose less than the full verification workload while selecting every check that the frozen full-catalog oracle demonstrates was relevant.", + "participating_repositories": [ + "affected-verification", + "research" + ], + "roles": { + "primary": "affected-verification", + "expected_support": [ + "research" + ], + "potential_support": [] + }, + "baseline": "The complete frozen 17-check Zustand verification catalog at b57db4f86ef179285da216eeb291266da82c361c, executed for every scenario and authoritative over all selector predictions.", + "experimental_arms": [ + "FULL frozen verification catalog", + "NATIVE Vitest 4.1.10 related selector", + "AV_CORE with normalized Git, catalog, source-graph, and policy evidence", + "AV_WITH_NATIVE_EVIDENCE with native output as one normalized evidence source" + ], + "primary_metric": "Selection misses reported individually against checks whose full-catalog outcome changed because of a frozen scenario.", + "secondary_metrics": [ + "relevant-check recall and scenario-level misses", + "exact selected and skipped test files, test executions, and non-test checks as separate units", + "uncertainty broadening and full-verification escalation", + "observed wall-clock telemetry without a causal time-saved claim", + "native-versus-AV selection differences" + ], + "correctness_gate": "Every frozen catalog check runs for every scenario; selector predictions never accept a scenario, and a miss is any omitted oracle-relevant failing check.", + "failure_classifications": [ + "selection miss", + "selected relevant failure", + "irrelevant full-catalog failure", + "incomplete or indeterminate full run", + "baseline instability", + "harness defect requiring a versioned amendment", + "evidence or identity drift", + "conservative broadening", + "insufficient evidence" + ], + "dataset_fixture_identity": "Zustand b57db4f86ef179285da216eeb291266da82c361c; preregistration commit 0544362d7659093b7f0b4f89ee8f68023fd269c3; catalog sha256:8c5b224deaa7077690341248a18a2155310e2b16072e1607a9b4cd546e3a0914; ten scenario identities and patches frozen in the preregistration.", + "model_provider_configuration": "NONE: no model/provider benchmark subject or external provider workload was used.", + "run_identities": [ + "sha256:68b8582a9ce7b86bfa5431d89d2dea07f8c34b88d1d0350bab25c99fa5b236df" + ], + "result_artifacts": [ + "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/REPORT.md", + "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/summary.json", + "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/analysis.json", + "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/evidence-manifest.json" + ], + "target": { + "repository": "https://github.com/pmndrs/zustand.git", + "sha": "b57db4f86ef179285da216eeb291266da82c361c", + "license": "MIT" + }, + "preregistration": { + "commit_sha": "0544362d7659093b7f0b4f89ee8f68023fd269c3", + "amendment_count": 3, + "comparative_outcomes_observed_before_commit": false + }, + "benchmark_result": { + "affected_verification_main_sha": "641aee9d29a89e2a8819f00817ccee8e5d234dcb", + "results_commit_sha": "97f0301c28e9840e18c5aa35a6cbf95b95f7c6cf", + "summary_identity": "sha256:68b8582a9ce7b86bfa5431d89d2dea07f8c34b88d1d0350bab25c99fa5b236df", + "analysis_identity": "sha256:c431d8849edce79d6121a290f49288ed2710e406600d8b66d3588e6b82c73a1d", + "evidence_bundle_identity": "sha256:1e176b7a40b5f16451797d87784f560f932b686f2fe261731526709331ff1172", + "native_selector_identity": "sha256:f9f8f41244923a6daa6f86b0818889c955b6b26362432b76ae4ee361f10171d6", + "scenario_count": 10, + "synthetic_fault_count": 6, + "synthetic_benign_change_shape_count": 4, + "relevant_check_count": 8, + "native_selected_relevant_check_count": 6, + "native_missed_relevant_check_count": 2, + "av_core_selected_relevant_check_count": 8, + "av_core_missed_relevant_check_count": 0, + "av_with_native_selected_relevant_check_count": 8, + "av_with_native_missed_relevant_check_count": 0, + "av_full_escalation_count": 3 + }, + "replication_status": "SAME_HOST_DETERMINISTIC_REPLAY_ONLY", + "verdict": "PASS for completing the preregistered shadow calibration: no AV selection miss was observed in the frozen corpus. This does not establish general safety, correctness equivalence, causal savings, or bounded trust.", + "blockers": [ + "Only one repository and one TypeScript/Vitest ecosystem were calibrated.", + "The corpus uses six synthetic faults and four synthetic benign change shapes rather than a historical real-change replay.", + "The adapter is benchmark-only and no independent qualifying replication exists.", + "Affected Verification remains OBSERVE/SHADOW; no TRUSTED_BOUNDED class is authorized." + ], + "lifecycle_impact": "PROMOTE_TO_VERIFIED_ONLY: the narrow scoped correctness and safety claims pass meaningful automated and revision-bound benchmark checks, but this run is explicitly capped below BENCHMARK_READY and does not authorize trusted execution.", + "next_task": "Preregister and run a second public-repository shadow calibration in a different ecosystem with a meaningful native selector and the same full-catalog oracle discipline." + }, { "id": "LEGACY-001", "title": "Graphify plus Antigravity semantic adapter integration observation", diff --git a/program/registry.json b/program/registry.json index 966575a..e2a7913 100644 --- a/program/registry.json +++ b/program/registry.json @@ -38,7 +38,7 @@ }, "current_highest_priority_workstream": "Independently review and release the provider-free EXP-001 live-authorization and current catalogue/pricing preflight; do not consume authorization or launch any provider/model subject.", "recommended_next_execution": "In opsle/research, create and provider-free validate one exact four-label LIVE_PROVIDER_RUN authorization set plus a model catalogue/pricing preflight artifact; do not consume authorization or launch a provider/model subject.", - "last_verified_at": "2026-08-31T01:29:54Z", + "last_verified_at": "2026-08-31T02:59:01Z", "repositories": [ { "name": "agent-trajectory-profiler", @@ -554,31 +554,31 @@ "name": "affected-verification", "github_url": "https://github.com/opsle/affected-verification", "default_branch": "main", - "last_verified_head_sha": "12076522c9b82501794d816f1fcc0b7775fad6e1", + "last_verified_head_sha": "641aee9d29a89e2a8819f00817ccee8e5d234dcb", "project_type": "concept", "purpose": "Select the smallest verification workload whose sufficiency can be defended from available change-impact, dependency, coverage, policy, and risk evidence.", - "lifecycle_stage": "PROTOTYPED", - "implementation_status": "dependency-free Node.js 20 deterministic planner with normalized change/impact evidence, reverse-dependency closure, explicit verification catalog and policy rules, selected and skipped check arguments, fail-closed escalation, canonical SHA-256 provenance, opsle.value-receipt.v1 output, and fixture-level shadow miss classification", + "lifecycle_stage": "VERIFIED", + "implementation_status": "dependency-free Node.js deterministic planner plus an AV-EXP-001 benchmark-only Git/catalog/source-graph/Vitest adapter, identity-bound SHADOW result validator, complete frozen-oracle harness, explainable skip records, fail-closed uncertainty handling, and opsle.value-receipt.v1 telemetry", "implementation_requirement": "A runnable deterministic planner, input validator, plan contract, conformance fixtures, and shadow classifier are sufficient for the prototype gate; real adapters and comparative evidence are later gates.", "specification_status": "versioned input and opsle.affected-verification.plan.v1 contracts define catalog entries, evidence providers, policy matching, selection and skip reasons, provenance, explicit sufficiency/uncertainty/escalation states, failure behavior, Visible Value limits, and shadow observations", - "test_status": "54 of 54 automated tests, 15 of 15 positive/boundary/negative/tampered conformance scenarios, and 5 of 5 determinism checks passed locally and in PR #1 CI; exact main CI, current Opsle value-receipt validation, actionlint, gitleaks, and clean local/remote parity passed", - "benchmark_status": "twelve public-safe synthetic positive/boundary scenarios plus three explicit negative fixtures and fixture-level shadow classification; benchmark plan only, with no real-repository baseline, native-selector comparison, measured execution, or frozen correctness oracle", - "measured_experiment_status": "none; no real-project verification experiment and zero provider/model runs", - "reproducibility_status": "planner tests, conformance, determinism, canonical plan/value output, fail-closed cases, and shadow classification reproduce locally and in public CI at the verified HEAD; cross-project adapter behavior and comparative safety remain untested", + "test_status": "70 of 70 automated tests, 15 of 15 conformance scenarios, and 6 of 6 determinism checks passed locally, in PR #2 CI, and from a fresh detached worktree at exact main; invalid-state coverage includes target, patch, catalog, selector, adapter, baseline, oracle, scenario, skip-reason, policy, trust-state, and result tampering", + "benchmark_status": "AV-EXP-001 preregistered a pinned Zustand b57db4f86ef179285da216eeb291266da82c361c SHADOW calibration with a 17-check full catalog, three stable clean baselines, four arms, ten frozen scenarios, a deterministic relevance oracle, explicit miss taxonomy, raw evidence, and content-addressed results; AV_CORE and AV_WITH_NATIVE_EVIDENCE selected 8/8 relevant checks with zero observed misses, while NATIVE selected 6/8 and missed relevant lint and typecheck checks", + "measured_experiment_status": "AV-EXP-001 RECORDED; exact plan-count and observed-runtime evidence exists for one repository and ecosystem, six synthetic fault scenarios and four benign synthetic change shapes; zero provider/model runs", + "reproducibility_status": "the public one-command harness and result verifier reproduce the frozen target preparation, catalog, scenarios, native and AV arms, full oracle, receipts, and summary; a second same-host clean run matched the baseline, all ten scenario, summary, and analysis semantic identities, but is not an independent qualifying replication", "documentation_status": "public canonical definition, normative specification, source-linked prior-art audit, architecture and independence boundaries, verification catalog and policy semantics, trust ramp, benchmark plan, limitations, usage, security, schema, and fixtures are present", "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["The prototype trusts normalized component, catalog, and policy claims; has no production Git, graph, coverage, ownership, schema, or native-selector adapters; executes no checks; uses a synthetic corpus; and has no evidence of real-repository sufficiency, reduced undetected-regression risk, comparative computation benefit, or independent reproduction."], + "known_limitations": ["AV-EXP-001 is one repository, one TypeScript/Vitest ecosystem, six synthetic faults and four benign synthetic change shapes; its adapters are benchmark-only, native evidence did not change AV's test count, the replay used the same host and implementation, and no result establishes general safety, correctness equivalence, causal time or cost savings, production trust, or a production-quality adapter."], "dependencies": [], "dependents": [], - "active_experiment_ids": [], - "blockers": ["A frozen real public-repository full-verification baseline, complete catalog, production evidence adapter, relevance oracle, native-selector comparison, shadow observations, and independent review are missing."], - "next_task": "Freeze one real public repository's full-verification baseline, catalog, adapter outputs, and relevance oracle, then run an OBSERVE/SHADOW comparison without replacing CI.", - "evidence": ["https://github.com/opsle/affected-verification/blob/12076522c9b82501794d816f1fcc0b7775fad6e1/SPEC.md", "https://github.com/opsle/affected-verification/blob/12076522c9b82501794d816f1fcc0b7775fad6e1/PRIOR_ART.md", "https://github.com/opsle/affected-verification/blob/12076522c9b82501794d816f1fcc0b7775fad6e1/src/planner.js", "https://github.com/opsle/affected-verification/blob/12076522c9b82501794d816f1fcc0b7775fad6e1/src/shadow.js", "https://github.com/opsle/affected-verification/blob/12076522c9b82501794d816f1fcc0b7775fad6e1/tests/planner.test.js", "https://github.com/opsle/affected-verification/blob/12076522c9b82501794d816f1fcc0b7775fad6e1/tests/value-shadow.test.js", "https://github.com/opsle/affected-verification/pull/1", "https://github.com/opsle/affected-verification/actions/runs/33347673715"], + "active_experiment_ids": ["AV-EXP-001"], + "blockers": ["A second ecosystem, historical real-change replay, production-quality evidence adapter, and independent qualifying replication remain missing; AV therefore remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized."], + "next_task": "Preregister and run a second public-repository shadow calibration in a different ecosystem with a meaningful native selector and the same full-catalog oracle discipline.", + "evidence": ["https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/REPORT.md", "https://github.com/opsle/affected-verification/blob/0544362d7659093b7f0b4f89ee8f68023fd269c3/benchmark/av-exp-001/preregistration-v1/preregistration.json", "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/summary.json", "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/analysis.json", "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/evidence-manifest.json", "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/verify-results.mjs", "https://github.com/opsle/affected-verification/pull/2", "https://github.com/opsle/affected-verification/actions/runs/33352268190"], "completion_criteria": ["Publish production evidence adapters and independently validate their completeness boundaries.", "Run correctness-first comparisons against full verification and native selectors with durable shadow miss evidence.", "Replicate bounded workload and safety claims on independent repositories or environments."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "active", - "last_verified_at": "2026-08-31T01:29:54Z" + "last_verified_at": "2026-08-31T02:59:01Z" }, { "name": "research", diff --git a/program/theory-registry.json b/program/theory-registry.json index 8749d7b..d6e968b 100644 --- a/program/theory-registry.json +++ b/program/theory-registry.json @@ -1,7 +1,7 @@ { "schema_version": 1, "registry_id": "opsle.theory-registry.v1", - "verified_at": "2026-08-31T01:29:54Z", + "verified_at": "2026-08-31T02:59:01Z", "source_repository_count": 21, "current_concept_repository_count": 18, "canonical_definitions": { @@ -714,7 +714,7 @@ "disposition_rationale": "Verification sufficiency, catalogs, risk policy, fail-closed escalation, and skip arguments are independently reusable by humans, agents, CI, Gearbox, and other developer tooling without belonging to an execution engine.", "disposition_gain": "A standalone evidence-composition boundary can reuse native affected selectors while remaining verification-class-neutral and independently benchmarkable.", "disposition_risk": "The project could duplicate mature test-impact systems, overclaim global minimality or safety, or drift into a CI scheduler unless adapters, claim ceilings, and execution boundaries remain explicit.", - "provenance_concerns": "The initial public prototype was created directly in opsle/affected-verification; preserve PR #1, exact revision 12076522c9b82501794d816f1fcc0b7775fad6e1, source-linked prior-art analysis, and synthetic-only evidence without retroactively claiming novelty or real-project safety.", + "provenance_concerns": "The initial public prototype was created directly in opsle/affected-verification; preserve PR #1 at revision 12076522c9b82501794d816f1fcc0b7775fad6e1, the AV-EXP-001 preregistration at 0544362d7659093b7f0b4f89ee8f68023fd269c3, PR #2 and final revision 641aee9d29a89e2a8819f00817ccee8e5d234dcb, source-linked prior art, and the one-repository synthetic-corpus claim boundary without retroactively claiming novelty or general safety.", "relationship_to_gearbox": "Independent provider: Gearbox may request a minimum defensible verification plan and choose execution gears, but it does not own the verification theory, catalog, policy, or implementation.", "relationship_to_context_firewall": "Complementary and sequential: Affected Verification decides what checks should execute; Context Firewall decides what check results should enter model context after execution.", "dependencies": [], @@ -724,22 +724,23 @@ "agent-trajectory-profiler" ], "current_implementation_fidelity": { - "status": "NARROW_PROTOTYPE", - "assessment": "The public dependency-free Node.js core validates normalized evidence/catalog/policy input, computes reverse impact, selects and explains checks, fails closed on uncertainty, emits canonical plans and Visible Value receipts, and classifies fixture-level shadow misses. It has no production adapter, check execution, real-repository benchmark, or trusted selective-verification evidence." + "status": "VERIFIED_NARROW_PROTOTYPE", + "assessment": "The public dependency-free Node.js core validates normalized evidence/catalog/policy input, computes reverse impact, selects and explains checks, fails closed on uncertainty, emits canonical plans and Visible Value receipts, and validates identity-bound SHADOW results. AV-EXP-001 adds one benchmark-only Zustand/Vitest adapter and a frozen full-catalog oracle: both AV arms selected all 8 relevant checks observed across ten scenarios, while the uncertainty case forced full verification. It has no production adapter, independent qualifying replication, or trusted selective-verification class." }, "drift_status": "NEW_ALIGNED_PUBLIC_HOME", "current_name_accuracy": "Accurate for the implemented verification-planning boundary and explicitly broader than affected-test selection.", "evidence_references": [ - "https://github.com/opsle/affected-verification/blob/12076522c9b82501794d816f1fcc0b7775fad6e1/SPEC.md", - "https://github.com/opsle/affected-verification/blob/12076522c9b82501794d816f1fcc0b7775fad6e1/PRIOR_ART.md", - "https://github.com/opsle/affected-verification/blob/12076522c9b82501794d816f1fcc0b7775fad6e1/src/planner.js", - "https://github.com/opsle/affected-verification/blob/12076522c9b82501794d816f1fcc0b7775fad6e1/tests/planner.test.js", - "https://github.com/opsle/affected-verification/pull/1" + "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/SPEC.md", + "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/PRIOR_ART.md", + "https://github.com/opsle/affected-verification/blob/0544362d7659093b7f0b4f89ee8f68023fd269c3/benchmark/av-exp-001/preregistration-v1/preregistration.json", + "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/REPORT.md", + "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/summary.json", + "https://github.com/opsle/affected-verification/pull/2" ], "confidence": "HIGH", "unresolved_questions": [ - "Which real public repository has a complete enough verification catalog and full-run oracle for the first shadow comparison?", - "Which native selector should be the first evidence adapter without duplicating its impact logic?", + "Does the zero-observed-miss result persist in a second public repository and different ecosystem?", + "Which benchmark-only evidence adapter is valuable enough to harden into a production-quality adapter?", "Which evidence-based promotion criteria justify TRUSTED_BOUNDED authority for a narrowly defined change class?" ] } diff --git a/tests/test_validate_program.py b/tests/test_validate_program.py index b7d9ce4..ee2eaa8 100644 --- a/tests/test_validate_program.py +++ b/tests/test_validate_program.py @@ -44,10 +44,25 @@ def test_affected_verification_is_the_twenty_first_repository(self): item for item in self.registry["repositories"] if item["name"] == "affected-verification" ) - self.assertEqual(affected["lifecycle_stage"], "PROTOTYPED") + self.assertEqual(affected["lifecycle_stage"], "VERIFIED") self.assertEqual( affected["last_verified_head_sha"], - "12076522c9b82501794d816f1fcc0b7775fad6e1", + "641aee9d29a89e2a8819f00817ccee8e5d234dcb", + ) + self.assertEqual(affected["active_experiment_ids"], ["AV-EXP-001"]) + + experiment = next( + item for item in self.experiments["experiments"] + if item["id"] == "AV-EXP-001" + ) + self.assertEqual(experiment["status"], "RECORDED") + self.assertEqual( + experiment["benchmark_result"]["summary_identity"], + "sha256:68b8582a9ce7b86bfa5431d89d2dea07f8c34b88d1d0350bab25c99fa5b236df", + ) + self.assertEqual( + experiment["benchmark_result"]["av_core_missed_relevant_check_count"], + 0, ) def test_missing_expected_repository_fails(self):