diff --git a/PROGRAM_STATUS.md b/PROGRAM_STATUS.md index 4efbc99..82aff81 100644 --- a/PROGRAM_STATUS.md +++ b/PROGRAM_STATUS.md @@ -4,7 +4,7 @@ **Coverage: 21/21 expected repositories; duplicates: 0.** -Last verified: `2026-08-31T02:59:01Z`. HEADs are the verified default-branch revisions, not an assumption about later changes. +Last verified: `2026-08-31T12:52:33Z`. HEADs are the verified default-branch revisions, not an assumption about later changes. ## Portfolio totals @@ -43,7 +43,7 @@ Program state totals: active 5; waiting 16; complete 0. | 15 | [agent-recovery-policy](https://github.com/opsle/agent-recovery-policy) | concept | `1b733a111e26` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/agent-recovery-policy/blob/1b733a111e26e0a409fee3b96f627048531daefe/THEORY.md); placeholder only; no automated tests | No shared failure schema, attempt ledger, route evaluator, or comparative fixture set. | After decision evidence and route schemas stabilize, define same-failure convergence on synthetic failures. | `agent-routing-policy`, `agent-state-ledger`, `decision-evidence-protocol` | waiting | | 16 | [ephemeral-agent-workers](https://github.com/opsle/ephemeral-agent-workers) | concept | `ad96fcfdfac0` | `THEORY` | [none; placeholder source directory only](https://github.com/opsle/ephemeral-agent-workers/blob/ad96fcfdfac06d340b5e96d369634980cee78ef4/THEORY.md); placeholder only; no automated tests | Portable authority, claim, and handoff contracts are not ready; no safe synthetic containment harness exists. | Wait for prerequisite contracts, then define a fake worker adapter and destruction receipt without infrastructure changes. | `agent-execution-authorization`, `agent-resource-claims`, `verifiable-agent-handoff` | waiting | | 17 | [gearbox](https://github.com/opsle/gearbox) | concept | `f3fab9f292cf` | `PROTOTYPED` | [provider-free Python reference core with strict authority-policy admission, exact deterministic argv execution, content-addressed staged helper context, injected one-shot helper transport, passive process waiting, compact results, raw-artifact accounting, fail-closed budgets, and Visible Value receipts](https://github.com/opsle/gearbox/blob/f3fab9f292cf4eabd7200615d444f98881f57d55/src/opsle_gearbox/core.py); 19 of 19 provider-free automated tests passed locally, in PR #1 CI, and in final-main CI; ruff, shellcheck, actionlint, gitleaks, wheel build, receipt validation, and public raw-locator/hash checks passed | A production-quality bounded helper transport, independently verified isolation and termination, full Context Firewall integration, and a frozen comparative benchmark remain missing. | Freeze a provider-free deterministic-versus-direct baseline and helper-transport conformance corpus before considering any live model/provider run. | `context-firewall`, `decision-evidence-protocol`, `agent-trajectory-profiler`, `agent-routing-policy`, `agent-execution-authorization` | waiting | -| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `641aee9d29a8` | `VERIFIED` | [dependency-free Node.js deterministic planner plus an AV-EXP-001 benchmark-only Git/catalog/source-graph/Vitest adapter, identity-bound SHADOW result validator, complete frozen-oracle harness, explainable skip records, fail-closed uncertainty handling, and opsle.value-receipt.v1 telemetry](https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/REPORT.md); 70 of 70 automated tests, 15 of 15 conformance scenarios, and 6 of 6 determinism checks passed locally, in PR #2 CI, and from a fresh detached worktree at exact main; invalid-state coverage includes target, patch, catalog, selector, adapter, baseline, oracle, scenario, skip-reason, policy, trust-state, and result tampering | A second ecosystem, historical real-change replay, production-quality evidence adapter, and independent qualifying replication remain missing; AV therefore remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized. | Preregister and run a second public-repository shadow calibration in a different ecosystem with a meaningful native selector and the same full-catalog oracle discipline. | — | active | +| 18 | [affected-verification](https://github.com/opsle/affected-verification) | concept | `3ff41688dded` | `VERIFIED` | [dependency-free Node.js deterministic planner plus benchmark-only Git/catalog/source-graph adapters for AV-EXP-001 Vitest and AV-EXP-002 Python/pytest-testmon, identity-bound SHADOW result validation, complete frozen-oracle harnesses, explainable skip records, fail-closed uncertainty handling, and opsle.value-receipt.v1 telemetry](https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/REPORT.md); 86 of 86 automated tests, 15 of 15 conformance scenarios, and 7 of 7 determinism checks passed locally, in PR #3 CI, and from a fresh detached worktree at exact main; the AV-EXP-002 result validator also passed at exact main, and invalid-state coverage includes Python/toolchain, target, patch, catalog, selector state/version/output, dynamic/conftest uncertainty, baseline, oracle, scenario, skip, policy, trust-state, and result tampering | AV-EXP-002 observed a false targeted-sufficiency claim for runtime/subprocess import behavior; historical real-change replay, a production-quality evidence adapter, and independent qualifying replication also remain missing. AV remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized. | Preregister and execute a selection-miss repair for AV2-006 that represents runtime/subprocess import uncertainty without changing the preserved AV-EXP-002 result. | — | active | | 19 | [research](https://github.com/opsle/research) | program infrastructure | `9ee43197880c` | `PROTOTYPED` | [authoritative 21-repository ledger, machine-readable 18-concept theory registry including Affected Verification, canonical theory map, normative Visible Value controls, and provider-free EXP-001 benchmark, launch, one-block coordinator, external four-label LIVE_PROVIDER_RUN authorization, and current catalogue/pricing preflight artifacts with six content-addressed tasks, deterministic oracle, four arm contracts, sealed blinded allocation, exact subject configuration and adapter, exact authorization admission, private boundaries, receipts, mutation tests, and integrity CI](program/THEORY_MAP.md); 88 of 88 repository tests pass locally after deliberate migration to the 21-repository, 18-concept anti-forgetting set, including 13 authorization validations and two byte-identical replays; generated status and registry validation pass | The exact live authorization set remains unconsumed and unreleased, account-specific API entitlement is unverified under the zero-provider-call policy, no immutable dated model snapshot is documented, and the program has no canonical measured concept experiment. | Independently review and release the provider-free live-authorization and catalogue/pricing preflight; do not consume authorization or launch a provider/model subject. | — | active | | 20 | [site](https://github.com/opsle/site) | program infrastructure | `28ad65be4750` | `PROTOTYPED` | [React/Vinext source implementation with content routes](https://github.com/opsle/site/blob/28ad65be4750dc849976fbf5c9eae9501c6bbb25/README.md); automated build/render tests present; not rerun because this reconciliation kept other repositories read-only | Wait for validated registry data and measured research; deployment requires separate authorization. | After registry merge, add a read-only registry ingestion design without deploying the site. | `research` | waiting | | 21 | [.github](https://github.com/opsle/.github) | program infrastructure | `01c38e726db7` | `THEORY` | [documentation-only organization profile](https://github.com/opsle/.github/blob/01c38e726db7c3e45059d25fccce55e071e35938/profile/README.md); not applicable to current single Markdown profile; consistency is unverified | No mechanical registry consistency check exists in this repository. | After registry merge, design a read-only consistency check for organization-profile repository links. | `research` | waiting | diff --git a/program/experiments.json b/program/experiments.json index 292884c..5312ead 100644 --- a/program/experiments.json +++ b/program/experiments.json @@ -1,6 +1,6 @@ { "schema_version": 1, - "last_verified_at": "2026-08-31T02:59:01Z", + "last_verified_at": "2026-08-31T12:52:33Z", "experiments": [ { "id": "EXP-001", @@ -343,6 +343,119 @@ "lifecycle_impact": "PROMOTE_TO_VERIFIED_ONLY: the narrow scoped correctness and safety claims pass meaningful automated and revision-bound benchmark checks, but this run is explicitly capped below BENCHMARK_READY and does not authorize trusted execution.", "next_task": "Preregister and run a second public-repository shadow calibration in a different ecosystem with a meaningful native selector and the same full-catalog oracle discipline." }, + { + "id": "AV-EXP-002", + "title": "Cross-Ecosystem Minimum Defensible Verification Shadow Calibration", + "status": "RECORDED", + "hypothesis": "On a pinned Python repository with an established ecosystem affected-test selector, Affected Verification preserves every oracle-relevant frozen catalog check while proposing less than FULL and broadening when evidence is incomplete.", + "participating_repositories": [ + "affected-verification", + "research" + ], + "roles": { + "primary": "affected-verification", + "expected_support": [ + "research" + ], + "potential_support": [] + }, + "baseline": "The complete frozen 2,024-check Click verification catalog at 36baa15ff831b939a22bc527cd76ce653ef6f66d, containing 2,016 pytest nodes and eight non-test checks, executed for every scenario and authoritative over all selector predictions.", + "experimental_arms": [ + "FULL frozen verification catalog", + "ECOSYSTEM_SELECTOR using pytest-testmon 2.2.0 under its test-selection contract", + "AV_CORE with normalized Git, Python import graph, pytest catalog, verification catalog, and policy evidence", + "AV_WITH_SELECTOR_EVIDENCE with pytest-testmon output as an additional normalized evidence source" + ], + "primary_metric": "Selection misses reported individually against checks whose full-catalog outcome changed because of a frozen scenario.", + "secondary_metrics": [ + "relevant-check recall and scenario-level misses", + "exact selected and skipped pytest nodes, test files, and non-test checks by compatible class", + "uncertainty broadening and full-verification escalation", + "observed wall-clock telemetry without a causal time-saved claim", + "static-graph versus runtime selector compensation", + "normalized comparison to AV-EXP-001 without aggregating incompatible units" + ], + "correctness_gate": "Every frozen catalog check runs for every scenario after selector and AV proposals are frozen; FULL remains authoritative, and a miss is any omitted oracle-relevant failing check.", + "failure_classifications": [ + "outside selector contract", + "dependency evidence miss", + "verification-class omission", + "policy omission", + "adapter defect", + "planner defect", + "oracle or harness defect", + "unresolved", + "conservative broadening", + "insufficient evidence" + ], + "dataset_fixture_identity": "Click 36baa15ff831b939a22bc527cd76ce653ef6f66d; preregistration commit f8a183c460535f3352fad2fb4990b0c54818d623; catalog sha256:28ed20abf60e7c785052308298dc6ed647b7a20a737513e8fb3c76aa62d9094c; corpus sha256:d5bc43405a5ab0ac34feef6d5fd5df111eace7f1a355400f963f8a7d4399640b; eleven frozen scenarios and patches.", + "model_provider_configuration": "NONE: no model/provider benchmark subject or external provider workload was used; one interactive Codex session used native shell and patch facilities without child agents.", + "run_identities": [ + "sha256:5b3f99bfbebd3a0d061651d66adfb5a6aaef899475c6267cb20e4040e6ed5768" + ], + "result_artifacts": [ + "https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/REPORT.md", + "https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/results-v1/summary.json", + "https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/results-v1/cross-experiment.json", + "https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/results-v1/evidence-manifest.json" + ], + "target": { + "repository": "https://github.com/pallets/click.git", + "sha": "36baa15ff831b939a22bc527cd76ce653ef6f66d", + "license": "BSD-3-Clause" + }, + "preregistration": { + "commit_sha": "f8a183c460535f3352fad2fb4990b0c54818d623", + "identity": "sha256:93743e15caee647de1807ee27d35518cd9520394983803e1b93a58e9a2841db6", + "amendment_count": 5, + "comparative_outcomes_observed_before_commit": false + }, + "benchmark_result": { + "affected_verification_main_sha": "3ff41688dded6e96e65da7cc44fe2608cf86d073", + "results_commit_sha": "5d126e0ae17557065b55ab84a46f6a3577a49989", + "summary_identity": "sha256:5b3f99bfbebd3a0d061651d66adfb5a6aaef899475c6267cb20e4040e6ed5768", + "evidence_bundle_identity": "sha256:435d8e5356ed6868edfc1747523ff75d1b327389bf14cc2675e867e45f4de705", + "cross_experiment_identity": "sha256:6ac2695ee80c9cb711e756846dad4c7138a3177c11d733456636259c36948da0", + "selector_baseline_identity": "sha256:9a3eb5630a868869179585210e9e6034e0c16039d97874921a09f63ef61aafd8", + "scenario_count": 11, + "synthetic_fault_count": 7, + "synthetic_benign_change_count": 1, + "uncertainty_scenario_count": 3, + "relevant_check_count": 87, + "ecosystem_selector_selected_relevant_check_count": 77, + "ecosystem_selector_missed_relevant_check_count": 10, + "ecosystem_selector_scenario_miss_count": 7, + "av_core_selected_relevant_check_count": 86, + "av_core_missed_relevant_check_count": 1, + "av_core_scenario_miss_count": 1, + "av_core_full_broadening_count": 6, + "av_with_selector_selected_relevant_check_count": 86, + "av_with_selector_missed_relevant_check_count": 1, + "av_with_selector_scenario_miss_count": 1, + "av_with_selector_full_broadening_count": 5, + "av_miss": { + "scenario_id": "AV2-006", + "check_id": "pytest:tests/test_imports.py::test_light_imports", + "classification": "PLANNER_OR_ADAPTER_MISS", + "finding": "A subprocess instrumented runtime imports through the public click package; the static Python graph declared completeness and pytest-testmon also omitted the relevant node." + } + }, + "major_findings": [ + "Catalog, policy escalation, explicit uncertainty, skip evidence, FULL fallback, shadow classification, and Visible Value generalized without changing the core plan schema.", + "Python dependency evidence, pytest node identity, conftest coupling, selector database lifecycle, and dynamic or subprocess imports required material ecosystem-specific logic.", + "Runtime selector evidence compensated for static dynamic-plugin uncertainty in AV2-009 but could not repair AV2-006 because the selector also omitted the relevant runtime-import test.", + "Both deliberately degraded cases denied aggressive skipping and required FULL." + ], + "replication_status": "SAME_HOST_RELEASE_AND_BUNDLE_VERIFICATION_ONLY", + "verdict": "FAIL for the safety hypothesis: AV_CORE and AV_WITH_SELECTOR_EVIDENCE each selected 86/87 oracle-relevant checks and omitted the same runtime/subprocess import test in AV2-006. The controlled shadow calibration itself completed and FULL exposed the miss.", + "blockers": [ + "The AV2-006 runtime/subprocess import miss invalidates a positive targeted-sufficiency claim for the frozen corpus.", + "The Python and pytest-testmon adapters are benchmark-only and no historical real-change replay or independent qualifying replication exists.", + "Affected Verification remains OBSERVE/SHADOW; no TRUSTED_BOUNDED class is authorized." + ], + "lifecycle_impact": "REMAIN_VERIFIED: the run adds revision-bound falsification evidence and failure modes, but the observed selection miss and explicit run cap do not establish BENCHMARK_READY or EXPERIMENTED lifecycle promotion for the repository.", + "next_task": "Preregister and execute a selection-miss repair for AV2-006 that represents runtime/subprocess import uncertainty without changing the preserved AV-EXP-002 result." + }, { "id": "LEGACY-001", "title": "Graphify plus Antigravity semantic adapter integration observation", diff --git a/program/registry.json b/program/registry.json index e2a7913..22a74e8 100644 --- a/program/registry.json +++ b/program/registry.json @@ -38,7 +38,7 @@ }, "current_highest_priority_workstream": "Independently review and release the provider-free EXP-001 live-authorization and current catalogue/pricing preflight; do not consume authorization or launch any provider/model subject.", "recommended_next_execution": "In opsle/research, create and provider-free validate one exact four-label LIVE_PROVIDER_RUN authorization set plus a model catalogue/pricing preflight artifact; do not consume authorization or launch a provider/model subject.", - "last_verified_at": "2026-08-31T02:59:01Z", + "last_verified_at": "2026-08-31T12:52:33Z", "repositories": [ { "name": "agent-trajectory-profiler", @@ -554,31 +554,31 @@ "name": "affected-verification", "github_url": "https://github.com/opsle/affected-verification", "default_branch": "main", - "last_verified_head_sha": "641aee9d29a89e2a8819f00817ccee8e5d234dcb", + "last_verified_head_sha": "3ff41688dded6e96e65da7cc44fe2608cf86d073", "project_type": "concept", "purpose": "Select the smallest verification workload whose sufficiency can be defended from available change-impact, dependency, coverage, policy, and risk evidence.", "lifecycle_stage": "VERIFIED", - "implementation_status": "dependency-free Node.js deterministic planner plus an AV-EXP-001 benchmark-only Git/catalog/source-graph/Vitest adapter, identity-bound SHADOW result validator, complete frozen-oracle harness, explainable skip records, fail-closed uncertainty handling, and opsle.value-receipt.v1 telemetry", + "implementation_status": "dependency-free Node.js deterministic planner plus benchmark-only Git/catalog/source-graph adapters for AV-EXP-001 Vitest and AV-EXP-002 Python/pytest-testmon, identity-bound SHADOW result validation, complete frozen-oracle harnesses, explainable skip records, fail-closed uncertainty handling, and opsle.value-receipt.v1 telemetry", "implementation_requirement": "A runnable deterministic planner, input validator, plan contract, conformance fixtures, and shadow classifier are sufficient for the prototype gate; real adapters and comparative evidence are later gates.", "specification_status": "versioned input and opsle.affected-verification.plan.v1 contracts define catalog entries, evidence providers, policy matching, selection and skip reasons, provenance, explicit sufficiency/uncertainty/escalation states, failure behavior, Visible Value limits, and shadow observations", - "test_status": "70 of 70 automated tests, 15 of 15 conformance scenarios, and 6 of 6 determinism checks passed locally, in PR #2 CI, and from a fresh detached worktree at exact main; invalid-state coverage includes target, patch, catalog, selector, adapter, baseline, oracle, scenario, skip-reason, policy, trust-state, and result tampering", - "benchmark_status": "AV-EXP-001 preregistered a pinned Zustand b57db4f86ef179285da216eeb291266da82c361c SHADOW calibration with a 17-check full catalog, three stable clean baselines, four arms, ten frozen scenarios, a deterministic relevance oracle, explicit miss taxonomy, raw evidence, and content-addressed results; AV_CORE and AV_WITH_NATIVE_EVIDENCE selected 8/8 relevant checks with zero observed misses, while NATIVE selected 6/8 and missed relevant lint and typecheck checks", - "measured_experiment_status": "AV-EXP-001 RECORDED; exact plan-count and observed-runtime evidence exists for one repository and ecosystem, six synthetic fault scenarios and four benign synthetic change shapes; zero provider/model runs", - "reproducibility_status": "the public one-command harness and result verifier reproduce the frozen target preparation, catalog, scenarios, native and AV arms, full oracle, receipts, and summary; a second same-host clean run matched the baseline, all ten scenario, summary, and analysis semantic identities, but is not an independent qualifying replication", + "test_status": "86 of 86 automated tests, 15 of 15 conformance scenarios, and 7 of 7 determinism checks passed locally, in PR #3 CI, and from a fresh detached worktree at exact main; the AV-EXP-002 result validator also passed at exact main, and invalid-state coverage includes Python/toolchain, target, patch, catalog, selector state/version/output, dynamic/conftest uncertainty, baseline, oracle, scenario, skip, policy, trust-state, and result tampering", + "benchmark_status": "AV-EXP-002 preregistered and recorded an eleven-scenario Click 36baa15ff831b939a22bc527cd76ce653ef6f66d SHADOW calibration with pytest-testmon 2.2.0, a 2,024-check full catalog, three stable clean baselines, four arms, fail-closed attacks, individual miss classification, raw evidence, and cross-experiment normalization; FULL selected 87/87 relevant checks, ECOSYSTEM_SELECTOR 77/87, and both AV arms 86/87, with the same runtime/subprocess import test omitted in AV2-006", + "measured_experiment_status": "AV-EXP-001 and AV-EXP-002 RECORDED; exact compatible-unit plan counts and observed runtime evidence exist for pinned JavaScript/Vitest and Python/pytest-testmon targets; AV-EXP-002 falsified perfect AV recall on its frozen corpus with one miss in each AV arm; zero provider/model runs", + "reproducibility_status": "the public AV-EXP-002 bounded harness and result verifier reproduce pinned target validation, locked environment, catalog, selector baseline, scenarios, four shadow arms, FULL oracle, receipts, semantic result, and aggregate report; the published bundle and exact merged SHA were independently validated on this host, but no independent qualifying replication exists", "documentation_status": "public canonical definition, normative specification, source-linked prior-art audit, architecture and independence boundaries, verification catalog and policy semantics, trust ramp, benchmark plan, limitations, usage, security, schema, and fixtures are present", "site_publication_status": "GitHub documentation only; no evidence-backed site publication", - "known_limitations": ["AV-EXP-001 is one repository, one TypeScript/Vitest ecosystem, six synthetic faults and four benign synthetic change shapes; its adapters are benchmark-only, native evidence did not change AV's test count, the replay used the same host and implementation, and no result establishes general safety, correctness equivalence, causal time or cost savings, production trust, or a production-quality adapter."], + "known_limitations": ["Two pinned repositories and ecosystems have synthetic shadow evidence, but the adapters remain benchmark-only. AV-EXP-002 exposed a runtime/subprocess import dependency miss that both the static Python graph and pytest-testmon omitted. There is no historical real-change replay, production-quality adapter, independent qualifying replication, general safety proof, correctness equivalence, causal time or cost saving, or production trust."], "dependencies": [], "dependents": [], - "active_experiment_ids": ["AV-EXP-001"], - "blockers": ["A second ecosystem, historical real-change replay, production-quality evidence adapter, and independent qualifying replication remain missing; AV therefore remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized."], - "next_task": "Preregister and run a second public-repository shadow calibration in a different ecosystem with a meaningful native selector and the same full-catalog oracle discipline.", - "evidence": ["https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/REPORT.md", "https://github.com/opsle/affected-verification/blob/0544362d7659093b7f0b4f89ee8f68023fd269c3/benchmark/av-exp-001/preregistration-v1/preregistration.json", "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/summary.json", "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/analysis.json", "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/results-v2/evidence-manifest.json", "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/verify-results.mjs", "https://github.com/opsle/affected-verification/pull/2", "https://github.com/opsle/affected-verification/actions/runs/33352268190"], + "active_experiment_ids": ["AV-EXP-001", "AV-EXP-002"], + "blockers": ["AV-EXP-002 observed a false targeted-sufficiency claim for runtime/subprocess import behavior; historical real-change replay, a production-quality evidence adapter, and independent qualifying replication also remain missing. AV remains OBSERVE/SHADOW and no TRUSTED_BOUNDED change class is authorized."], + "next_task": "Preregister and execute a selection-miss repair for AV2-006 that represents runtime/subprocess import uncertainty without changing the preserved AV-EXP-002 result.", + "evidence": ["https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/REPORT.md", "https://github.com/opsle/affected-verification/blob/f8a183c460535f3352fad2fb4990b0c54818d623/benchmark/av-exp-002/preregistration-v1/preregistration.json", "https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/results-v1/summary.json", "https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/results-v1/cross-experiment.json", "https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/results-v1/evidence-manifest.json", "https://github.com/opsle/affected-verification/blob/3ff41688dded6e96e65da7cc44fe2608cf86d073/benchmark/av-exp-002/verify-results.mjs", "https://github.com/opsle/affected-verification/pull/3", "https://github.com/opsle/affected-verification/actions/runs/33393661890", "https://github.com/opsle/affected-verification/blob/641aee9d29a89e2a8819f00817ccee8e5d234dcb/benchmark/av-exp-001/REPORT.md"], "completion_criteria": ["Publish production evidence adapters and independently validate their completeness boundaries.", "Run correctness-first comparisons against full verification and native selectors with durable shadow miss evidence.", "Replicate bounded workload and safety claims on independent repositories or environments."], "completion_evidence": [], "completion_status": "INCOMPLETE", "program_state": "active", - "last_verified_at": "2026-08-31T02:59:01Z" + "last_verified_at": "2026-08-31T12:52:33Z" }, { "name": "research", diff --git a/tests/test_validate_program.py b/tests/test_validate_program.py index ee2eaa8..a5425bd 100644 --- a/tests/test_validate_program.py +++ b/tests/test_validate_program.py @@ -47,9 +47,12 @@ def test_affected_verification_is_the_twenty_first_repository(self): self.assertEqual(affected["lifecycle_stage"], "VERIFIED") self.assertEqual( affected["last_verified_head_sha"], - "641aee9d29a89e2a8819f00817ccee8e5d234dcb", + "3ff41688dded6e96e65da7cc44fe2608cf86d073", + ) + self.assertEqual( + affected["active_experiment_ids"], + ["AV-EXP-001", "AV-EXP-002"], ) - self.assertEqual(affected["active_experiment_ids"], ["AV-EXP-001"]) experiment = next( item for item in self.experiments["experiments"] @@ -65,6 +68,24 @@ def test_affected_verification_is_the_twenty_first_repository(self): 0, ) + experiment = next( + item for item in self.experiments["experiments"] + if item["id"] == "AV-EXP-002" + ) + self.assertEqual(experiment["status"], "RECORDED") + self.assertEqual( + experiment["benchmark_result"]["summary_identity"], + "sha256:5b3f99bfbebd3a0d061651d66adfb5a6aaef899475c6267cb20e4040e6ed5768", + ) + self.assertEqual( + experiment["benchmark_result"]["av_core_missed_relevant_check_count"], + 1, + ) + self.assertEqual( + experiment["benchmark_result"]["av_miss"]["check_id"], + "pytest:tests/test_imports.py::test_light_imports", + ) + def test_missing_expected_repository_fails(self): registry = copy.deepcopy(self.registry) registry["repositories"].pop()