Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
12 changes: 12 additions & 0 deletions .gitleaks.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
title = "Affected Verification secret-scan policy"

[extend]
useDefault = true

[[allowlists]]
description = "Public immutable Git object identities used by AV-EXP-001"
regexTarget = "match"
regexes = [
'''b57db4f86ef179285da216eeb291266da82c361c''',
'''0544362d7659093b7f0b4f89ee8f68023fd269c3''',
]
13 changes: 10 additions & 3 deletions BENCHMARK.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,9 @@
# Benchmark and research plan
# Benchmark and research record

No benchmark has run and no result is claimed.
AV-EXP-001 completed the first preregistered real-repository `SHADOW`
calibration. See [benchmark/av-exp-001/REPORT.md](benchmark/av-exp-001/REPORT.md)
for the bounded result and immutable artifact identities. The method below
remains authoritative for future runs.

## Central question

Expand Down Expand Up @@ -59,6 +62,10 @@ Unlike verification types are not added into one deceptive percentage. Counterfa
- How frequently do native and composed selectors miss relevant full-run failures?
- Does adding policy and non-test checks produce value above native selectors?

## Initial experiment sequence
## Experiment sequence

First, freeze a real public repository with a trustworthy full-verification baseline and run in `OBSERVE`/`SHADOW`. Compare plan identities and full outcomes without replacing CI. Only after the oracle, catalog completeness, adapter fidelity, and miss classifications are independently reviewable should a bounded trust decision be considered.

AV-EXP-001 satisfied that first calibration step for Zustand at
`b57db4f86ef179285da216eeb291266da82c361c`. It remains `SHADOW`; its absence
of observed AV misses in ten scenarios is not a bounded-trust decision.
8 changes: 4 additions & 4 deletions LIMITATIONS.md
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
# Limitations and non-goals

- The prototype consumes synthetic normalized evidence; it has no Git, graph, coverage, ownership, schema, Nx, Turbo, Jest, Vitest, or testmon adapter.
- The reusable prototype still consumes normalized evidence. AV-EXP-001 adds one narrow benchmark-only Git/catalog/source-graph/Vitest adapter for pinned Zustand; it is not a general or production adapter.
- The component graph and catalog completeness flags are caller claims, not independently attested.
- The selection algorithm uses declared component scope and policy tags. It has no symbol/data-flow analysis, runtime coverage collection, weighted set cover, probabilistic model, or learned judgment.
- Commands are opaque identities. The project does not execute, schedule, cache, distribute, retry, or report CI work.
- `SUFFICIENT_*` means sufficient under the supplied model and policy, not globally safe, formally sound, or mathematically minimal.
- Full catalog selection cannot compensate for an incomplete catalog; that state remains `INSUFFICIENT_EVIDENCE`.
- Shadow relevance is caller supplied. A robust real benchmark needs an independent, reproducible relevance oracle.
- AV-EXP-001 derives shadow relevance reproducibly from changed outcomes in its frozen full catalog. Other callers can still supply relevance, and one corpus does not validate every catalog or oracle.
- Test execution counts are declared catalog metadata. They are exact relative to input, not observed executions.
- No real-project benchmark, comparative baseline, selection-miss corpus, computation measurement, or independent reproduction exists.
- No production trust stage is justified. This repository is at most `PROTOTYPED`.
- One preregistered real-project shadow benchmark, comparative native arm, synthetic selection-miss corpus, and same-host deterministic replay exist. There is no second ecosystem, historical real-change replay, production-quality adapter, or independent qualifying replication.
- No production trust stage is justified. AV-EXP-001 remains `SHADOW`; any lifecycle promotion is bounded to the verified prototype and frozen benchmark evidence.
- No time, token, monetary, correctness, failure-prevention, or causal savings claim is supported.
- Context Firewall, Decision Evidence, Agent Trajectory Profiler, Gearbox, and Opsle Tasks are external consumers or validators, not dependencies.
5 changes: 3 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,13 +25,14 @@ npm run verify
- [SPEC.md](SPEC.md) — normative prototype contract and sufficiency states
- [PRIOR_ART.md](PRIOR_ART.md) — source-linked reconciliation with existing selectors
- [ARCHITECTURE.md](ARCHITECTURE.md) — adapters, project boundaries, shadow mode, and trust ramp
- [BENCHMARK.md](BENCHMARK.md) — controlled research plan; no results are claimed
- [BENCHMARK.md](BENCHMARK.md) — controlled research method and AV-EXP-001 result
- [benchmark/av-exp-001/REPORT.md](benchmark/av-exp-001/REPORT.md) — preregistered real-repository shadow calibration
- [LIMITATIONS.md](LIMITATIONS.md) — current claim ceiling and non-goals
- [fixtures/scenarios.js](fixtures/scenarios.js) and [fixtures/negative-cases.js](fixtures/negative-cases.js) — twelve positive/boundary scenarios plus explicit conflicting, malformed, impossible, and tampered cases
- [schemas/plan-v1.schema.json](schemas/plan-v1.schema.json) — plan shape

## Status

This is a narrow research prototype, not a trusted replacement for full verification. It has not been benchmarked on a real repository, and it has no production adapters.
This is a narrow research prototype, not a trusted replacement for full verification. AV-EXP-001 completed a preregistered shadow calibration on one pinned public repository with no observed AV miss in its frozen corpus. The result remains `SHADOW`, does not establish general safety, and does not provide a production adapter.

Apache-2.0.
177 changes: 177 additions & 0 deletions benchmark/av-exp-001/REPORT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,177 @@
# AV-EXP-001 benchmark report

Experiment: **Minimum Defensible Verification — Real Repository Shadow
Calibration**

## Verdict and claim boundary

The controlled benchmark completed without a design-invalidating defect.
Affected Verification selected all eight checks that the frozen full-catalog
oracle identified as relevant in this ten-scenario corpus. The native Vitest
related-test arm selected six of eight and omitted two relevant non-test
checks. Full verification remained authoritative for every scenario.

This is a `SHADOW` calibration result. It does not establish general safety,
correctness equivalence, causal time or cost savings, or production trust. The
precise safety claim is: **no Affected Verification selection miss was
observed in this frozen calibration corpus**.

## Provenance

- Starting release: `12076522c9b82501794d816f1fcc0b7775fad6e1`
- Preregistration commit: `0544362d7659093b7f0b4f89ee8f68023fd269c3`
- Target: `https://github.com/pmndrs/zustand.git`
- Target commit: `b57db4f86ef179285da216eeb291266da82c361c`
- License: MIT
- Catalog: `sha256:8c5b224deaa7077690341248a18a2155310e2b16072e1607a9b4cd546e3a0914`
- Baseline: `sha256:8844d3be4b27f9abd9bab7c6b04d34fc6a1e7cfe36a21e6deec2b315633d3465`
- Result summary: `sha256:68b8582a9ce7b86bfa5431d89d2dea07f8c34b88d1d0350bab25c99fa5b236df`
- Analysis: `sha256:c431d8849edce79d6121a290f49288ed2710e406600d8b66d3588e6b82c73a1d`
- Evidence bundle: `sha256:1e176b7a40b5f16451797d87784f560f932b686f2fe261731526709331ff1172`

The original preregistration is immutable. Three versioned amendments preserve
the initial failed attempts: worktree dependency preparation, Vitest JSON
report interpretation and timing-excluded semantic identities, and output-path
normalization. None changed the target, corpus, oracle, arms, selection policy,
metrics, or analysis rules.

## Target, environment, and baseline

Zustand was selected after evaluating Redux Toolkit and Immer against the
preregistered criteria. It offered a pinned pnpm lockfile and toolchain, a
bounded primary CI catalog, enumerated Vitest checks, and Vitest's supported
`related` selector. Redux Toolkit's multi-matrix/package surface and Immer's
built-artifact, Flow, and performance workflow ambiguity made them less
suitable for this first bounded calibration; they were not benchmarked.

The benchmark ran on Linux x86-64 with Node 24.20.0 and pnpm 11.3.0. Frozen
target tools included Vitest 4.1.10, TypeScript 6.0.3, ESLint 9.39.4, Prettier
3.9.6, and Rollup 4.62.4. Exact environment and source hashes are in the
preregistration.

The full catalog contained 17 compatible check units: 13 test-file checks
representing 224 test executions, plus format, typecheck, lint, and build.
Three clean baseline repetitions passed all checks with the same semantic
outcome and no observed flaky check. Full-catalog wall time was 29.329–32.807
seconds across the three repetitions (`OBSERVED`).

## Corpus and oracle

The corpus contains ten frozen scenarios: six deterministic synthetic faults
and four synthetic benign change shapes. It covers isolated and direct changes,
reverse/transitive dependencies, shared code, test-only changes,
configuration, dependency/build metadata, public contracts, multi-file
changes, and deliberately incomplete impact evidence. Each scenario binds its
base SHA, patch hash, changed paths, classification, setup, and materialization
procedure. No selector input contains the full-run outcome.

For each scenario, the oracle first relies on the passing clean baseline,
materializes the frozen patch, executes all 17 frozen catalog checks, and marks
as relevant the checks whose outcome changed because of that scenario. The
claim is only **relevant within this frozen verification catalog**; the catalog
does not prove complete software correctness.

## Arms and safety result

`FULL` ran every catalog check and remained authoritative. `NATIVE` used
Vitest 4.1.10 `related` with the actual changed paths. `AV_CORE` used normalized
Git change, catalog, package/source import graph, and policy evidence.
`AV_WITH_NATIVE_EVIDENCE` additionally consumed native related-test output as
one evidence source.

| Selector | Relevant | Selected | Missed | Scenario misses | Full escalations |
|---|---:|---:|---:|---:|---:|
| FULL | 8 | 8 | 0 | 0 | n/a |
| NATIVE | 8 | 6 | 2 | 2 | 0 |
| AV_CORE | 8 | 8 | 0 | 0 | 3 |
| AV_WITH_NATIVE_EVIDENCE | 8 | 8 | 0 | 0 | 3 |

Every observed selection miss:

- `NATIVE`, `AVS-003`: omitted relevant `check:lint`.
- `NATIVE`, `AVS-008`: omitted relevant `check:typecheck`.

The misses demonstrate the boundary of a tests-only native selector; they are
not characterized as Vitest defects. Both AV arms added the four policy-required
non-test checks in every scenario and selected both relevant checks. AV
broadened for public-contract changes in AVS-003 and AVS-008. It required the
full catalog for verification metadata/configuration in AVS-006 and AVS-007,
and for deliberately incomplete dependency evidence in AVS-010. Aggressive
targeted selection was therefore denied in the uncertainty scenario.

## Workload result

These are `EXACT` plan counts aggregated over ten scenarios. Unlike units are
kept separate and the counts are not correctness, cost, or causal savings.

| Arm | Test files selected / available | Test executions selected / available | Non-test checks selected / available |
|---|---:|---:|---:|
| FULL | 130 / 130 | 2,240 / 2,240 | 40 / 40 |
| NATIVE | 70 / 130 | 1,358 / 2,240 | 0 / 40 |
| AV_CORE | 89 / 130 | 1,640 / 2,240 | 40 / 40 |
| AV_WITH_NATIVE_EVIDENCE | 89 / 130 | 1,640 / 2,240 | 40 / 40 |

AV proposed skipping 41 of 130 test-file executions and 600 of 2,240 test
executions while retaining every oracle-relevant check observed in the corpus.
Native evidence did not change AV's test count in this corpus. Native selected
fewer tests in AVS-006 and AVS-010 because AV correctly broadened under policy
or uncertainty. Per-scenario counts and every skip explanation are preserved
in `results-v2/scenarios/`.

Scenario full-catalog time was 29.088–33.442 seconds and native selection time
was 1.716–5.683 seconds (`OBSERVED`). They are telemetry, not a causal
time-saved estimate.

## Shadow evidence, visibility, and reproducibility

Every result binds the experiment, benchmark revision, target and scenario,
catalog, planner and policy, native selector, four arm selections, full oracle,
misses, workload, uncertainty, evidence hashes, and result identity. The
validator rejects target, patch, catalog, selector, baseline, oracle, policy,
scenario, skip-reason, trust-state, and result tampering, including attempts to
represent `SHADOW` as trusted.

Twenty `opsle.value-receipt.v1` records provide one concise Affected
Verification indicator per AV arm and scenario. Test/check counts are `EXACT`;
wall time is `OBSERVED`. Raw full-run outputs remain in `results-v2/raw/` even
though the published summary reduces what an operator needs to inspect. Thus
Affected Verification answers what should run, while the Context Firewall
pattern answers what result should be shown; output reduction never changes
oracle truth.

The deterministic reproduction command is:

```bash
npm run benchmark:av-exp-001
```

It prepares the exact target in disposable storage, validates identity,
installs locked dependencies, validates the catalog, materializes each patch,
runs native and AV selection, executes the full catalog, computes relevance,
and emits results. A second clean run at a different result path reproduced the
baseline, all ten scenario, summary, and analysis identities exactly. This was
a same-host, same-implementation replay, not an independent qualifying
replication. Raw bundle identities intentionally include variable observed
timing and output locations.

## Opsle mechanisms and decision

Actually exercised implementations: Affected Verification planning and
evidence validation; Visible Value receipts and the authoritative receipt
validator. The Context Firewall separation was exercised as a manual artifact
pattern: raw truth was retained while summaries constrained supervising
context. Decision Evidence concepts were implemented through content-addressed
bindings and validation, but its separate runtime was not invoked. Agent
Trajectory Profiler, Gearbox, and Opsle Tasks were not used. No Codex child
agents or external model/provider workloads were used.

Affected Verification remains `OBSERVE/SHADOW`. This one-repository,
one-ecosystem, synthetic-fault-heavy corpus and same-host replay do not justify
`TRUSTED_BOUNDED`. The evidence can support lifecycle `VERIFIED` for the narrow,
revision-bound prototype and benchmark artifacts if all release and program
gates pass; it cannot support a higher stage.

Exactly one next execution: preregister and run a second public-repository
shadow calibration in a different ecosystem with a meaningful native selector
and the same full-catalog oracle discipline. Do not begin it as part of this
experiment.
34 changes: 34 additions & 0 deletions benchmark/av-exp-001/amendments/001-worktree-install.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# AV-EXP-001 benchmark amendment 001

Status: harness-only amendment before comparative scenario execution

The first baseline attempt after preregistration commit
`0544362d7659093b7f0b4f89ee8f68023fd269c3` stopped before any scenario.
All three clean repetitions were incomplete because the runner symlinked the
prepared clone's `node_modules` into Git worktrees. pnpm correctly detected that
the modules directory belonged to a different workspace and attempted a
modules-dir repair. In the non-interactive session it refused that repair with
`ERR_PNPM_ABORTED_REMOVE_MODULES_DIR_NO_TTY`.

This is a harness materialization defect. It does not change the target,
catalog, native selector, scenarios, patches, oracle, arms, metrics, analysis
rules, or stop conditions frozen in preregistration v1. No native selection,
AV plan, scenario full run, oracle, or comparative result was produced.

Amendment:

1. each disposable worktree performs
`corepack pnpm install --frozen-lockfile --offline` against the already
populated content-addressed pnpm store;
2. `CI=true` is set for non-interactive command execution;
3. no `node_modules` symlink is created;
4. the failed baseline artifacts are retained under
`attempts/harness-defect-001/`;
5. baseline stability is rerun from three new clean worktrees before any
scenario.

Result bundles bind this amendment's content identity in addition to the
original preregistration commit.

Amendment identity:
`sha256:80a04b99ad73e86ecb2e7c85dda3a11ddbe99cdf200ca38e2c1effe498184357`
32 changes: 32 additions & 0 deletions benchmark/av-exp-001/amendments/002-reporter-semantics.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# AV-EXP-001 benchmark amendment 002

Status: harness-only amendment before comparative scenario execution

After amendment 001, three new clean baseline repetitions each completed all
17 catalog checks, passed 224 of 224 tests, and had no failed checks. The gate
still stopped before scenarios because the runner treated Vitest's
`numTotalTestSuites` as a test-file count. Vitest 4.1.10 reported 66 nested
suites and separately emitted 13 per-file `testResults`, matching the frozen
catalog.

Review of the same baseline artifacts also found that raw stdout/stderr hashes
were included in semantic full-run and benchmark-result identities. Raw reports
may contain observed timestamps or durations, so this violated the frozen rule
that semantic identities must not depend on timing.

Amendment:

1. derive the test-file count from normalized `testResults.length` and continue
requiring exactly 13 frozen file outcomes and 224 test cases;
2. retain and hash raw output as observational evidence, but exclude raw
evidence hashes from semantic full-run and benchmark-result identities;
3. continue binding command identities, exit states, normalized per-check
outcomes, catalog, scenario, oracle, planner, policy, and selector identities;
4. retain this stopped baseline attempt under `attempts/harness-defect-002/`;
5. rerun three fresh clean baselines before any scenario.

This amendment does not change any target, catalog, selector, scenario, patch,
oracle, metric, analysis, or stop rule frozen at preregistration.

Amendment identity:
`sha256:51437d6c3bfdd1461a9679a5055ab5c34d9eb603ba95b67f86f8516f8b93a25e`
29 changes: 29 additions & 0 deletions benchmark/av-exp-001/amendments/003-artifact-path-normalization.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
# AV-EXP-001 benchmark amendment 003

Status: semantic-identity amendment after first complete shadow run

The first complete result bundle and an immediate second clean reproduction had
identical baseline pass states, normalized check outcomes, relevance sets,
selector selections, misses, workloads, planner version, and policy behavior.
Their semantic identities differed because the full-run command record included
the absolute Vitest `--outputFile` path. The published run wrote inside the
repository while the reproduction wrote under `/tmp`.

Amendment:

1. preserve exact commands and absolute output locators in raw execution
evidence;
2. normalize only the semantic command token to `--outputFile=<artifact>`;
3. preserve the first complete bundle under
`attempts/semantic-identity-defect-003/`;
4. publish the corrected benchmark revision as `results-v2/`;
5. rerun the complete baseline and scenario corpus, then run a second complete
reproduction and require identical semantic baseline, scenario, analysis,
and summary identities.

No target, catalog, selector, scenario, patch, oracle, metric, analysis, stop,
selection, workload, or trust rule changes. The first complete outcome remains
evidence but its path-dependent identities are superseded.

Amendment identity:
`sha256:44f420409b719b5ddfaac04928e215572873475c8d8e720e4755e043b878dfde`
Loading
Loading