Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 42 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
name: CI

on:
pull_request:
push:
branches:
- main

permissions:
contents: read

jobs:
verify:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: actions/setup-node@v4
with:
node-version: 20
- run: npm run verify

secret-scan:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Install pinned gitleaks
env:
GITLEAKS_VERSION: 8.30.1
run: |
base="https://github.com/gitleaks/gitleaks/releases/download/v${GITLEAKS_VERSION}"
archive="gitleaks_${GITLEAKS_VERSION}_linux_x64.tar.gz"
checksums="gitleaks_${GITLEAKS_VERSION}_checksums.txt"
curl -fsSLO "${base}/${archive}"
curl -fsSLO "${base}/${checksums}"
grep " ${archive}$" "${checksums}" | sha256sum -c -
tar -xzf "${archive}"
sudo install -m 0755 gitleaks /usr/local/bin/gitleaks
- run: gitleaks git --no-banner --redact .
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
node_modules/
coverage/
*.tmp
66 changes: 66 additions & 0 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
# Architecture and boundaries

```text
change + normalized evidence + catalog + policy
|
v
deterministic verification planner
|
v
plan: selected + skipped + argument + state
|
execution belongs to an external system
```

## Core and adapters

The core owns validation, reverse-impact closure, deterministic policy matching, fail-closed escalation, selection/skip arguments, canonical hashing, Visible Value planning counts, and shadow-result classification.

Adapters may translate Git diffs, package/project graphs, imports, coverage, test selectors, CODEOWNERS, schemas, or declared critical boundaries into normalized input. V1 implements no production adapter. This prevents the core from embedding Nx, Turbo, Jest, Vitest, testmon, or any one repository layout.

The verification catalog is the planner's universe. It is not a scheduler. Commands are opaque identities; the planner never shells out to them.

## Independence from Gearbox

Affected Verification owns the question: “What is the minimum defensible verification workload for this change under current evidence and policy?” Gearbox may ask that question and choose where or how to execute the returned work. Gearbox does not own the theory, catalog, or planner.

The prototype has no Gearbox dependency and is independently usable by people, coding agents, CI, PR bots, Opsle Tasks, or other tooling.

## Context Firewall boundary

Affected Verification decides **what should execute**. After execution, Context Firewall decides **what result should enter model context**. For example, the planner may require 17 of 1,000 tests; after those run, Context Firewall may retain only aggregate success plus one failure. Neither mechanism substitutes for the other.

## Decision Evidence and Trajectory boundaries

Decision Evidence may later validate change identity, evidence provenance, plan identity, selected/skip arguments, and execution-result bindings. Agent Trajectory Profiler may measure planned selections, actual executions, shadow comparisons, and observed latency/cost when genuinely recorded. V1 records compatible identities but creates no cross-repository package coupling.

## Shadow observations

`classifyShadow` first recomputes and verifies the plan identity, then compares a predicted plan with an externally supplied full-run record. Its durable observation binds:

- plan, change, policy, and planner schema identities;
- predicted selected and skipped check IDs;
- full-run executed IDs and failures;
- whether full execution was complete;
- relevant failures in skipped checks and their exact supplied reason;
- `SELECTION_MISS`, `NO_SELECTION_MISS`, or `INDETERMINATE_FULL_RUN_INCOMPLETE`.

A selection miss occurs only when a check omitted by the targeted plan fails in full execution and that failure is classified as relevant to the change. Relevance is an external oracle claim in v1; the classifier records it and does not infer causality.

No telemetry service exists. The fixture-level classifier proves the observation contract only.

## Trust ramp

`OBSERVE`
: Generate plans for inspection; existing verification remains authoritative.

`SHADOW`
: Generate targeted plans while full verification runs; retain complete, identity-bound miss observations.

`TRUSTED_BOUNDED`
: Permit targeted plans only for repository-specific low-risk classes whose impact/catalog completeness, policy behavior, miss history, oracle quality, and rollback path are defended.

`TRUSTED_POLICY`
: Extend authority to additional explicitly bounded classes only after their evidence is comparable and their promotion criteria are met.

Promotion depends on evidence quality and coverage, not an arbitrary run count. Critical/global classes may permanently require full verification. A miss, evidence drift, adapter change, catalog drift, policy change, or loss of shadow completeness can demote trust.
64 changes: 64 additions & 0 deletions BENCHMARK.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# Benchmark and research plan

No benchmark has run and no result is claimed.

## Central question

Can Affected Verification substantially reduce verification workload without increasing undetected regression risk relative to existing full verification?

Correctness and selection-miss evidence are evaluated before computation reduction.

## Arms

For each frozen repository/change fixture where applicable:

1. full verification;
2. native related/affected tooling;
3. Affected Verification using only normalized repository evidence;
4. Affected Verification consuming native tooling as evidence.

Native arms should include the repository's actual Jest, Vitest, Nx, Turbo, testmon, Bazel, Pants, or other selector rather than a synthetic strawman.

## Required freeze before a run

- exact repository and base/target revisions;
- immutable change set and changed-region derivation;
- complete verification catalog and units;
- evidence-provider versions and outputs;
- policy identity and trust stage;
- commands, environment, and deterministic plan outputs;
- full-verification correctness oracle;
- relevance adjudication protocol for full-run failures;
- failure classes, stopping rules, and excluded changes.

## Primary safety metrics

- relevant selection misses found by full/shadow execution;
- incomplete or indeterminate full runs;
- targeted plans that escalated because sufficiency was indefensible;
- false targeted-sufficiency claims discovered by audit;
- correctness-oracle disagreements.

Any relevant selection miss is reported before reduction. “No miss observed” is bounded to the executed corpus and is not proof of safety.

## Workload metrics

- checks available, selected, skipped, and actually executed, partitioned by check type;
- test executions available, selected, skipped, and actually executed;
- computation, latency, and cost only when directly observed in comparable arms;
- uncertainty/escalation frequency by change class;
- marginal reduction from each evidence source by controlled ablation.

Unlike verification types are not added into one deceptive percentage. Counterfactual avoided computation requires controlled comparability; plan counts alone are `EXACT`, not causal savings.

## Secondary questions

- Which evidence sources materially reduce workload?
- How often does uncertainty force escalation?
- Which change classes remain defensibly targetable?
- How frequently do native and composed selectors miss relevant full-run failures?
- Does adding policy and non-test checks produce value above native selectors?

## Initial experiment sequence

First, freeze a real public repository with a trustworthy full-verification baseline and run in `OBSERVE`/`SHADOW`. Compare plan identities and full outcomes without replacing CI. Only after the oracle, catalog completeness, adapter fidelity, and miss classifications are independently reviewable should a bounded trust decision be considered.
3 changes: 3 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
# Contributing

Keep the core deterministic, dependency-light, and separate from execution. Changes to selection semantics require exact tests, updated contract text, and new or revised conformance fixtures. Prior-art claims require primary sources. Do not add benchmark results without immutable inputs, a full-verification oracle, and explicit failure reporting.
14 changes: 14 additions & 0 deletions LIMITATIONS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
# Limitations and non-goals

- The prototype consumes synthetic normalized evidence; it has no Git, graph, coverage, ownership, schema, Nx, Turbo, Jest, Vitest, or testmon adapter.
- The component graph and catalog completeness flags are caller claims, not independently attested.
- The selection algorithm uses declared component scope and policy tags. It has no symbol/data-flow analysis, runtime coverage collection, weighted set cover, probabilistic model, or learned judgment.
- Commands are opaque identities. The project does not execute, schedule, cache, distribute, retry, or report CI work.
- `SUFFICIENT_*` means sufficient under the supplied model and policy, not globally safe, formally sound, or mathematically minimal.
- Full catalog selection cannot compensate for an incomplete catalog; that state remains `INSUFFICIENT_EVIDENCE`.
- Shadow relevance is caller supplied. A robust real benchmark needs an independent, reproducible relevance oracle.
- Test execution counts are declared catalog metadata. They are exact relative to input, not observed executions.
- No real-project benchmark, comparative baseline, selection-miss corpus, computation measurement, or independent reproduction exists.
- No production trust stage is justified. This repository is at most `PROTOTYPED`.
- No time, token, monetary, correctness, failure-prevention, or causal savings claim is supported.
- Context Firewall, Decision Evidence, Agent Trajectory Profiler, Gearbox, and Opsle Tasks are external consumers or validators, not dependencies.
27 changes: 27 additions & 0 deletions PRIOR_ART.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# Prior-art reconciliation

This audit uses project-owned documentation or source and established research literature. It is a boundary analysis, not a novelty claim. Several existing systems already provide sophisticated affected selection, conservative fallbacks, explanations, and shadow prediction. Affected Verification should reuse them as evidence providers where they are authoritative.

## Comparison

| System | Input evidence | Selection unit | Static/runtime | Uncertainty behavior | Non-test verification | Skip explanation | Shadow validation | Sufficiency claim | Intended reuse |
|---|---|---|---|---|---|---|---|---|---|
| [Nx affected](https://nx.dev/docs/features/ci-features/affected) | Git base/head or supplied files, project graph, source/config analysis, lockfile analysis | Projects, then requested targets/tasks | Primarily static workspace and project/task graph | Conservative lockfile default marks all projects; graph/plugin behavior can broaden | Yes: any Nx target such as lint, test, build, or custom task | Affected graph and task graph explain inclusion; no first-class reason for every omitted task | No general OSS CLI shadow/full miss contract found | Claims a minimum affected project set, not acceptance sufficiency across verification classes | Consume affected projects/task graph and its failsafe signals |
| [Turborepo `--affected`](https://github.com/vercel/turborepo/blob/main/apps/docs/content/docs/reference/configuration.mdx) | Git range, package graph, global dependencies, optionally task input globs | Packages by default; tasks with `affectedUsingTaskInputs` | Static package/task/input graph | Global configuration and lockfile changes select all; missing Git history can fall back broadly | Yes: caller-named Turbo tasks | `--dry=json` shows planned tasks, but not an argument for every skipped task | No built-in miss classifier found | No cross-class acceptance sufficiency claim | Consume affected tasks/packages and global-change signals |
| [Jest `--findRelatedTests`](https://jestjs.io/docs/30.0/cli#--findrelatedtests-spaceseparatedlistofsourcefiles) | Supplied source files and Jest's module dependency information | Test files/tests | Static module-resolution evidence, not runtime coverage | No general external uncertainty/policy model; caller owns source list and configuration | No; tests only | Lists selected tests, not a first-class reason for each omitted test | No | “Related” test selection, not acceptance sufficiency | Use its related-test result as one test-evidence provider |
| [Vitest `related` / `--changed`](https://vitest.dev/guide/cli) | Supplied files or Git changes plus static imports | Test files | Static imports; documented dynamic-import limitation | `forceRerunTriggers` and config/package changes can force the full suite | No; tests only | No durable per-skip argument in the core CLI | No general miss observation contract | No acceptance sufficiency claim | Use related/changed output and propagate its limitations |
| [pytest-testmon](https://www.testmon.org/) | Per-test executed-code dependencies from Coverage.py plus source/block changes and persisted `.testmondata` | Pytest tests | Runtime coverage plus source analysis | First qualifying run executes all; failed tests rerun; mode conflicts can disable selection/collection | No; tests only | Selection can be inspected, but no cross-catalog reason for every skip | `--testmon-noselect` runs all while prioritizing likely failures, but no durable generic miss contract | Makes scoped affected-test claims, not whole-change verification sufficiency | Use runtime test-to-code evidence and explicit database readiness state |
| [vitest-affected](https://github.com/craigvandotcom/vitest-affected) | Git changes, cached Vitest runtime import data, delta static parsing, explicit full-suite triggers | Vitest test files | Runtime-observed imports plus static delta parsing | Cache/git/graph failure falls back to full; first run is full; documented non-import gaps need triggers; stale cache warns but does not force full | No; tests only | Yes: selected chains and a why-not explanation based on absence from the cached graph | Yes: predicts selection while the full suite runs and emits decision data | Explicitly advises retaining full/periodic truth; no cross-class sufficiency claim | Reuse runtime graph, explain trails, fallback signals, and shadow observations rather than rebuilding them |
| [Bazel query](https://bazel.build/query/quickstart) | Declared target/build graph and query universe | Build/test targets | Static declared build graph | Query scope and graph completeness are caller responsibilities; `--keep_going` tolerates errors but is not sufficiency | Queries can select any target; Bazel can build/test selected targets | `rdeps`, `somepath`, and `allpaths` can explain graph relationships | No generic selection-miss contract | No change-acceptance sufficiency claim | Consume target/reverse-dependency paths and configuration boundaries |
| [Pants changed targets](https://www.pantsbuild.org/stable/docs/using-pants/advanced-target-selection) | Git changes, inferred/declared target dependencies, `changed` options | Targets supplied to any goal | Static target/dependency graph | Can include direct or transitive dependees; no general evidence-sufficiency state | Yes: selected targets can feed test, lint, package, and other goals | Target introspection can explain graph membership; no per-skip verification argument | No generic miss classifier | No whole-change sufficiency claim | Consume changed target closure and goal/task metadata |
| [Azure Pipelines Test Impact Analysis](https://learn.microsoft.com/en-us/azure/devops/pipelines/test/test-impact-analysis) | Managed-code test impact data and source changes | Automated tests | Runtime instrumentation/impact data | Unknown file types fall back to all tests; periodic full runs are configurable and recommended | No; tests only | Reporting exposes TIA outcome, not a generic per-skip argument | Periodic full runs provide validation opportunity, but no portable durable miss schema | Scoped platform TIA, not cross-class acceptance sufficiency | Reuse impacted-test result, unknown-type fallback, and full-run cadence evidence |

## Research boundary

Regression-test selection is a mature research field. Yoo and Harman's [survey](https://doi.org/10.1002/stvr.430) distinguishes minimization, modification-aware selection, and prioritization. Rothermel and Harrold's [safe regression-test selection work](https://doi.org/10.1145/248233.248262) uses “safe” in a formal, controlled-program sense. This project must not borrow that term for a multi-language, multi-check planner without equivalent proof.

The prototype's defensible distinction is narrower: it combines potentially heterogeneous selector evidence with an explicit verification catalog and risk policy, then emits an acceptance-oriented argument covering tests and non-test checks, uncertainty, escalation, and every catalogued skip. The individual graph, coverage, affected-target, explanation, and shadow mechanisms are prior art.

## Novelty ceiling

No novelty is claimed. Future comparison must include full verification, native affected tooling, Affected Verification, and Affected Verification consuming native tooling. Value exists only if the composed argument improves workload without increasing relevant misses under controlled shadow evidence.
39 changes: 37 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,2 +1,37 @@
# affected-verification
Minimum defensible verification planning from change-impact, coverage, policy, and risk evidence
# Affected Verification

Affected Verification deterministically selects the smallest verification workload whose sufficiency can be defended from the available change-impact, dependency, coverage, policy, and risk evidence.

The operative claim is **minimum defensible verification**, not mathematical global minimality. Unknown impact is never permission to skip work.

This repository contains a dependency-free Node.js 20 prototype. It consumes normalized change, impact, verification-catalog, and policy data and emits an `opsle.affected-verification.plan.v1` argument containing selected checks, skipped checks, exact reasons, provenance hashes, uncertainty, escalation, and sufficiency. It plans work; it does not run CI.

## Try it

```bash
node bin/affected-verification.js \
fixture unrelated-large-suite \
--receipt /tmp/av-receipt.json
```

Canonical plan JSON is written to stdout. The `opsle.value-receipt.v1` is written only to the requested sidecar, and one `[Affected Verification]` indicator is written to stderr. The fixture reports exactly 14 of 1,043 test executions selected, 1,029 skipped, plus one lint and one typecheck; the test-execution reduction is an `EXACT` calculation, not a time, cost, token, or correctness claim.

```bash
npm run verify
```

## Contract and evidence

- [SPEC.md](SPEC.md) — normative prototype contract and sufficiency states
- [PRIOR_ART.md](PRIOR_ART.md) — source-linked reconciliation with existing selectors
- [ARCHITECTURE.md](ARCHITECTURE.md) — adapters, project boundaries, shadow mode, and trust ramp
- [BENCHMARK.md](BENCHMARK.md) — controlled research plan; no results are claimed
- [LIMITATIONS.md](LIMITATIONS.md) — current claim ceiling and non-goals
- [fixtures/scenarios.js](fixtures/scenarios.js) and [fixtures/negative-cases.js](fixtures/negative-cases.js) — twelve positive/boundary scenarios plus explicit conflicting, malformed, impossible, and tampered cases
- [schemas/plan-v1.schema.json](schemas/plan-v1.schema.json) — plan shape

## Status

This is a narrow research prototype, not a trusted replacement for full verification. It has not been benchmarked on a real repository, and it has no production adapters.

Apache-2.0.
Loading
Loading