A released UsageBench report produced by reference environment version 1 can
be rebuilt and rerun from the report alone. The canonical platform is
linux/amd64; version 1 covers Bifrost and gopls. The reproduction command
selects the exact UsageBench release and analyzer input, builds the matching
local image, runs the analyzer without network access, and compares the new
report semantically with the published report.
This repository publishes a versioned build contract, not ready-built images. CI builds and exercises both images ephemerally, but does not log in to a registry, push an image, export an OCI archive, or upload an image artifact. Built images may be archived with a future Zenodo deposit after the artifact has been reviewed.
The checked-in fixture corpus remains a development and diagnosis corpus. The
separate real-project-v1 partition is preregistered, independently reviewed,
and published as evaluation release v0.2.0. Reproducible execution does not
change the ground-truth status recorded by either partition.
The wider public comparison uses one pinned native report when a canonical image is not available. The freeze validates the selected release, profile checksum, resolved analyzer identity, executable provenance, environment, case coverage, and report checksum. Native rows are not presented as a cross-host reproducibility or cryptographic software-supply-chain claim.
- Docker Engine or Docker Desktop with
docker buildxand amd64 container support; - Bash, Git, and jq;
- outbound network access while building images; and
- enough free disk for Rust, Go, Bifrost, and gopls build layers.
Native profiles additionally need their exact executables and package launchers
available, plus permission for language servers such as Roslyn to create local
project-build sockets. See docs/src/content/docs/reproduce.md and
adapters/lsp/README.md for profile and hydration requirements.
Runtime analysis is network-disabled. On an Apple Silicon development machine, the clean amd64 Bifrost analyzer build took about nine minutes under emulation; the first gopls image build took about five minutes. Native amd64 systems and cached rebuilds can be substantially faster.
From an extracted usagebench-vMAJOR.MINOR.PATCH.tar.gz release bundle, run:
./scripts/reproduce-report.sh /path/to/published-report.json reproduced-report.jsonThe command verifies that the input is a canonical container report, then:
- reads its UsageBench release, exact revision, environment version, runner, case selection, and inclusion policy;
- uses the current bundle when it matches, or fetches the exact tag, verifies its commit, and exports it into a release-shaped corpus;
- builds the local
linux/amd64image from digest-pinned definitions; - reruns the same selection with networking disabled, as a non-root user, with the release corpus mounted read-only; and
- runs
compare-reportsinside the image.
Success ends with:
reports are semantically equivalent
reproduced report: .../reproduced-report.json
The comparator ignores only timestamps, temporary workspace roots, local filesystem paths, and the locally rebuilt image identity. It still compares the release and revision, reference-environment definition digest, executable checksum, requested and resolved analyzer versions, capabilities, case statuses, locations, diagnostics, and totals. A changed outcome exits nonzero and identifies the case-level field that differs.
Use the release tag recorded in the report:
./scripts/reference-image.sh bifrost vMAJOR.MINOR.PATCH
./scripts/reference-image.sh gopls vMAJOR.MINOR.PATCHBuild metadata is written under target/reference/. It contains the canonical
platform, stable definition and full identity digests, analyzer identity,
source release and revision, loaded image ID, immutable registry digest when
available, and BuildKit's output digest. The script restores by immutable
digest only after verifying every identity label. Set
USAGEBENCH_REFERENCE_IMAGE_FORCE_REBUILD=1 to bypass reuse; forced rebuilds
are deliberately incompatible with publication.
To run one released case directly:
./scripts/run-reference.sh \
gopls \
/path/to/extracted/usagebench-vMAJOR.MINOR.PATCH \
gopls-report.json \
benchmarks/cases/go-baseline.yaml \
go-package-function-callThe corpus root must contain .usagebench-release.json. The wrapper verifies
that its release and revision match the image labels and embedded identity,
then enforces --network none, a read-only root filesystem, a non-root UID/GID,
a read-only corpus mount, and isolated writable tmpfs work directories. A fresh
private directory receives the container output; only the completed report is
copied to the requested host path.
Inspect the evidence envelope with:
jq '{usagebenchRelease, usagebenchRevision, runner, invocation, environment}' \
reproduced-report.jsonusagebenchReleaseandusagebenchRevisionidentify the immutable corpus and harness source.environment.referenceEnvironment.versionidentifies the container contract independently of the benchmark and YAML schema versions.definitionDigestcovers the environment manifest, schema, Dockerfile, and build/run wrappers.- Dockerfile frontend, builder, and runtime images are pinned by amd64 digest.
- Bifrost is fetched at an exact Git commit. gopls is fetched at an exact module version and its Go module checksum is verified before compilation.
- Cargo lockfile checksums and the pinned Go module graph protect transitive build inputs.
- Reports checksum the analyzer executable that actually ran.
- The wrapper verifies the full identity envelope, resolves registry tags to an immutable digest, and executes the verified immutable local image ID.
Repeated cached builds during development produced identical image IDs for both version 1 images. Cross-builder byte identity is not the scientific claim: the semantic comparator intentionally permits the local image identity to differ while requiring the stable definition digest, executable checksum, and result semantics to match.
Cannot connect to the Docker daemon: start Docker Engine or Docker Desktop.- amd64 warnings or very slow builds on arm64: ensure Docker's amd64 emulation is enabled; the canonical platform is intentionally fixed.
- build download failures: construction needs network access even though the benchmark runtime does not.
reference definition mismatch: use the release bundle named by the report; do not mix container scripts from another revision.- a semantic diff after a successful run: preserve both JSON reports. The diff is evidence of a result, capability, executable, or environment discrepancy, not a reason to relax the benchmark assertion.