Skip to content

Latest commit

 

History

History
156 lines (125 loc) · 6.91 KB

File metadata and controls

156 lines (125 loc) · 6.91 KB

UsageBench artifact review guide

Reproducibility claim

A released UsageBench report produced by reference environment version 1 can be rebuilt and rerun from the report alone. The canonical platform is linux/amd64; version 1 covers Bifrost and gopls. The reproduction command selects the exact UsageBench release and analyzer input, builds the matching local image, runs the analyzer without network access, and compares the new report semantically with the published report.

This repository publishes a versioned build contract, not ready-built images. CI builds and exercises both images ephemerally, but does not log in to a registry, push an image, export an OCI archive, or upload an image artifact. Built images may be archived with a future Zenodo deposit after the artifact has been reviewed.

The checked-in fixture corpus remains a development and diagnosis corpus. The separate real-project-v1 partition is preregistered, independently reviewed, and published as evaluation release v0.2.0. Reproducible execution does not change the ground-truth status recorded by either partition.

The wider public comparison uses one pinned native report when a canonical image is not available. The freeze validates the selected release, profile checksum, resolved analyzer identity, executable provenance, environment, case coverage, and report checksum. Native rows are not presented as a cross-host reproducibility or cryptographic software-supply-chain claim.

Requirements

  • Docker Engine or Docker Desktop with docker buildx and amd64 container support;
  • Bash, Git, and jq;
  • outbound network access while building images; and
  • enough free disk for Rust, Go, Bifrost, and gopls build layers.

Native profiles additionally need their exact executables and package launchers available, plus permission for language servers such as Roslyn to create local project-build sockets. See docs/src/content/docs/reproduce.md and adapters/lsp/README.md for profile and hydration requirements.

Runtime analysis is network-disabled. On an Apple Silicon development machine, the clean amd64 Bifrost analyzer build took about nine minutes under emulation; the first gopls image build took about five minutes. Native amd64 systems and cached rebuilds can be substantially faster.

Reproduce a published report

From an extracted usagebench-vMAJOR.MINOR.PATCH.tar.gz release bundle, run:

./scripts/reproduce-report.sh /path/to/published-report.json reproduced-report.json

The command verifies that the input is a canonical container report, then:

  1. reads its UsageBench release, exact revision, environment version, runner, case selection, and inclusion policy;
  2. uses the current bundle when it matches, or fetches the exact tag, verifies its commit, and exports it into a release-shaped corpus;
  3. builds the local linux/amd64 image from digest-pinned definitions;
  4. reruns the same selection with networking disabled, as a non-root user, with the release corpus mounted read-only; and
  5. runs compare-reports inside the image.

Success ends with:

reports are semantically equivalent
reproduced report: .../reproduced-report.json

The comparator ignores only timestamps, temporary workspace roots, local filesystem paths, and the locally rebuilt image identity. It still compares the release and revision, reference-environment definition digest, executable checksum, requested and resolved analyzer versions, capabilities, case statuses, locations, diagnostics, and totals. A changed outcome exits nonzero and identifies the case-level field that differs.

Build and inspect the environments directly

Use the release tag recorded in the report:

./scripts/reference-image.sh bifrost vMAJOR.MINOR.PATCH
./scripts/reference-image.sh gopls vMAJOR.MINOR.PATCH

Build metadata is written under target/reference/. It contains the canonical platform, stable definition and full identity digests, analyzer identity, source release and revision, loaded image ID, immutable registry digest when available, and BuildKit's output digest. The script restores by immutable digest only after verifying every identity label. Set USAGEBENCH_REFERENCE_IMAGE_FORCE_REBUILD=1 to bypass reuse; forced rebuilds are deliberately incompatible with publication.

To run one released case directly:

./scripts/run-reference.sh \
  gopls \
  /path/to/extracted/usagebench-vMAJOR.MINOR.PATCH \
  gopls-report.json \
  benchmarks/cases/go-baseline.yaml \
  go-package-function-call

The corpus root must contain .usagebench-release.json. The wrapper verifies that its release and revision match the image labels and embedded identity, then enforces --network none, a read-only root filesystem, a non-root UID/GID, a read-only corpus mount, and isolated writable tmpfs work directories. A fresh private directory receives the container output; only the completed report is copied to the requested host path.

Inspect the evidence envelope with:

jq '{usagebenchRelease, usagebenchRevision, runner, invocation, environment}' \
  reproduced-report.json

Version and integrity boundaries

  • usagebenchRelease and usagebenchRevision identify the immutable corpus and harness source.
  • environment.referenceEnvironment.version identifies the container contract independently of the benchmark and YAML schema versions.
  • definitionDigest covers the environment manifest, schema, Dockerfile, and build/run wrappers.
  • Dockerfile frontend, builder, and runtime images are pinned by amd64 digest.
  • Bifrost is fetched at an exact Git commit. gopls is fetched at an exact module version and its Go module checksum is verified before compilation.
  • Cargo lockfile checksums and the pinned Go module graph protect transitive build inputs.
  • Reports checksum the analyzer executable that actually ran.
  • The wrapper verifies the full identity envelope, resolves registry tags to an immutable digest, and executes the verified immutable local image ID.

Repeated cached builds during development produced identical image IDs for both version 1 images. Cross-builder byte identity is not the scientific claim: the semantic comparator intentionally permits the local image identity to differ while requiring the stable definition digest, executable checksum, and result semantics to match.

Troubleshooting

  • Cannot connect to the Docker daemon: start Docker Engine or Docker Desktop.
  • amd64 warnings or very slow builds on arm64: ensure Docker's amd64 emulation is enabled; the canonical platform is intentionally fixed.
  • build download failures: construction needs network access even though the benchmark runtime does not.
  • reference definition mismatch: use the release bundle named by the report; do not mix container scripts from another revision.
  • a semantic diff after a successful run: preserve both JSON reports. The diff is evidence of a result, capability, executable, or environment discrepancy, not a reason to relax the benchmark assertion.