Skip to content

Commit d69db58

Browse files
christsoclaude
andcommitted
docs(adr-0017): cross-check exploitbench — confirms split-bundle/no-DB; adopt provenance field; audit/anti-reward-hacking future scope
exploitbench confirms: split filesystem run-tree = source of truth, SQLite is a derived rebuildable view (import/export bijection) not required, image pinned by sha256 digest, config_snapshot=bundle.json. Borrow: (1) provenance field on result rows (native/mock/replay/imported_*) — adopt; (2) eval-integrity/anti-reward-hacking (read-only grader container, audit re-grade + red-flag scan + model-identity check) — future scope. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1 parent c5dc6f7 commit d69db58

1 file changed

Lines changed: 3 additions & 0 deletions

File tree

‎docs/adr/0017-output-artifact-and-workspace-resolver-contract.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -131,6 +131,9 @@ the repo's tests in the workspace and its exit code is the verdict (exactly marg
131131
(#5), this is how AgentV runs SWE-bench natively — same `repo`+`commit` provenance, no
132132
schema change, no new grader type.
133133

134+
### Cross-check: exploitbench (confirms + two borrowables)
135+
exploitbench (security-exploit benchmark; AgentV research `entities/exploitbench.md`) **confirms** this contract: split filesystem run-tree is the source of truth (`job.json`/`score.json`/`cost.json`/`transcript.jsonl`/`tool_calls.jsonl`/`config_snapshot.yaml`); its SQLite is a **derived, rebuildable view** (`import`/`export` bijection), not required — validating our no-DB core (jq + `index.jsonl` is the query surface; a SQLite view stays an optional post-run adapter, Phoenix boundary intact). Docker images are pinned by `sha256:` digest at run start (reinforces resolver backend #5); `config_snapshot` = our `bundle.json`. **Borrow:** (1) a **`provenance`** field on result rows (`native`/`mock`/`replay`/`imported_from_*`) — durable, fits AgentV's replay/transcript/mock providers; adopt now. (2) **Eval-integrity / anti-reward-hacking — future scope**: run high-stakes graders in a fresh container with the workspace mounted **read-only**; an `audit` pass that re-grades from the stored transcript, scans for reward-hacking red flags, and verifies model identity (the provider served the requested model). "Post-hoc audit as part of benchmark validity."
136+
134137
## Consequences
135138

136139
- Refines ADR-0011/0012 (bundle layout, `index_path`, timing→metrics merge, `.internal/`);

0 commit comments

Comments
 (0)