Skip to content

Commit aa0c87b

Browse files
committed
docs: clarify eval and experiment boundaries
Entire-Checkpoint: 335841c2d678
1 parent e2f865a commit aa0c87b

2 files changed

Lines changed: 104 additions & 11 deletions

File tree

‎docs/adr/2026-06-17-harbor-runner-boundary.md‎

Lines changed: 39 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -37,19 +37,46 @@ Harbor should own:
3737
- Harbor `task.toml` files and Harbor YAML config;
3838
- Opik trace upload through Harbor when enabled.
3939

40+
## Alignment with experiment separation
41+
42+
The 2026-06-23 experiment/eval separation decision makes runtime binding an
43+
experiment concern. Harbor execution should follow the same split:
44+
45+
- AgentV eval YAML remains the authoring or selection surface for what benchmark
46+
suite is being evaluated.
47+
- AgentV experiment YAML selects the Harbor runner, candidate agent/model, run
48+
policy, and other runtime binding.
49+
- Harbor-authored YAML remains Harbor's own config surface when the standard
50+
suite needs Harbor-specific task packaging or verifier settings.
51+
52+
This means the examples below describe the desired logical fields, but new
53+
runtime fields should be placed on an experiment unless they are genuinely part
54+
of the benchmark suite identity. Do not put candidate agent/model binding in the
55+
eval file for new AgentV-native examples.
56+
4057
## Minimal future config surface
4158

42-
The AgentV eval file should select Harbor with a nested runner config:
59+
An AgentV eval suite can select the benchmark identity without copying Harbor's
60+
task schema:
4361

4462
```yaml
45-
name: swebench-verified-codex
63+
name: swebench-verified
4664

4765
execution:
4866
runner: harbor
4967
harbor:
5068
dataset: swebench-verified
51-
agent: codex
52-
model: openai/gpt-5-mini
69+
```
70+
71+
The corresponding experiment selects how that suite runs:
72+
73+
```yaml
74+
name: swebench-verified-codex
75+
target: codex-gpt5-mini
76+
evals: evals/swebench-verified.eval.yaml
77+
runner:
78+
type: harbor
79+
harbor:
5380
opik:
5481
enabled: true
5582
```
@@ -74,6 +101,11 @@ or Docker/Compose adapter settings. If a Harbor option becomes too specific to
74101
standardize, users should put it in the referenced Harbor YAML file instead of
75102
AgentV adding a pass-through field.
76103

104+
If the Harbor integration later chooses to move `runner.type: harbor` entirely
105+
into experiment files, this ADR should be updated with the final schema. The
106+
boundary decision is stable: Harbor runtime binding is not an eval-case schema
107+
extension.
108+
77109
## CLI invocation strategy
78110

79111
Native evals continue to run with the existing command:
@@ -105,9 +137,9 @@ agentv results import harbor --job <harbor-job-id>
105137
```
106138

107139
Do not overload native `--target` semantics in the first Harbor runner slice.
108-
Harbor `agent`, `model`, and matrix behavior should come from
109-
`execution.harbor` or the referenced Harbor YAML until repeated usage proves a
110-
shared AgentV flag is needed.
140+
Harbor `agent`, `model`, and matrix behavior should come from the experiment or
141+
the referenced Harbor YAML until repeated usage proves a shared AgentV flag is
142+
needed.
111143

112144
## Unsupported fields and non-goals
113145

‎docs/adr/2026-06-23-experiments-vs-eval-separation.md‎

Lines changed: 65 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -37,10 +37,22 @@ This decision must also preserve AgentV's existing product boundary:
3737

3838
## Vocabulary
3939

40-
An eval is a frozen task definition. It includes the prompt or dataset, expected
41-
behavior, task-owned workspace fixtures, and assertions. AgentV's LLM-judge,
42-
code-grader, deterministic assertions, and hidden or explicit evaluation
43-
criteria belong here.
40+
An eval suite is a frozen task-definition boundary. It includes suite metadata,
41+
shared prompt/context, case references, shared assertions or graders, and
42+
task-owned workspace fixtures. AgentV's LLM-judge, code-grader, deterministic
43+
assertions, and hidden or explicit evaluation criteria belong here.
44+
45+
An eval case is one atomic task inside a suite. It includes the case id, prompt
46+
or input, criteria, expected output or reference behavior, case metadata, and
47+
case-specific workspace overrides. A suite can inline cases, point to
48+
`cases.yaml`/JSONL, or use a directory convention where each case owns files such
49+
as `TASK.txt`, `PROMPT.md`, `answer/`, or `grader.test.ts`.
50+
51+
In that directory-convention form, `EVAL.yaml` may be thin or inferred by a
52+
loader, but the suite layer is still present conceptually: the directory
53+
convention plus runner adapter is the suite contract. This distinction matters
54+
because AgentV is a reusable framework, not a single benchmark harness whose
55+
suite semantics can live only in code.
4456

4557
An experiment is a committed or generated run definition. It declares which
4658
agent, target, provider, model, harness options, setup steps, run count, timeout,
@@ -62,6 +74,13 @@ task inputs, datasets, assertions, and task fixtures. They should not be the
6274
canonical place for which agent, model, harness, setup injection, sandbox, or run
6375
matrix executes the task.
6476

77+
For simple projects, an eval-only run remains valid. AgentV treats the implicit
78+
experiment label as `default` unless a committed experiment is configured. For
79+
specialized harnesses that already have a strong directory contract, AgentV may
80+
support loaders that infer the suite from the directory instead of requiring a
81+
separate YAML file, but those loaders must still lower into the same suite/case
82+
concepts.
83+
6584
Experiment files will live under `experiments/` by convention. AgentV will
6685
support YAML as the canonical authoring path for the abstraction story and TypeScript
6786
as the power-user escape hatch:
@@ -85,6 +104,48 @@ setup:
85104
- script: cp skills/copilot/AGENTS.md AGENTS.md
86105
```
87106
107+
## Workspace boundary
108+
109+
Workspace config belongs with the eval suite or case when it defines the task
110+
scenario being replayed. Examples:
111+
112+
- clone `org/repo` at a specific `commit` or `base_commit`;
113+
- copy starter files, failing tests, fixtures, or issue prompts;
114+
- run task-owned setup hooks that prepare the repo state required by the case;
115+
- declare per-case repo pins or fixture overrides.
116+
117+
Experiment setup belongs with the experiment when it changes the runtime
118+
condition being compared. Examples:
119+
120+
- choose `codex` versus `claude` targets;
121+
- inject an `AGENTS.md`, skill, guideline file, or tool config for an A/B run;
122+
- choose repeat/run policy, timeout, workers, budget, or sandbox mode;
123+
- select a subset of suites or cases for a run campaign.
124+
125+
Rule of thumb: if changing it changes the task being evaluated, put it in the
126+
suite or case workspace. If changing it changes the candidate or run condition
127+
measured against the same task, put it in the experiment.
128+
129+
## Directory-style evals
130+
131+
Convex-style harnesses are a useful counterexample to requiring YAML for every
132+
case. A product-specific benchmark can encode each case as a directory with a
133+
task prompt, reference solution, and executable grader. In AgentV terms, that is
134+
not `experiment -> eval case` with no suite; it is an implicit suite contract
135+
provided by the loader:
136+
137+
```text
138+
evals/<category>/<case>/
139+
TASK.txt # case input
140+
answer/ # reference fixture
141+
grader.test.ts # code-grader assertion
142+
```
143+
144+
AgentV should support this as an import/loader shape when useful, but the core
145+
contract remains `experiment -> eval suite -> eval case`. The experiment applies
146+
runtime bindings to the selected suites/cases; it does not own the prompt,
147+
expected behavior, or grading contract.
148+
88149
`config.yaml` will gain a default experiment pointer so existing `agentv eval`
89150
usage keeps working:
90151

0 commit comments

Comments
 (0)