Skip to content

Commit 2fcd72f

Browse files
authored
docs: clarify eval and experiment boundaries (#1517)
Entire-Checkpoint: 335841c2d678
1 parent 462a03e commit 2fcd72f

2 files changed

Lines changed: 159 additions & 28 deletions

File tree

‎docs/adr/2026-06-17-harbor-runner-boundary.md‎

Lines changed: 94 additions & 24 deletions
Original file line numberDiff line numberDiff line change
@@ -37,42 +37,110 @@ Harbor should own:
3737
- Harbor `task.toml` files and Harbor YAML config;
3838
- Opik trace upload through Harbor when enabled.
3939

40+
## Alignment with experiment separation
41+
42+
The 2026-06-23 experiment/eval separation decision makes runtime binding an
43+
experiment concern. Harbor execution should follow the same split:
44+
45+
- AgentV eval YAML remains the authoring or selection surface for what benchmark
46+
suite is being evaluated.
47+
- AgentV experiment YAML selects or pins the Harbor runner, candidate
48+
agent/model, run policy, and other runtime binding.
49+
- Harbor-authored YAML remains Harbor's own config surface when the standard
50+
suite needs Harbor-specific task packaging or verifier settings.
51+
52+
This means the examples below describe the desired logical fields, but new
53+
runtime fields should be placed on an experiment unless they are genuinely part
54+
of the benchmark suite identity. Do not put candidate agent/model binding in the
55+
eval file for new AgentV-native examples.
56+
4057
## Minimal future config surface
4158

42-
The AgentV eval file should select Harbor with a nested runner config:
59+
An AgentV eval suite can select the benchmark source without copying Harbor's
60+
task schema or claiming to be the runtime runner:
4361

4462
```yaml
45-
name: swebench-verified-codex
63+
name: swebench-verified
4664

47-
execution:
48-
runner: harbor
49-
harbor:
50-
dataset: swebench-verified
51-
agent: codex
52-
model: openai/gpt-5-mini
65+
source:
66+
type: harbor
67+
dataset: swebench-verified
68+
```
69+
70+
The corresponding experiment selects how that suite runs:
71+
72+
```yaml
73+
name: swebench-verified-codex
74+
target: codex-gpt5-mini
75+
evals: evals/swebench-verified.eval.yaml
76+
runner:
77+
type: harbor
78+
options:
5379
opik:
5480
enabled: true
5581
```
5682
5783
For a Harbor-authored YAML file, use `config` instead of `dataset`:
5884

5985
```yaml
60-
execution:
61-
runner: harbor
62-
harbor:
63-
config: ./harbor/swebench-verified.yaml
86+
source:
87+
type: harbor
88+
config: ./harbor/swebench-verified.yaml
6489
```
6590

6691
The first implementation should accept exactly one Harbor source selector:
6792
`dataset` for a known Harbor dataset id, or `config` for an existing Harbor YAML
6893
file. There should be no precedence rule between them. If both are set, fail
6994
validation and ask the user to choose one.
7095

71-
Keep Harbor-specific options nested under `execution.harbor`. Do not add
72-
top-level AgentV fields for Harbor task packaging, verifier images, task patches,
73-
or Docker/Compose adapter settings. If a Harbor option becomes too specific to
74-
standardize, users should put it in the referenced Harbor YAML file instead of
75-
AgentV adding a pass-through field.
96+
Do not combine Harbor suite selection with candidate binding in the eval file:
97+
98+
```yaml
99+
# Avoid in eval.yaml
100+
execution:
101+
runner: harbor
102+
harbor:
103+
dataset: swebench-verified
104+
agent: codex
105+
model: openai/gpt-5-mini
106+
```
107+
108+
Split that shape across the suite and experiment instead:
109+
110+
```yaml
111+
# evals/swebench-verified.eval.yaml
112+
name: swebench-verified
113+
source:
114+
type: harbor
115+
dataset: swebench-verified
116+
```
117+
118+
```yaml
119+
# experiments/swebench-verified-codex.yaml
120+
name: swebench-verified-codex
121+
target: codex
122+
model: openai/gpt-5-mini
123+
evals: evals/swebench-verified.eval.yaml
124+
runner:
125+
type: harbor
126+
```
127+
128+
Keep Harbor suite source selection under `source` in the eval suite. Keep
129+
experiment-side runner selection under `runner.type`, with runner knobs under
130+
`runner.options`. The eval suite answers "where do these cases come from?"; the
131+
experiment answers "how is this run executed?" Do not use `execution.runner` in
132+
new eval-suite examples because that name collides with the experiment runner.
133+
Do not repeat the runner discriminator as `runner.harbor.options`; `type:
134+
harbor` already provides that namespace.
135+
136+
Do not add top-level AgentV fields for Harbor task packaging, verifier images,
137+
task patches, or Docker/Compose adapter settings. If a Harbor option becomes too
138+
specific to standardize, users should put it in the referenced Harbor YAML file
139+
instead of AgentV adding a pass-through field.
140+
141+
If the Harbor integration later changes the eval source or experiment runner
142+
schema, this ADR should be updated with the final shape. The boundary decision
143+
is stable: Harbor runtime binding is not an eval-case schema extension.
76144

77145
## CLI invocation strategy
78146

@@ -82,8 +150,9 @@ Native evals continue to run with the existing command:
82150
agentv eval evals/native.eval.yaml --target codex
83151
```
84152

85-
Harbor-backed evals should use the same top-level entrypoint and dispatch based
86-
on `execution.runner`:
153+
Harbor-backed evals should use the same top-level entrypoint. If no explicit
154+
experiment runner is configured, AgentV may infer Harbor execution from
155+
`source.type: harbor`:
87156

88157
```bash
89158
agentv eval evals/swebench-harbor.eval.yaml
@@ -105,9 +174,9 @@ agentv results import harbor --job <harbor-job-id>
105174
```
106175

107176
Do not overload native `--target` semantics in the first Harbor runner slice.
108-
Harbor `agent`, `model`, and matrix behavior should come from
109-
`execution.harbor` or the referenced Harbor YAML until repeated usage proves a
110-
shared AgentV flag is needed.
177+
Harbor `agent`, `model`, and matrix behavior should come from the experiment or
178+
the referenced Harbor YAML until repeated usage proves a shared AgentV flag is
179+
needed.
111180

112181
## Unsupported fields and non-goals
113182

@@ -128,8 +197,9 @@ standard-suite path.
128197
## Implementation sequencing
129198

130199
1. Document the native-vs-Harbor boundary and commit alias rules.
131-
2. Add schema validation for optional `execution.runner` and
132-
`execution.harbor`, with no changes to native workspace acquisition.
200+
2. Add schema validation for eval-suite `source.type: harbor` and exactly one of
201+
`source.dataset` or `source.config`, plus experiment `runner.type` and
202+
`runner.options`, with no changes to native workspace acquisition.
133203
3. Add a Harbor launch adapter that records job identity and status.
134204
4. Add a Harbor result importer that maps rewards, exceptions, timings,
135205
artifacts, and Opik trace URLs into AgentV run bundles.

‎docs/adr/2026-06-23-experiments-vs-eval-separation.md‎

Lines changed: 65 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -37,10 +37,22 @@ This decision must also preserve AgentV's existing product boundary:
3737

3838
## Vocabulary
3939

40-
An eval is a frozen task definition. It includes the prompt or dataset, expected
41-
behavior, task-owned workspace fixtures, and assertions. AgentV's LLM-judge,
42-
code-grader, deterministic assertions, and hidden or explicit evaluation
43-
criteria belong here.
40+
An eval suite is a frozen task-definition boundary. It includes suite metadata,
41+
shared prompt/context, case references, shared assertions or graders, and
42+
task-owned workspace fixtures. AgentV's LLM-judge, code-grader, deterministic
43+
assertions, and hidden or explicit evaluation criteria belong here.
44+
45+
An eval case is one atomic task inside a suite. It includes the case id, prompt
46+
or input, criteria, expected output or reference behavior, case metadata, and
47+
case-specific workspace overrides. A suite can inline cases, point to
48+
`cases.yaml`/JSONL, or use a directory convention where each case owns files such
49+
as `TASK.txt`, `PROMPT.md`, `answer/`, or `grader.test.ts`.
50+
51+
In that directory-convention form, `EVAL.yaml` may be thin or inferred by a
52+
loader, but the suite layer is still present conceptually: the directory
53+
convention plus runner adapter is the suite contract. This distinction matters
54+
because AgentV is a reusable framework, not a single benchmark harness whose
55+
suite semantics can live only in code.
4456

4557
An experiment is a committed or generated run definition. It declares which
4658
agent, target, provider, model, harness options, setup steps, run count, timeout,
@@ -62,6 +74,13 @@ task inputs, datasets, assertions, and task fixtures. They should not be the
6274
canonical place for which agent, model, harness, setup injection, sandbox, or run
6375
matrix executes the task.
6476

77+
For simple projects, an eval-only run remains valid. AgentV treats the implicit
78+
experiment label as `default` unless a committed experiment is configured. For
79+
specialized harnesses that already have a strong directory contract, AgentV may
80+
support loaders that infer the suite from the directory instead of requiring a
81+
separate YAML file, but those loaders must still lower into the same suite/case
82+
concepts.
83+
6584
Experiment files will live under `experiments/` by convention. AgentV will
6685
support YAML as the canonical authoring path for the abstraction story and TypeScript
6786
as the power-user escape hatch:
@@ -85,6 +104,48 @@ setup:
85104
- script: cp skills/copilot/AGENTS.md AGENTS.md
86105
```
87106
107+
## Workspace boundary
108+
109+
Workspace config belongs with the eval suite or case when it defines the task
110+
scenario being replayed. Examples:
111+
112+
- clone `org/repo` at a specific `commit` or `base_commit`;
113+
- copy starter files, failing tests, fixtures, or issue prompts;
114+
- run task-owned setup hooks that prepare the repo state required by the case;
115+
- declare per-case repo pins or fixture overrides.
116+
117+
Experiment setup belongs with the experiment when it changes the runtime
118+
condition being compared. Examples:
119+
120+
- choose `codex` versus `claude` targets;
121+
- inject an `AGENTS.md`, skill, guideline file, or tool config for an A/B run;
122+
- choose repeat/run policy, timeout, workers, budget, or sandbox mode;
123+
- select a subset of suites or cases for a run campaign.
124+
125+
Rule of thumb: if changing it changes the task being evaluated, put it in the
126+
suite or case workspace. If changing it changes the candidate or run condition
127+
measured against the same task, put it in the experiment.
128+
129+
## Directory-style evals
130+
131+
Convex-style harnesses are a useful counterexample to requiring YAML for every
132+
case. A product-specific benchmark can encode each case as a directory with a
133+
task prompt, reference solution, and executable grader. In AgentV terms, that is
134+
not `experiment -> eval case` with no suite; it is an implicit suite contract
135+
provided by the loader:
136+
137+
```text
138+
evals/<category>/<case>/
139+
TASK.txt # case input
140+
answer/ # reference fixture
141+
grader.test.ts # code-grader assertion
142+
```
143+
144+
AgentV should support this as an import/loader shape when useful, but the core
145+
contract remains `experiment -> eval suite -> eval case`. The experiment applies
146+
runtime bindings to the selected suites/cases; it does not own the prompt,
147+
expected behavior, or grading contract.
148+
88149
`config.yaml` will gain a default experiment pointer so existing `agentv eval`
89150
usage keeps working:
90151

0 commit comments

Comments
 (0)