@@ -37,42 +37,110 @@ Harbor should own:
3737- Harbor ` task.toml ` files and Harbor YAML config;
3838- Opik trace upload through Harbor when enabled.
3939
40+ ## Alignment with experiment separation
41+
42+ The 2026-06-23 experiment/eval separation decision makes runtime binding an
43+ experiment concern. Harbor execution should follow the same split:
44+
45+ - AgentV eval YAML remains the authoring or selection surface for what benchmark
46+ suite is being evaluated.
47+ - AgentV experiment YAML selects or pins the Harbor runner, candidate
48+ agent/model, run policy, and other runtime binding.
49+ - Harbor-authored YAML remains Harbor's own config surface when the standard
50+ suite needs Harbor-specific task packaging or verifier settings.
51+
52+ This means the examples below describe the desired logical fields, but new
53+ runtime fields should be placed on an experiment unless they are genuinely part
54+ of the benchmark suite identity. Do not put candidate agent/model binding in the
55+ eval file for new AgentV-native examples.
56+
4057## Minimal future config surface
4158
42- The AgentV eval file should select Harbor with a nested runner config:
59+ An AgentV eval suite can select the benchmark source without copying Harbor's
60+ task schema or claiming to be the runtime runner:
4361
4462``` yaml
45- name : swebench-verified-codex
63+ name : swebench-verified
4664
47- execution :
48- runner : harbor
49- harbor :
50- dataset : swebench-verified
51- agent : codex
52- model : openai/gpt-5-mini
65+ source :
66+ type : harbor
67+ dataset : swebench-verified
68+ ` ` `
69+
70+ The corresponding experiment selects how that suite runs:
71+
72+ ` ` ` yaml
73+ name : swebench-verified-codex
74+ target : codex-gpt5-mini
75+ evals : evals/swebench-verified.eval.yaml
76+ runner :
77+ type : harbor
78+ options :
5379 opik :
5480 enabled : true
5581` ` `
5682
5783For a Harbor-authored YAML file, use ` config` instead of `dataset`:
5884
5985` ` ` yaml
60- execution:
61- runner: harbor
62- harbor:
63- config: ./harbor/swebench-verified.yaml
86+ source:
87+ type: harbor
88+ config: ./harbor/swebench-verified.yaml
6489` ` `
6590
6691The first implementation should accept exactly one Harbor source selector :
6792` dataset` for a known Harbor dataset id, or `config` for an existing Harbor YAML
6893file. There should be no precedence rule between them. If both are set, fail
6994validation and ask the user to choose one.
7095
71- Keep Harbor-specific options nested under `execution.harbor`. Do not add
72- top-level AgentV fields for Harbor task packaging, verifier images, task patches,
73- or Docker/Compose adapter settings. If a Harbor option becomes too specific to
74- standardize, users should put it in the referenced Harbor YAML file instead of
75- AgentV adding a pass-through field.
96+ Do not combine Harbor suite selection with candidate binding in the eval file :
97+
98+ ` ` ` yaml
99+ # Avoid in eval.yaml
100+ execution:
101+ runner: harbor
102+ harbor:
103+ dataset: swebench-verified
104+ agent: codex
105+ model: openai/gpt-5-mini
106+ ` ` `
107+
108+ Split that shape across the suite and experiment instead :
109+
110+ ` ` ` yaml
111+ # evals/swebench-verified.eval.yaml
112+ name: swebench-verified
113+ source:
114+ type: harbor
115+ dataset: swebench-verified
116+ ` ` `
117+
118+ ` ` ` yaml
119+ # experiments/swebench-verified-codex.yaml
120+ name: swebench-verified-codex
121+ target: codex
122+ model: openai/gpt-5-mini
123+ evals: evals/swebench-verified.eval.yaml
124+ runner:
125+ type: harbor
126+ ` ` `
127+
128+ Keep Harbor suite source selection under `source` in the eval suite. Keep
129+ experiment-side runner selection under `runner.type`, with runner knobs under
130+ ` runner.options` . The eval suite answers "where do these cases come from?"; the
131+ experiment answers "how is this run executed?" Do not use `execution.runner` in
132+ new eval-suite examples because that name collides with the experiment runner.
133+ Do not repeat the runner discriminator as `runner.harbor.options`; `type :
134+ harbor` already provides that namespace.
135+
136+ Do not add top-level AgentV fields for Harbor task packaging, verifier images,
137+ task patches, or Docker/Compose adapter settings. If a Harbor option becomes too
138+ specific to standardize, users should put it in the referenced Harbor YAML file
139+ instead of AgentV adding a pass-through field.
140+
141+ If the Harbor integration later changes the eval source or experiment runner
142+ schema, this ADR should be updated with the final shape. The boundary decision
143+ is stable : Harbor runtime binding is not an eval-case schema extension.
76144
77145# # CLI invocation strategy
78146
@@ -82,8 +150,9 @@ Native evals continue to run with the existing command:
82150agentv eval evals/native.eval.yaml --target codex
83151` ` `
84152
85- Harbor-backed evals should use the same top-level entrypoint and dispatch based
86- on `execution.runner` :
153+ Harbor-backed evals should use the same top-level entrypoint. If no explicit
154+ experiment runner is configured, AgentV may infer Harbor execution from
155+ `source.type : harbor`:
87156
88157` ` ` bash
89158agentv eval evals/swebench-harbor.eval.yaml
@@ -105,9 +174,9 @@ agentv results import harbor --job <harbor-job-id>
105174` ` `
106175
107176Do not overload native `--target` semantics in the first Harbor runner slice.
108- Harbor `agent`, `model`, and matrix behavior should come from
109- ` execution.harbor ` or the referenced Harbor YAML until repeated usage proves a
110- shared AgentV flag is needed.
177+ Harbor `agent`, `model`, and matrix behavior should come from the experiment or
178+ the referenced Harbor YAML until repeated usage proves a shared AgentV flag is
179+ needed.
111180
112181# # Unsupported fields and non-goals
113182
@@ -128,8 +197,9 @@ standard-suite path.
128197# # Implementation sequencing
129198
1301991. Document the native-vs-Harbor boundary and commit alias rules.
131- 2. Add schema validation for optional `execution.runner` and
132- ` execution.harbor` , with no changes to native workspace acquisition.
200+ 2. Add schema validation for eval-suite `source.type : harbor` and exactly one of
201+ ` source.dataset` or `source.config`, plus experiment `runner.type` and
202+ ` runner.options` , with no changes to native workspace acquisition.
1332033. Add a Harbor launch adapter that records job identity and status.
1342044. Add a Harbor result importer that maps rewards, exceptions, timings,
135205 artifacts, and Opik trace URLs into AgentV run bundles.
0 commit comments