@@ -37,10 +37,22 @@ This decision must also preserve AgentV's existing product boundary:
3737
3838## Vocabulary
3939
40- An eval is a frozen task definition. It includes the prompt or dataset, expected
41- behavior, task-owned workspace fixtures, and assertions. AgentV's LLM-judge,
42- code-grader, deterministic assertions, and hidden or explicit evaluation
43- criteria belong here.
40+ An eval suite is a frozen task-definition boundary. It includes suite metadata,
41+ shared prompt/context, case references, shared assertions or graders, and
42+ task-owned workspace fixtures. AgentV's LLM-judge, code-grader, deterministic
43+ assertions, and hidden or explicit evaluation criteria belong here.
44+
45+ An eval case is one atomic task inside a suite. It includes the case id, prompt
46+ or input, criteria, expected output or reference behavior, case metadata, and
47+ case-specific workspace overrides. A suite can inline cases, point to
48+ ` cases.yaml ` /JSONL, or use a directory convention where each case owns files such
49+ as ` TASK.txt ` , ` PROMPT.md ` , ` answer/ ` , or ` grader.test.ts ` .
50+
51+ In that directory-convention form, ` EVAL.yaml ` may be thin or inferred by a
52+ loader, but the suite layer is still present conceptually: the directory
53+ convention plus runner adapter is the suite contract. This distinction matters
54+ because AgentV is a reusable framework, not a single benchmark harness whose
55+ suite semantics can live only in code.
4456
4557An experiment is a committed or generated run definition. It declares which
4658agent, target, provider, model, harness options, setup steps, run count, timeout,
@@ -62,6 +74,13 @@ task inputs, datasets, assertions, and task fixtures. They should not be the
6274canonical place for which agent, model, harness, setup injection, sandbox, or run
6375matrix executes the task.
6476
77+ For simple projects, an eval-only run remains valid. AgentV treats the implicit
78+ experiment label as ` default ` unless a committed experiment is configured. For
79+ specialized harnesses that already have a strong directory contract, AgentV may
80+ support loaders that infer the suite from the directory instead of requiring a
81+ separate YAML file, but those loaders must still lower into the same suite/case
82+ concepts.
83+
6584Experiment files will live under ` experiments/ ` by convention. AgentV will
6685support YAML as the canonical authoring path for the abstraction story and TypeScript
6786as the power-user escape hatch:
@@ -85,6 +104,48 @@ setup:
85104 - script : cp skills/copilot/AGENTS.md AGENTS.md
86105` ` `
87106
107+ ## Workspace boundary
108+
109+ Workspace config belongs with the eval suite or case when it defines the task
110+ scenario being replayed. Examples:
111+
112+ - clone ` org/repo` at a specific `commit` or `base_commit`;
113+ - copy starter files, failing tests, fixtures, or issue prompts;
114+ - run task-owned setup hooks that prepare the repo state required by the case;
115+ - declare per-case repo pins or fixture overrides.
116+
117+ Experiment setup belongs with the experiment when it changes the runtime
118+ condition being compared. Examples :
119+
120+ - choose `codex` versus `claude` targets;
121+ - inject an `AGENTS.md`, skill, guideline file, or tool config for an A/B run;
122+ - choose repeat/run policy, timeout, workers, budget, or sandbox mode;
123+ - select a subset of suites or cases for a run campaign.
124+
125+ Rule of thumb : if changing it changes the task being evaluated, put it in the
126+ suite or case workspace. If changing it changes the candidate or run condition
127+ measured against the same task, put it in the experiment.
128+
129+ # # Directory-style evals
130+
131+ Convex-style harnesses are a useful counterexample to requiring YAML for every
132+ case. A product-specific benchmark can encode each case as a directory with a
133+ task prompt, reference solution, and executable grader. In AgentV terms, that is
134+ not `experiment -> eval case` with no suite; it is an implicit suite contract
135+ provided by the loader :
136+
137+ ` ` ` text
138+ evals/<category>/<case>/
139+ TASK.txt # case input
140+ answer/ # reference fixture
141+ grader.test.ts # code-grader assertion
142+ ` ` `
143+
144+ AgentV should support this as an import/loader shape when useful, but the core
145+ contract remains `experiment -> eval suite -> eval case`. The experiment applies
146+ runtime bindings to the selected suites/cases; it does not own the prompt,
147+ expected behavior, or grading contract.
148+
88149` config.yaml` will gain a default experiment pointer so existing `agentv eval`
89150usage keeps working :
90151
0 commit comments