You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: README.md
+13-6Lines changed: 13 additions & 6 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -19,7 +19,7 @@ Test AI targets on real repo tasks and measure what actually works.
19
19
-**Workspace / fixtures / graders** are task-owned context: repos, setup scripts, files, fixtures, isolation, deterministic checks, and LLM grading prompts.
20
20
-**Target** is the system under test: an agent, provider, gateway, replay target, CLI wrapper, transcript provider, or future app/service wrapper. Each eval selects one `target`, either by name from `targets.yaml` or with an eval-local target object.
21
21
-**Experiment** is the run/result grouping label being measured over that corpus, such as `backend-with-skills` or `backend-without-skills`.
22
-
-**Run controls** configure repeats, exits, timeouts, budgets, thresholds, and completion hooks with top-level fields such as `runs`, `early_exit`, `timeout_seconds`, `budget_usd`, `threshold`, and `on_run_complete`.
22
+
-**Run controls** configure repeats, timeouts, budgets, thresholds, and completion hooks with fields such as `repeat`, `timeout_seconds`, `budget_usd`, `threshold`, and `on_run_complete`.
23
23
-**Run** is one concrete execution of an experiment against a resolved target that writes portable artifacts for readers such as Dashboard, compare, and trend.
Copy file name to clipboardExpand all lines: apps/web/src/content/docs/docs/evaluation/eval-files.mdx
+7-6Lines changed: 7 additions & 6 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -5,7 +5,7 @@ sidebar:
5
5
order: 1
6
6
---
7
7
8
-
Evaluation files define the test cases, graders, workspace lifecycle, and run controls for an evaluation run. Top-level `experiment` is the run/result grouping label, top-level `target` identifies the system under test, and flat fields such as `runs`, `threshold`, `timeout_seconds`, and `budget_usd` control repeated attempts and gates. Workspace reuse belongs under `workspace.isolation`; Docker/container binding belongs under `workspace.docker`. Install, build, and reset commands belong under `workspace.hooks`; runner-specific setup belongs in the `target` object or `targets.yaml`. AgentV supports two eval data formats: YAML and JSONL.
8
+
Evaluation files define the test cases, graders, workspace lifecycle, and run controls for an evaluation run. Top-level `experiment` is the run/result grouping label, top-level `target` identifies the system under test, and fields such as `repeat`, `threshold`, `timeout_seconds`, and `budget_usd` control repeated attempts and gates. Workspace reuse belongs under `workspace.isolation`; Docker/container binding belongs under `workspace.docker`. Install, build, and reset commands belong under `workspace.hooks`; runner-specific setup belongs in the `target` object or `targets.yaml`. AgentV supports two eval data formats: YAML and JSONL.
9
9
10
10
YAML is the canonical portable model. TypeScript helpers, generated fixtures, and Python scripts should lower to the same YAML/JSONL shapes rather than inventing a separate eval contract.
11
11
Eval files describe the task, target binding, and run controls. Concurrency is an operator/run setting: pass `--workers` or set `execution.workers` in `agentv.config.*` / `.agentv/config.yaml` instead of authoring `workers` in eval YAML.
@@ -24,7 +24,7 @@ experiment format.
24
24
it with `imports.tests`, `tests: ./cases.yaml`, or string shorthand; parent
25
25
suite context applies because raw cases do not carry their own suite context.
26
26
- A **wrapper eval** is eval YAML that imports one or more suites with
27
-
`imports.suites` and binds run controls with top-level `target`, `runs`,
27
+
`imports.suites` and binds run controls with top-level `target`, `repeat`,
28
28
`threshold`, `timeout_seconds`, and `budget_usd`.
29
29
Wrapper evals can live anywhere in the repo. A wrapper that imports suites
30
30
with `imports.suites` must not define parent `workspace`; imported suites own
@@ -64,7 +64,9 @@ A wrapper eval stays ordinary eval YAML while choosing a target and run controls
64
64
# experiments/refunds-codex.eval.yaml
65
65
name: refunds-codex
66
66
target: codex-gpt5
67
-
runs: 2
67
+
repeat:
68
+
count: 2
69
+
strategy: pass_any
68
70
69
71
imports:
70
72
suites:
@@ -115,13 +117,12 @@ tests:
115
117
| `category` | Optional slash-delimited analytics taxonomy path. Overrides the category derived from the eval file path. |
116
118
| `target` | Named system under test from `.agentv/targets.yaml` or `--targets` |
| `early_exit` | Optional early exit for repeated attempts |
120
+
| `repeat` | Optional repeat policy with `count`, `strategy`, and `early_exit` |
120
121
| `timeout_seconds` | Optional per-case timeout |
121
122
| `budget_usd` | Optional suite budget |
122
123
| `threshold` | Optional suite quality threshold |
123
124
| `workspace` | Suite-level task environment — inline object or string path to an [external workspace file](/docs/guides/workspace-pool/#external-workspace-config). Repo entries declare identity and checkout pins; acquisition is covered in [Workspace Architecture](/docs/guides/workspace-architecture/#repo-provenance-vs-acquisition). |
124
-
| `imports` | Optional import groups. `imports.suites` imports full child eval suites with their task context. `imports.tests` imports raw test rows into this file's context. Import entries may use scoped `run:` overrides for `threshold`, `timeout_seconds`, and `budget_usd`. |
125
+
| `imports` | Optional import groups. `imports.suites` imports full child eval suites with their task context. `imports.tests` imports raw test rows into this file's context. Import entries may use scoped `run:` overrides for `threshold`, `repeat`, `timeout_seconds`, and `budget_usd`. |
125
126
| `tests` | Inline raw tests or a string path to an external raw-case file or directory. Legacy `tests[].include` entries still load with a migration warning; prefer `imports.suites` or `imports.tests`. |
126
127
| `assertions` | Suite-level graders appended to each test unless `execution.skip_defaults: true` is set on the test |
127
128
| `input` | Suite-level input messages prepended to each test's input unless `execution.skip_defaults: true` is set on the test |
{"timestamp":"2026-02-20T21:40:25.928Z","test_id":"capital-knowledge","suite":"dataset","score":1,"target":"default","trials":[{"attempt":0,"score":1,"verdict":"pass"}],"aggregation":{"strategy":"pass_at_k","passed_attempts":1,"total_attempts":1},"assertions":[{"text":"Correctly identifies Canberra as the capital of Australia","passed":true,"evidence":"The candidate answer provides the correct and complete information, fully matching the reference answer."}]}
2
-
{"timestamp":"2026-02-20T21:40:26.593Z","test_id":"math-basics","suite":"dataset","score":1,"target":"default","trials":[{"attempt":0,"score":1,"verdict":"pass"}],"aggregation":{"strategy":"pass_at_k","passed_attempts":1,"total_attempts":1},"assertions":[{"text":"Explains step-by-step reasoning","passed":true,"evidence":"The candidate answer breaks down the calculation clearly, explains each step, and arrives at the correct answer, matching the reference reasoning."},{"text":"Splits 15 into 10 and 5 for easier calculation","passed":true},{"text":"Calculates partial products (10\u00d77 and 5\u00d77)","passed":true},{"text":"Arrives at correct final answer (105)","passed":true}]}
1
+
{"timestamp":"2026-02-20T21:40:25.928Z","test_id":"capital-knowledge","suite":"dataset","score":1,"target":"default","trials":[{"attempt":0,"score":1,"verdict":"pass"},{"attempt":1,"score":1,"verdict":"pass"}],"aggregation":{"strategy":"pass_any","passed_attempts":2,"total_attempts":2},"assertions":[{"text":"Correctly identifies Canberra as the capital of Australia","passed":true,"evidence":"The candidate answer provides the correct and complete information, fully matching the reference answer."}]}
2
+
{"timestamp":"2026-02-20T21:40:26.593Z","test_id":"math-basics","suite":"dataset","score":1,"target":"default","trials":[{"attempt":0,"score":1,"verdict":"pass"},{"attempt":1,"score":1,"verdict":"pass"}],"aggregation":{"strategy":"pass_any","passed_attempts":2,"total_attempts":2},"assertions":[{"text":"Explains step-by-step reasoning","passed":true,"evidence":"The candidate answer breaks down the calculation clearly, explains each step, and arrives at the correct answer, matching the reference reasoning."},{"text":"Splits 15 into 10 and 5 for easier calculation","passed":true},{"text":"Calculates partial products (10\u00d77 and 5\u00d77)","passed":true},{"text":"Arrives at correct final answer (105)","passed":true}]}
0 commit comments