Skip to content

Commit 103cefe

Browse files
committed
docs(readme): align eval policy examples
1 parent 1877ceb commit 103cefe

1 file changed

Lines changed: 5 additions & 6 deletions

File tree

‎README.md‎

Lines changed: 5 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -19,7 +19,7 @@ Test AI targets on real repo tasks and measure what actually works.
1919
- **Workspace / fixtures / graders** are task-owned context: repos, setup scripts, files, fixtures, isolation, deterministic checks, and LLM grading prompts.
2020
- **Target** is the system under test: an agent, provider, gateway, replay target, CLI wrapper, transcript provider, or future app/service wrapper. Use `model` when you need to override the target's default model for a run.
2121
- **Experiment** is the named condition being measured over that corpus, such as `backend-with-skills` or `backend-without-skills`.
22-
- **Policy** controls how AgentV executes and gates the eval: runs, early exit, thresholds, timeouts, and budgets. It is not the experiment identity.
22+
- **Policy** controls how AgentV executes and gates the eval: runs, thresholds, timeouts, and budgets. It is not the experiment identity.
2323
- **Run** is one concrete execution of an experiment against a target/model that writes portable artifacts for readers such as Dashboard, compare, and trend.
2424

2525
```mermaid
@@ -57,7 +57,7 @@ agentv init
5757

5858
**3. Create an eval** in `evals/`:
5959
```yaml
60-
experiment: backend-with-skills
60+
name: backend-with-skills
6161
description: Code generation quality
6262
target: copilot-sdk
6363
model: claude-sonnet-4.6
@@ -67,7 +67,6 @@ workspace:
6767

6868
policy:
6969
runs: 3
70-
early_exit: false
7170
timeout_seconds: 600
7271
threshold: 0.8
7372
budget_usd: 5
@@ -97,7 +96,7 @@ agentv compare .agentv/results/backend-without-skills/<timestamp>/copilot-sdk--c
9796

9897
## Results
9998

100-
Each run writes a timestamped invocation directory under `.agentv/results/<experiment>/<timestamp>/`. In this example, `experiment: backend-with-skills` names the condition being measured, `target: copilot-sdk` selects the system under test, and `model: claude-sonnet-4.6` overrides that target's default model. The resolved target identity is still `copilot-sdk--claude-sonnet-4.6` so CI baselines can distinguish model changes. The flat `index.jsonl` manifest is the portable surface used by scripts, CI, and `agentv compare`:
99+
Each run writes a timestamped invocation directory under `.agentv/results/<experiment>/<timestamp>/`. In this example, `name: backend-with-skills` names the condition being measured, `target: copilot-sdk` selects the system under test, and `model: claude-sonnet-4.6` overrides that target's default model. The resolved target identity is still `copilot-sdk--claude-sonnet-4.6` so CI baselines can distinguish model changes. The flat `index.jsonl` manifest is the portable surface used by scripts, CI, and `agentv compare`:
101100

102101
```bash
103102
agentv eval evals/my-eval.yaml
@@ -163,13 +162,13 @@ Use `defineEval()` when you want AgentV to run the TypeScript eval file:
163162
import { defineEval } from '@agentv/sdk';
164163

165164
export default defineEval({
166-
experiment: 'backend-with-skills',
165+
name: 'backend-with-skills',
167166
description: 'Code generation quality',
168167
target: 'copilot-sdk',
169168
model: 'claude-sonnet-4.6',
170169
policy: {
171170
runs: 3,
172-
earlyExit: false,
171+
timeoutSeconds: 600,
173172
threshold: 0.8,
174173
budgetUsd: 5,
175174
},

0 commit comments

Comments
 (0)