Skip to content

Commit 7f2418b

Browse files
committed
docs(readme): clarify experiment layout and rubric schema
1 parent c39c8d8 commit 7f2418b

4 files changed

Lines changed: 944 additions & 814 deletions

File tree

‎README.md‎

Lines changed: 32 additions & 42 deletions
Original file line numberDiff line numberDiff line change
@@ -2,35 +2,9 @@
22

33
**Evaluate AI targets against real repos from the terminal. No server. No signup.**
44

5-
```bash
6-
npm install -g agentv
7-
agentv init
8-
agentv eval evals/example.yaml
9-
```
10-
11-
That's it. Results in seconds, not minutes.
12-
13-
## What it does
14-
15-
AgentV runs evaluation cases against configured targets and scores them with deterministic code graders + customizable LLM graders. Everything lives in Git — YAML eval files, markdown judge prompts, JSONL results.
16-
17-
```yaml
18-
# evals/math.yaml
19-
description: Math problem solving
20-
tests:
21-
- id: addition
22-
input: What is 15 + 27?
23-
expected_output: "42"
24-
assertions:
25-
- type: contains
26-
value: "42"
27-
```
5+
AgentV is a repo-native eval runner for comparing AI targets on real workspace tasks with portable results for local development, CI, Dashboard, compare, and trend.
286

29-
```bash
30-
agentv eval evals/math.yaml
31-
```
32-
33-
## Why AgentV?
7+
## Why?
348

359
- **Local-first** — runs on your machine, no cloud accounts or API keys for eval infrastructure
3610
- **Repo-backed workspaces** — reuse real repos, setup scripts, and existing harnesses instead of rebuilding synthetic tasks
@@ -44,15 +18,15 @@ agentv eval evals/math.yaml
4418

4519
- **Suite / imports / tests** are the task corpus: the prompts, cases, datasets, and imported benchmarks you want to evaluate.
4620
- **Workspace / fixtures / graders** are task-owned context: repos, setup scripts, files, fixtures, deterministic checks, and LLM grading prompts.
47-
- **Target** is the system under test: an agent, model/provider, gateway, replay target, CLI wrapper, transcript provider, or future app/service wrapper.
48-
- **Experiment** names comparison intent: target/model, variant, repeats, gates, timeout/runtime policy, and result grouping.
21+
- **Target** is the system under test: an agent, model/provider, gateway, replay target, CLI wrapper, transcript provider, or future app/service wrapper. Use target ids such as `copilot--claude-opus-4.8` to name the concrete provider/model/variant being compared.
22+
- **Experiment** names comparison intent: target id, repeats, gates, timeout/runtime policy, and result grouping.
4923
- **Run** is one concrete execution that writes portable artifacts for readers such as Dashboard, compare, and trend.
5024

5125
```mermaid
5226
flowchart LR
5327
corpus["Suite / imports / tests<br/>task corpus"]
5428
context["Workspace / fixtures / graders<br/>task-owned context"]
55-
experiment["Experiment<br/>target + variant + runtime policy"]
29+
experiment["Experiment<br/>target id + runtime policy"]
5630
run["Run<br/>concrete execution"]
5731
artifacts["Run artifacts<br/>summary.json + index.jsonl + sidecars"]
5832
readers["Dashboard / compare / trend<br/>derived readers"]
@@ -76,11 +50,20 @@ agentv init
7650

7751
**3. Create an eval** in `evals/`:
7852
```yaml
53+
name: backend-with-skills
7954
description: Code generation quality
8055

8156
experiment:
82-
target: copilot
57+
name: backend-with-skills
58+
target: copilot--claude-opus-4.8
59+
repeat:
60+
count: 3
61+
strategy: pass_at_k
62+
early_exit: true
63+
timeout_seconds: 600
8364
threshold: 0.8
65+
budget_usd: 5
66+
sandbox: auto
8467

8568
tests:
8669
- id: fizzbuzz
@@ -104,30 +87,30 @@ agentv eval evals/my-eval.yaml
10487

10588
**5. Compare two runs** (pass two `index.jsonl` manifests — e.g. before and after a change):
10689
```bash
107-
agentv compare .agentv/results/<experiment>/<before-timestamp>/default/index.jsonl .agentv/results/<experiment>/<after-timestamp>/default/index.jsonl
90+
agentv compare .agentv/results/backend-without-skills/<timestamp>/copilot--claude-opus-4.8/index.jsonl .agentv/results/backend-with-skills/<timestamp>/copilot--claude-opus-4.8/index.jsonl
10891
```
10992

11093
## Results
11194

112-
Each run writes a timestamped bundle under `.agentv/results/<experiment>/<timestamp>/<run-id>/`. The flat `index.jsonl` manifest is the portable surface used by scripts, CI, and `agentv compare`:
95+
Each run writes a timestamped invocation directory under `.agentv/results/<experiment>/<timestamp>/`. In this example, the experiment is `backend-with-skills` and the resolved target id is `copilot--claude-opus-4.8`; model or variant choices are encoded in that target id, not in an extra results folder. The flat `index.jsonl` manifest is the portable surface used by scripts, CI, and `agentv compare`:
11396

11497
```bash
115-
agentv eval evals/my-eval.yaml --output ./run # writes ./run/default/index.jsonl
116-
cat ./run/default/index.jsonl # JSONL results for scripts/CI
98+
agentv eval evals/my-eval.yaml
99+
cat .agentv/results/backend-with-skills/<timestamp>/copilot--claude-opus-4.8/index.jsonl
117100
```
118101

119102
Run bundle layout:
120103

121104
```
122105
.agentv/results/
123-
└── my-eval/ # <experiment> — comparison/run grouping
106+
└── backend-with-skills/ # <experiment> — comparison/run grouping
124107
└── 2026-06-30T08-30-00-000Z/ # <timestamp> — one run
125-
└── default/ # <run-id>
108+
└── copilot--claude-opus-4.8/ # <target> — resolved system under test
126109
├── index.jsonl # flat per-test results (scripts/CI, `agentv compare`)
127110
├── summary.json # run rollup: pass rate, counts, cost
128-
└── fizzbuzz--a1b2c3d4/ # <case-allocation>
111+
└── fizzbuzz--a1b2c3d4/ # <result_dir> for one test case
129112
├── summary.json # per-test rollup across runs
130-
├── task/ # frozen inputs, for reproducibility
113+
├── test/ # generated test bundle: frozen inputs for reproducibility
131114
│ ├── EVAL.yaml # resolved eval spec
132115
│ ├── targets.yaml # resolved target config
133116
│ └── graders/ # grader files used
@@ -149,7 +132,7 @@ Use `evaluate()` when your application owns the run:
149132
import { evaluate } from '@agentv/sdk';
150133

151134
const { results, summary } = await evaluate({
152-
target: { name: 'copilot', provider: 'copilot' },
135+
target: { name: 'copilot--claude-opus-4.8', provider: 'copilot' },
153136
threshold: 0.8,
154137
tests: [
155138
{
@@ -179,8 +162,15 @@ import { defineEval } from '@agentv/sdk';
179162
export default defineEval({
180163
description: 'Code generation quality',
181164
experiment: {
182-
target: 'copilot',
165+
name: 'backend-with-skills',
166+
target: 'copilot--claude-opus-4.8',
167+
repeat: {
168+
count: 3,
169+
strategy: 'pass_at_k',
170+
},
183171
threshold: 0.8,
172+
budgetUsd: 5,
173+
sandbox: 'auto',
184174
},
185175
tests: [
186176
{

‎packages/core/src/evaluation/validation/eval-file.schema.ts‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -87,6 +87,8 @@ const RubricItemSchema = z.object({
8787
score_ranges: z.array(ScoreRangeSchema).optional(),
8888
});
8989

90+
const RubricCriterionSchema = z.union([z.string().min(1), RubricItemSchema]);
91+
9092
// --- Type-specific evaluator schemas ---
9193

9294
const CodeGraderSchema = EvaluatorCommonSchema.extend({
@@ -237,7 +239,7 @@ const EqualsSchema = EvaluatorCommonSchema.extend({
237239

238240
const RubricsSchema = EvaluatorCommonSchema.extend({
239241
type: z.literal('rubrics'),
240-
criteria: z.array(RubricItemSchema).min(1),
242+
criteria: z.array(RubricCriterionSchema).min(1),
241243
});
242244

243245
/** Union of all grader types */

‎packages/core/test/evaluation/validation/eval-file-schema.test.ts‎

Lines changed: 18 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -116,6 +116,24 @@ describe('EvalFileSchema input shorthand', () => {
116116
expect(result.success).toBe(true);
117117
});
118118

119+
it('accepts explicit rubrics criteria string shorthand', () => {
120+
const result = EvalFileSchema.safeParse({
121+
tests: [
122+
{
123+
...baseTest,
124+
assertions: [
125+
{
126+
type: 'rubrics',
127+
criteria: ['Must be polite', 'Must be accurate'],
128+
},
129+
],
130+
},
131+
],
132+
});
133+
134+
expect(result.success).toBe(true);
135+
});
136+
119137
it('accepts flatter imports with optional inline tests', () => {
120138
const result = EvalFileSchema.safeParse({
121139
name: 'wrapper',

0 commit comments

Comments
 (0)