Skip to content

Commit bcf3584

Browse files
committed
docs(readme): clarify experiment layout and rubric schema
1 parent c39c8d8 commit bcf3584

4 files changed

Lines changed: 959 additions & 824 deletions

File tree

‎README.md‎

Lines changed: 47 additions & 52 deletions
Original file line numberDiff line numberDiff line change
@@ -1,36 +1,8 @@
11
# AgentV
22

3-
**Evaluate AI targets against real repos from the terminal. No server. No signup.**
3+
Test AI targets on real repo tasks and measure what actually works.
44

5-
```bash
6-
npm install -g agentv
7-
agentv init
8-
agentv eval evals/example.yaml
9-
```
10-
11-
That's it. Results in seconds, not minutes.
12-
13-
## What it does
14-
15-
AgentV runs evaluation cases against configured targets and scores them with deterministic code graders + customizable LLM graders. Everything lives in Git — YAML eval files, markdown judge prompts, JSONL results.
16-
17-
```yaml
18-
# evals/math.yaml
19-
description: Math problem solving
20-
tests:
21-
- id: addition
22-
input: What is 15 + 27?
23-
expected_output: "42"
24-
assertions:
25-
- type: contains
26-
value: "42"
27-
```
28-
29-
```bash
30-
agentv eval evals/math.yaml
31-
```
32-
33-
## Why AgentV?
5+
## Why?
346

357
- **Local-first** — runs on your machine, no cloud accounts or API keys for eval infrastructure
368
- **Repo-backed workspaces** — reuse real repos, setup scripts, and existing harnesses instead of rebuilding synthetic tasks
@@ -42,23 +14,28 @@ agentv eval evals/math.yaml
4214

4315
## Core Concepts
4416

45-
- **Suite / imports / tests** are the task corpus: the prompts, cases, datasets, and imported benchmarks you want to evaluate.
46-
- **Workspace / fixtures / graders** are task-owned context: repos, setup scripts, files, fixtures, deterministic checks, and LLM grading prompts.
47-
- **Target** is the system under test: an agent, model/provider, gateway, replay target, CLI wrapper, transcript provider, or future app/service wrapper.
48-
- **Experiment** names comparison intent: target/model, variant, repeats, gates, timeout/runtime policy, and result grouping.
17+
- **Imports / tests** are the task corpus: the prompts, cases, datasets, and imported benchmarks you want to evaluate.
18+
- **Workspace / fixtures / graders** are task-owned context: repos, setup scripts, files, fixtures, isolation, deterministic checks, and LLM grading prompts.
19+
- **Target** is the system under test: an agent, model/provider, gateway, replay target, CLI wrapper, transcript provider, or future app/service wrapper. Use target ids such as `copilot--claude-opus-4.8` to name the concrete provider/model/variant being compared.
20+
- **Policy** controls how AgentV runs and gates the eval: repeats, thresholds, timeouts, and budgets.
21+
- **Experiment** is created automatically from the eval definition. The top-level `name` gives the experiment namespace, `target` identifies the system under test, and AgentV groups concrete runs under that resolved identity.
4922
- **Run** is one concrete execution that writes portable artifacts for readers such as Dashboard, compare, and trend.
5023

5124
```mermaid
5225
flowchart LR
53-
corpus["Suite / imports / tests<br/>task corpus"]
26+
corpus["Imports / tests<br/>task corpus"]
5427
context["Workspace / fixtures / graders<br/>task-owned context"]
55-
experiment["Experiment<br/>target + variant + runtime policy"]
28+
target["Target<br/>system under test"]
29+
policy["Policy<br/>runtime + gates"]
30+
experiment["Experiment<br/>auto-created grouping"]
5631
run["Run<br/>concrete execution"]
5732
artifacts["Run artifacts<br/>summary.json + index.jsonl + sidecars"]
5833
readers["Dashboard / compare / trend<br/>derived readers"]
5934
6035
corpus --> run
6136
context --> run
37+
target --> experiment
38+
policy --> experiment
6239
experiment --> run
6340
run --> artifacts
6441
artifacts --> readers
@@ -76,21 +53,29 @@ agentv init
7653

7754
**3. Create an eval** in `evals/`:
7855
```yaml
56+
name: backend-with-skills
7957
description: Code generation quality
58+
target: copilot--claude-opus-4.8
59+
60+
workspace:
61+
isolation: per_case
8062

81-
experiment:
82-
target: copilot
63+
policy:
64+
repeat:
65+
count: 3
66+
strategy: pass_at_k
67+
early_exit: false
68+
timeout_seconds: 600
8369
threshold: 0.8
70+
budget_usd: 5
8471

8572
tests:
8673
- id: fizzbuzz
8774
input: Write FizzBuzz in Python
8875
assertions:
8976
- type: contains
9077
value: "fizz"
91-
- type: rubrics
92-
criteria:
93-
- Implements correct FizzBuzz logic for multiples of 3, 5, and 15
78+
- Implements correct FizzBuzz logic for multiples of 3, 5, and 15
9479
- type: code-grader
9580
command: ["python3", "./validators/check_syntax.py"]
9681
- type: llm-grader
@@ -104,30 +89,30 @@ agentv eval evals/my-eval.yaml
10489

10590
**5. Compare two runs** (pass two `index.jsonl` manifests — e.g. before and after a change):
10691
```bash
107-
agentv compare .agentv/results/<experiment>/<before-timestamp>/default/index.jsonl .agentv/results/<experiment>/<after-timestamp>/default/index.jsonl
92+
agentv compare .agentv/results/backend-without-skills/<timestamp>/copilot--claude-opus-4.8/index.jsonl .agentv/results/backend-with-skills/<timestamp>/copilot--claude-opus-4.8/index.jsonl
10893
```
10994

11095
## Results
11196

112-
Each run writes a timestamped bundle under `.agentv/results/<experiment>/<timestamp>/<run-id>/`. The flat `index.jsonl` manifest is the portable surface used by scripts, CI, and `agentv compare`:
97+
Each run writes a timestamped invocation directory under `.agentv/results/<experiment>/<timestamp>/`. In this example, the top-level eval `name` creates the `backend-with-skills` experiment namespace and the resolved target id is `copilot--claude-opus-4.8`; model or variant choices are encoded in that target id, not in an extra results folder. The flat `index.jsonl` manifest is the portable surface used by scripts, CI, and `agentv compare`:
11398

11499
```bash
115-
agentv eval evals/my-eval.yaml --output ./run # writes ./run/default/index.jsonl
116-
cat ./run/default/index.jsonl # JSONL results for scripts/CI
100+
agentv eval evals/my-eval.yaml
101+
cat .agentv/results/backend-with-skills/<timestamp>/copilot--claude-opus-4.8/index.jsonl
117102
```
118103

119104
Run bundle layout:
120105

121106
```
122107
.agentv/results/
123-
└── my-eval/ # <experiment> — comparison/run grouping
108+
└── backend-with-skills/ # <experiment> — comparison/run grouping
124109
└── 2026-06-30T08-30-00-000Z/ # <timestamp> — one run
125-
└── default/ # <run-id>
110+
└── copilot--claude-opus-4.8/ # <target> — resolved system under test
126111
├── index.jsonl # flat per-test results (scripts/CI, `agentv compare`)
127112
├── summary.json # run rollup: pass rate, counts, cost
128-
└── fizzbuzz--a1b2c3d4/ # <case-allocation>
113+
└── fizzbuzz--a1b2c3d4/ # <result_dir> for one test case
129114
├── summary.json # per-test rollup across runs
130-
├── task/ # frozen inputs, for reproducibility
115+
├── test/ # generated test bundle: frozen inputs for reproducibility
131116
│ ├── EVAL.yaml # resolved eval spec
132117
│ ├── targets.yaml # resolved target config
133118
│ └── graders/ # grader files used
@@ -149,7 +134,7 @@ Use `evaluate()` when your application owns the run:
149134
import { evaluate } from '@agentv/sdk';
150135

151136
const { results, summary } = await evaluate({
152-
target: { name: 'copilot', provider: 'copilot' },
137+
task: async (input) => runMyAppTarget(input),
153138
threshold: 0.8,
154139
tests: [
155140
{
@@ -177,10 +162,20 @@ Use `defineEval()` when you want AgentV to run the TypeScript eval file:
177162
import { defineEval } from '@agentv/sdk';
178163

179164
export default defineEval({
165+
name: 'backend-with-skills',
180166
description: 'Code generation quality',
181-
experiment: {
182-
target: 'copilot',
167+
target: 'copilot--claude-opus-4.8',
168+
policy: {
169+
repeat: {
170+
count: 3,
171+
strategy: 'pass_at_k',
172+
},
173+
earlyExit: false,
183174
threshold: 0.8,
175+
budgetUsd: 5,
176+
},
177+
workspace: {
178+
isolation: 'per_case',
184179
},
185180
tests: [
186181
{

‎packages/core/src/evaluation/validation/eval-file.schema.ts‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -87,6 +87,8 @@ const RubricItemSchema = z.object({
8787
score_ranges: z.array(ScoreRangeSchema).optional(),
8888
});
8989

90+
const RubricCriterionSchema = z.union([z.string().min(1), RubricItemSchema]);
91+
9092
// --- Type-specific evaluator schemas ---
9193

9294
const CodeGraderSchema = EvaluatorCommonSchema.extend({
@@ -237,7 +239,7 @@ const EqualsSchema = EvaluatorCommonSchema.extend({
237239

238240
const RubricsSchema = EvaluatorCommonSchema.extend({
239241
type: z.literal('rubrics'),
240-
criteria: z.array(RubricItemSchema).min(1),
242+
criteria: z.array(RubricCriterionSchema).min(1),
241243
});
242244

243245
/** Union of all grader types */

‎packages/core/test/evaluation/validation/eval-file-schema.test.ts‎

Lines changed: 18 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -116,6 +116,24 @@ describe('EvalFileSchema input shorthand', () => {
116116
expect(result.success).toBe(true);
117117
});
118118

119+
it('accepts explicit rubrics criteria string shorthand', () => {
120+
const result = EvalFileSchema.safeParse({
121+
tests: [
122+
{
123+
...baseTest,
124+
assertions: [
125+
{
126+
type: 'rubrics',
127+
criteria: ['Must be polite', 'Must be accurate'],
128+
},
129+
],
130+
},
131+
],
132+
});
133+
134+
expect(result.success).toBe(true);
135+
});
136+
119137
it('accepts flatter imports with optional inline tests', () => {
120138
const result = EvalFileSchema.safeParse({
121139
name: 'wrapper',

0 commit comments

Comments
 (0)