You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: README.md
+32-42Lines changed: 32 additions & 42 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -2,35 +2,9 @@
2
2
3
3
**Evaluate AI targets against real repos from the terminal. No server. No signup.**
4
4
5
-
```bash
6
-
npm install -g agentv
7
-
agentv init
8
-
agentv eval evals/example.yaml
9
-
```
10
-
11
-
That's it. Results in seconds, not minutes.
12
-
13
-
## What it does
14
-
15
-
AgentV runs evaluation cases against configured targets and scores them with deterministic code graders + customizable LLM graders. Everything lives in Git — YAML eval files, markdown judge prompts, JSONL results.
16
-
17
-
```yaml
18
-
# evals/math.yaml
19
-
description: Math problem solving
20
-
tests:
21
-
- id: addition
22
-
input: What is 15 + 27?
23
-
expected_output: "42"
24
-
assertions:
25
-
- type: contains
26
-
value: "42"
27
-
```
5
+
AgentV is a repo-native eval runner for comparing AI targets on real workspace tasks with portable results for local development, CI, Dashboard, compare, and trend.
28
6
29
-
```bash
30
-
agentv eval evals/math.yaml
31
-
```
32
-
33
-
## Why AgentV?
7
+
## Why?
34
8
35
9
-**Local-first** — runs on your machine, no cloud accounts or API keys for eval infrastructure
36
10
-**Repo-backed workspaces** — reuse real repos, setup scripts, and existing harnesses instead of rebuilding synthetic tasks
@@ -44,15 +18,15 @@ agentv eval evals/math.yaml
44
18
45
19
-**Suite / imports / tests** are the task corpus: the prompts, cases, datasets, and imported benchmarks you want to evaluate.
46
20
-**Workspace / fixtures / graders** are task-owned context: repos, setup scripts, files, fixtures, deterministic checks, and LLM grading prompts.
47
-
-**Target** is the system under test: an agent, model/provider, gateway, replay target, CLI wrapper, transcript provider, or future app/service wrapper.
48
-
-**Experiment** names comparison intent: target/model, variant, repeats, gates, timeout/runtime policy, and result grouping.
21
+
-**Target** is the system under test: an agent, model/provider, gateway, replay target, CLI wrapper, transcript provider, or future app/service wrapper. Use target ids such as `copilot--claude-opus-4.8` to name the concrete provider/model/variant being compared.
22
+
-**Experiment** names comparison intent: target id, repeats, gates, timeout/runtime policy, and result grouping.
49
23
-**Run** is one concrete execution that writes portable artifacts for readers such as Dashboard, compare, and trend.
Each run writes a timestamped bundle under `.agentv/results/<experiment>/<timestamp>/<run-id>/`. The flat `index.jsonl` manifest is the portable surface used by scripts, CI, and `agentv compare`:
95
+
Each run writes a timestamped invocation directory under `.agentv/results/<experiment>/<timestamp>/`. In this example, the experiment is `backend-with-skills` and the resolved target id is `copilot--claude-opus-4.8`; model or variant choices are encoded in that target id, not in an extra results folder. The flat `index.jsonl` manifest is the portable surface used by scripts, CI, and `agentv compare`:
0 commit comments