You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: README.md
+47-52Lines changed: 47 additions & 52 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,36 +1,8 @@
1
1
# AgentV
2
2
3
-
**Evaluate AI targets against real repos from the terminal. No server. No signup.**
3
+
Test AI targets on real repo tasks and measure what actually works.
4
4
5
-
```bash
6
-
npm install -g agentv
7
-
agentv init
8
-
agentv eval evals/example.yaml
9
-
```
10
-
11
-
That's it. Results in seconds, not minutes.
12
-
13
-
## What it does
14
-
15
-
AgentV runs evaluation cases against configured targets and scores them with deterministic code graders + customizable LLM graders. Everything lives in Git — YAML eval files, markdown judge prompts, JSONL results.
16
-
17
-
```yaml
18
-
# evals/math.yaml
19
-
description: Math problem solving
20
-
tests:
21
-
- id: addition
22
-
input: What is 15 + 27?
23
-
expected_output: "42"
24
-
assertions:
25
-
- type: contains
26
-
value: "42"
27
-
```
28
-
29
-
```bash
30
-
agentv eval evals/math.yaml
31
-
```
32
-
33
-
## Why AgentV?
5
+
## Why?
34
6
35
7
-**Local-first** — runs on your machine, no cloud accounts or API keys for eval infrastructure
36
8
-**Repo-backed workspaces** — reuse real repos, setup scripts, and existing harnesses instead of rebuilding synthetic tasks
@@ -42,23 +14,28 @@ agentv eval evals/math.yaml
42
14
43
15
## Core Concepts
44
16
45
-
-**Suite / imports / tests** are the task corpus: the prompts, cases, datasets, and imported benchmarks you want to evaluate.
46
-
-**Workspace / fixtures / graders** are task-owned context: repos, setup scripts, files, fixtures, deterministic checks, and LLM grading prompts.
47
-
-**Target** is the system under test: an agent, model/provider, gateway, replay target, CLI wrapper, transcript provider, or future app/service wrapper.
48
-
-**Experiment** names comparison intent: target/model, variant, repeats, gates, timeout/runtime policy, and result grouping.
17
+
-**Imports / tests** are the task corpus: the prompts, cases, datasets, and imported benchmarks you want to evaluate.
18
+
-**Workspace / fixtures / graders** are task-owned context: repos, setup scripts, files, fixtures, isolation, deterministic checks, and LLM grading prompts.
19
+
-**Target** is the system under test: an agent, model/provider, gateway, replay target, CLI wrapper, transcript provider, or future app/service wrapper. Use target ids such as `copilot--claude-opus-4.8` to name the concrete provider/model/variant being compared.
20
+
-**Policy** controls how AgentV runs and gates the eval: repeats, thresholds, timeouts, and budgets.
21
+
-**Experiment** is created automatically from the eval definition. The top-level `name` gives the experiment namespace, `target` identifies the system under test, and AgentV groups concrete runs under that resolved identity.
49
22
-**Run** is one concrete execution that writes portable artifacts for readers such as Dashboard, compare, and trend.
Each run writes a timestamped bundle under `.agentv/results/<experiment>/<timestamp>/<run-id>/`. The flat `index.jsonl` manifest is the portable surface used by scripts, CI, and `agentv compare`:
97
+
Each run writes a timestamped invocation directory under `.agentv/results/<experiment>/<timestamp>/`. In this example, the top-level eval `name` creates the `backend-with-skills` experiment namespace and the resolved target id is `copilot--claude-opus-4.8`; model or variant choices are encoded in that target id, not in an extra results folder. The flat `index.jsonl` manifest is the portable surface used by scripts, CI, and `agentv compare`:
0 commit comments