You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: apps/web/src/content/docs/docs/reference/result-artifacts.mdx
+4Lines changed: 4 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -45,6 +45,7 @@ The default local layout is:
45
45
transcript-raw.jsonl
46
46
outputs/
47
47
answer.md
48
+
file_changes.diff
48
49
run-2/
49
50
result.json
50
51
grading.json
@@ -54,6 +55,7 @@ The default local layout is:
54
55
transcript-raw.jsonl
55
56
outputs/
56
57
answer.md
58
+
file_changes.diff
57
59
```
58
60
59
61
The `<experiment>` and `<run_id>` directories are storage allocation. They help
@@ -83,6 +85,7 @@ query.
83
85
|`result.json`| Compact per-attempt manifest for one attempt directory. | Loading one attempt without scanning the whole run index. |
84
86
|`grading.json`| Grader outputs, assertions, rubric evidence, execution-metric grader facts, and scoring provenance. | Explaining why a row passed or failed. |
85
87
|`metrics.json`| Derived executor behavior summary, such as tool calls, files touched, shell commands, errors, turns, and output sizes. | Dashboard behavior views, metric-style graders, adapter projections, and lightweight analysis. |
88
+
|`outputs/file_changes.diff`| Full unified diff of workspace file changes when file changes are captured. | Human review and external artifact inspection; LLM and code graders still receive the same full diff through `file_changes`. |
86
89
|`timing.json`| Duration, token usage, cost usage, and source labels such as `provider_reported`, `token_estimated`, `aggregate`, or `unavailable`. | Cost/latency reporting and provider-accounting audits. |
87
90
|`transcript.jsonl`| AgentV-normalized transcript/timeline rows. | Portable human review, replay, transcript-aware graders, and tool-trajectory analysis. |
88
91
|`transcript-raw.jsonl`| Native provider or harness evidence when available. | Parser debugging, forensic review, and preserving source bytes without making provider schemas public AgentV fields. |
|`tool_call_events`, `tool_call_counts`, `tool_category_counts`, `shell_commands`, `files_read`, `files_modified`, `web_fetches`, `errors`, `reasoning_blocks`, `thinking_blocks`, `total_turns`| AgentV/Vercel-style behavior summary when source data includes it |
159
161
160
162
Vercel `@vercel/agent-eval``results.o11y` maps into AgentV like this:
0 commit comments