diff --git a/.agents/verification.md b/.agents/verification.md index 6dbff41b4..0f02e1237 100644 --- a/.agents/verification.md +++ b/.agents/verification.md @@ -120,6 +120,18 @@ bun apps/cli/src/cli.ts eval examples/features/rubric/evals/dataset.eval.yaml -- 4. Update baseline files if output format changes. Baselines live next to eval YAML files as `*.baseline.jsonl`. 5. `--dry-run` returns schema-valid mock responses, but the scores are not meaningful. Use it only for plumbing and harness checks. +## Live Dogfood for Eval and Experiment Changes + +Use live dogfood before marking PRs ready when they affect eval execution, experiments, repeat runs, targets, providers, graders, or artifact provenance. + +- Live means both sides are real: a live agent/provider target and a live grader target. Do not count `mock`, `--dry-run`, or deterministic-only assertions as dogfood for these changes. +- Prefer the smallest realistic eval: one or two cases, bounded timeouts, and `workers: 1` for heavyweight agent providers. +- For native experiment changes, run through `agentv eval run ... --experiment ` so resolution, setup, scripts, target selection, run knobs, and artifact metadata are exercised together. +- For repeat-run changes, use an experiment-level repeat config with `count >= 2`, `early_exit: false` when validating all attempts are persisted. Inspect root `index.jsonl`, root `benchmark.json`, and the repeated case folder. The repeated case folder should carry aggregate `summary.json` with flattened snake_case timing fields plus AgentV aggregate `grading.json`; attempt-specific outputs and transcripts live under `run-N/`. Each `run-N/` folder should contain `result.json`, `grading.json`, `transcript.json`, `transcript-raw.jsonl`, and `outputs/answer.md`. Do not write per-run `metrics.json`; timing and o11y fields belong in `result.json`, and `result.json` points at `./grading.json` through `grading_path`. +- For local OpenAI-compatible grading through the OAuth proxy, use `endpoint: http://127.0.0.1:10531/v1`, but still route `api_key` and `model` through environment references such as `${{ LOCAL_OPENAI_PROXY_API_KEY }}` and `${{ LOCAL_OPENAI_PROXY_MODEL }}`. Literal secrets and literal model values are intentionally rejected by target validation unless a resolver explicitly allows them. +- Preserve review evidence in `agentv-private` on an `evidence/` branch. Include the run bundle, source eval/experiment/targets files, a short README, an artifact tree, and screenshots when folder structure or UI behavior is under review. +- If comparing against an external convention such as Vercel `agent-eval`, verify both semantic provenance and the physical `run-N` artifact layout for repeat runs. + ## Checking Grader Score Ranges Use `scripts/check-grader-scores.ts` as a post-processor after an eval run. diff --git a/AGENTS.md b/AGENTS.md index 110f0d3b5..da2e54680 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -43,6 +43,7 @@ Read the full rationale and examples in [.agents/product-boundary.md](.agents/pr - Non-trivial work needs a plan or task list. If the implementation surface starts to balloon, stop and re-plan. - Large or high-risk PRs need meaningful, reviewable commits for each coherent change. Rewrite only the PR branch with `git push --force-with-lease` when needed to replace WIP or accidental squashed history before review. - Manual red/green UAT is blocking before a branch is ready for review. GitHub Actions is the authoritative merge gate. +- For eval execution, experiments, repeat runs, providers, graders, or artifact-layout changes, dogfood with a live provider and a real LLM grader before marking ready. Mock graders, dry-run, and deterministic-only smoke tests are useful plumbing checks, but they are not live dogfood. Use canonical `.agentv/results//` output and publish private evidence. See [.agents/verification.md](.agents/verification.md). - For browser or screenshot UAT, keep evidence out of the public repo and publish reviewable artifacts to an `agentv-private` evidence branch. See [.agents/verification.md](.agents/verification.md). - Wire formats are `snake_case`; internal TypeScript is `camelCase`. Translate only at the boundary. - In AgentV, a `project` holds runs, traces, and experiments; a `benchmark` is a curated eval suite. Do not collapse those terms. diff --git a/apps/cli/src/commands/eval/artifact-writer.ts b/apps/cli/src/commands/eval/artifact-writer.ts index 1d6ce52e6..09e2d098e 100644 --- a/apps/cli/src/commands/eval/artifact-writer.ts +++ b/apps/cli/src/commands/eval/artifact-writer.ts @@ -6,6 +6,7 @@ import { type BenchmarkArtifact, type EvalTest, type EvaluationResult, + type ExperimentArtifactMetadata, type ExportDuplicatePolicy, type GradingArtifact, type IndexArtifactEntry, @@ -228,6 +229,7 @@ export async function writeArtifactsFromResults( options?: { evalFile?: string; experiment?: string; + experimentMetadata?: ExperimentArtifactMetadata; plannedTestCount?: number; runId?: string; duplicatePolicy?: ExportDuplicatePolicy; @@ -245,6 +247,7 @@ export async function writeArtifactsFromResults( return writeCoreArtifactsFromResults(results, outputDir, { evalFile: options?.evalFile, experiment: options?.experiment, + experimentMetadata: options?.experimentMetadata, plannedTestCount: options?.plannedTestCount, runId: options?.runId, duplicatePolicy: options?.duplicatePolicy, diff --git a/apps/cli/src/commands/eval/commands/run.ts b/apps/cli/src/commands/eval/commands/run.ts index 5b74c3884..a498f4ccc 100644 --- a/apps/cli/src/commands/eval/commands/run.ts +++ b/apps/cli/src/commands/eval/commands/run.ts @@ -11,7 +11,6 @@ import { } from 'cmd-ts'; import { runEvalCommand } from '../run-eval.js'; -import { resolveEvalPaths } from '../shared.js'; export const evalRunCommand = command({ name: 'eval', @@ -264,14 +263,6 @@ export const evalRunCommand = command({ }), }, handler: async (args) => { - // Launch interactive wizard when no eval paths and stdin is a TTY - if (args.evalPaths.length === 0 && process.stdin.isTTY) { - const { launchInteractiveWizard } = await import('../interactive.js'); - await launchInteractiveWizard(); - return; - } - - const resolvedPaths = await resolveEvalPaths(args.evalPaths, process.cwd()); if (args.budgetUsd !== undefined && args.budgetUsd <= 0) { console.error('Error: --budget-usd must be a positive number.'); process.exit(2); @@ -330,7 +321,7 @@ export const evalRunCommand = command({ recordReplay: args.recordReplay, recordReplayVariant: args.recordReplayVariant, }; - const result = await runEvalCommand({ testFiles: resolvedPaths, rawOptions }); + const result = await runEvalCommand({ testFiles: args.evalPaths, rawOptions }); if (result?.allExecutionErrors) { process.exit(2); } diff --git a/apps/cli/src/commands/eval/run-eval.ts b/apps/cli/src/commands/eval/run-eval.ts index 47a6d4940..5e2eb1058 100644 --- a/apps/cli/src/commands/eval/run-eval.ts +++ b/apps/cli/src/commands/eval/run-eval.ts @@ -1,27 +1,38 @@ +import { spawn } from 'node:child_process'; import { constants, existsSync, mkdirSync } from 'node:fs'; import { access, readFile } from 'node:fs/promises'; import path from 'node:path'; import { pathToFileURL } from 'node:url'; import { + DEFAULT_EVAL_PATTERNS, DEFAULT_THRESHOLD, + type EvalTargetRef, type EvalTest, type EvaluationCache, type EvaluationResult, type ExecutionDefaults, + type ExperimentArtifactMetadata, + type ExperimentConfig, + type ExperimentScript, type FailOnError, type OtelTraceExporter as OtelTraceExporterType, type ResolvedTarget, ResponseCache, RunBudgetTracker, type TrialsConfig, + buildExperimentArtifactMetadata, buildTraceFromMessages, runEvaluation as defaultRunEvaluation, deriveCategory, + deriveExperimentNameFromPath, ensureVSCodeSubagents, + isExperimentFileReference, loadConfig, + loadExperimentConfig, loadTestSuite, loadTsConfig, + resolveDefaultExperimentReference, resolveTargetDefinition, shouldEnableCache, shouldSkipCacheForTemperature, @@ -67,7 +78,7 @@ import { loadNonErrorResults, } from './retry-errors.js'; import { resolveCachedRunDir, saveRunCache } from './run-cache.js'; -import { findRepoRoot } from './shared.js'; +import { findRepoRoot, resolveEvalPaths } from './shared.js'; import { calculateEvaluationSummary, formatEvaluationSummary, @@ -141,6 +152,10 @@ interface NormalizedOptions { readonly recordReplay?: string; readonly recordReplayVariant?: string; readonly experiment?: string; + readonly experimentConfig?: ExperimentConfig; + readonly experimentMetadata?: ExperimentArtifactMetadata; + readonly experimentTargetRefs?: readonly EvalTargetRef[]; + readonly experimentTrialsConfig?: TrialsConfig; readonly budgetUsd?: number; readonly sourceMetadataByEvalFile?: ReadonlyMap>; readonly resultsOverrides?: ResultsPublishOverrides; @@ -554,6 +569,321 @@ function buildDefaultOutputPathForExperiment( return path.join(runDir, 'index.jsonl'); } +function normalizeTsDefaultExperiment( + config: Awaited> | null, +): string | undefined { + return ( + normalizeString(config?.experiments?.default) ?? normalizeString(config?.defaultExperiment) + ); +} + +type ResolvedExperimentForRun = { + readonly name?: string; + readonly config?: ExperimentConfig; +}; + +async function resolveExperimentForRun(params: { + readonly cwd: string; + readonly explicitExperiment?: string; + readonly yamlDefaultExperiment?: string; + readonly tsDefaultExperiment?: string; +}): Promise { + const experimentRef = + params.explicitExperiment ?? params.yamlDefaultExperiment ?? params.tsDefaultExperiment; + if (!experimentRef) { + return {}; + } + + const experimentPath = resolveExperimentFilePath(params.cwd, experimentRef); + if (!experimentPath) { + if (isExperimentFileReference(experimentRef)) { + throw new Error(`Experiment file not found: ${experimentRef}`); + } + return { name: experimentRef }; + } + + const config = await loadExperimentConfig(experimentPath); + return { + name: config.name ?? deriveExperimentNameFromPath(experimentPath), + config, + }; +} + +function resolveExperimentFilePath(cwd: string, experimentRef: string): string | undefined { + if (isExperimentFileReference(experimentRef)) { + const experimentPath = path.isAbsolute(experimentRef) + ? experimentRef + : path.resolve(cwd, experimentRef); + return existsSync(experimentPath) ? experimentPath : undefined; + } + + for (const ext of ['yaml', 'yml', 'ts', 'js', 'mts', 'mjs']) { + const candidate = path.resolve(cwd, 'experiments', `${experimentRef}.${ext}`); + if (existsSync(candidate)) { + return candidate; + } + } + return undefined; +} + +function applyExperimentOptions( + options: NormalizedOptions, + experiment: ExperimentConfig | undefined, +): NormalizedOptions { + if (!experiment) { + return options; + } + + const experimentTargetRefs = buildExperimentTargetRefs(experiment); + const experimentTargetNames = experimentTargetRefs?.map((target) => target.name) ?? []; + const experimentTarget = + experiment.target && experiment.target.trim().length > 0 ? experiment.target : undefined; + const nextCliTargets = + options.cliTargets.length > 0 + ? options.cliTargets + : experimentTargetNames.length > 0 + ? experimentTargetNames + : experimentTarget + ? [experimentTarget] + : options.cliTargets; + + const workspaceMode = + options.workspaceMode ?? readExperimentWorkspaceMode(experiment.workspace?.mode); + const workspacePath = options.workspacePath ?? readExperimentWorkspacePath(experiment.workspace); + const experimentFilter = normalizeExperimentCaseFilter(experiment.evals); + + return { + ...options, + target: options.target ?? (nextCliTargets.length === 1 ? nextCliTargets[0] : undefined), + cliTargets: nextCliTargets, + agentTimeoutSeconds: options.agentTimeoutSeconds ?? experiment.timeoutSeconds, + workers: options.workers ?? experiment.workers, + workspaceMode: workspacePath ? 'static' : workspaceMode, + workspacePath, + filter: options.filter ?? experimentFilter, + budgetUsd: options.budgetUsd ?? experiment.budgetUsd, + experimentConfig: experiment, + experimentMetadata: buildExperimentArtifactMetadata(experiment), + experimentTargetRefs: options.cliTargets.length === 0 ? experimentTargetRefs : undefined, + experimentTrialsConfig: buildExperimentTrialsConfig(experiment), + }; +} + +function buildExperimentTargetRefs( + experiment: ExperimentConfig, +): readonly EvalTargetRef[] | undefined { + if (!experiment.targets || experiment.targets.length === 0) { + return undefined; + } + return experiment.targets.map((target) => { + if (typeof target === 'string') { + return { name: target }; + } + return { + name: target.name, + ...(target.useTarget !== undefined && { use_target: target.useTarget }), + ...(target.hooks !== undefined && { + hooks: target.hooks as EvalTargetRef['hooks'], + }), + }; + }); +} + +function buildExperimentTrialsConfig(experiment: ExperimentConfig): TrialsConfig | undefined { + if (experiment.repeat) { + if (experiment.repeat.count <= 1) { + return undefined; + } + return { + count: experiment.repeat.count, + strategy: experiment.repeat.strategy, + ...(experiment.repeat.costLimitUsd !== undefined && { + costLimitUsd: experiment.repeat.costLimitUsd, + }), + ...(experiment.earlyExit !== undefined && { earlyExit: experiment.earlyExit }), + }; + } + if (!experiment.runs || experiment.runs <= 1) { + return undefined; + } + return { + count: experiment.runs, + strategy: 'pass_at_k', + ...(experiment.earlyExit !== undefined && { earlyExit: experiment.earlyExit }), + }; +} + +function readExperimentWorkspaceMode(value: unknown): 'pooled' | 'temp' | 'static' | undefined { + return value === 'pooled' || value === 'temp' || value === 'static' ? value : undefined; +} + +function readExperimentWorkspacePath( + workspace: Record | undefined, +): string | undefined { + const value = workspace?.path; + return typeof value === 'string' && value.trim().length > 0 ? value.trim() : undefined; +} + +function normalizeExperimentCaseFilter( + evals: ExperimentConfig['evals'] | undefined, +): NormalizedOptions['filter'] { + if (typeof evals === 'string') { + return evals; + } + if (!Array.isArray(evals)) { + return undefined; + } + return evals.length === 1 ? evals[0] : [...evals]; +} + +async function runExperimentSteps(params: { + readonly label: 'setup' | 'script'; + readonly steps: readonly ExperimentScript[] | undefined; + readonly cwd: string; + readonly experimentConfig?: ExperimentConfig; +}): Promise { + const steps = params.steps ?? []; + if (steps.length === 0) { + return; + } + + for (let index = 0; index < steps.length; index++) { + const step = steps[index]; + const command = buildExperimentStepCommand(step); + const cwd = resolveExperimentStepCwd(params.cwd, params.experimentConfig, step.cwd); + console.log(`Experiment ${params.label} ${index + 1}/${steps.length}: ${command.display}`); + await runExperimentCommand(command.argv, { + cwd, + env: step.env, + timeoutMs: step.timeoutSeconds ? step.timeoutSeconds * 1000 : undefined, + label: `experiment ${params.label}`, + }); + } +} + +async function runExperimentSetup(params: { + readonly config: ExperimentConfig | undefined; + readonly cwd: string; + readonly runDir: string; +}): Promise { + const setup = params.config?.setup; + if (typeof setup === 'function') { + console.log('Experiment setup: running TypeScript setup()'); + await setup({ + cwd: params.cwd, + runDir: params.runDir, + experiment: params.config, + env: process.env, + }); + return; + } + await runExperimentSteps({ + label: 'setup', + steps: setup, + cwd: params.cwd, + experimentConfig: params.config, + }); +} + +function buildExperimentStepCommand(step: ExperimentScript): { + readonly argv: readonly string[]; + readonly display: string; +} { + if (step.command && step.command.length > 0) { + return { argv: step.command, display: step.command.join(' ') }; + } + if (typeof step.script === 'string' && step.script.trim().length > 0) { + return { + argv: shellCommand(step.script), + display: step.script, + }; + } + if (Array.isArray(step.script) && step.script.length > 0) { + return { argv: step.script, display: step.script.join(' ') }; + } + throw new Error('Experiment step must define command or script.'); +} + +function shellCommand(script: string): readonly string[] { + return process.platform === 'win32' ? ['cmd', '/c', script] : ['sh', '-c', script]; +} + +function resolveExperimentStepCwd( + cwd: string, + experimentConfig: ExperimentConfig | undefined, + stepCwd: string | undefined, +): string { + const base = experimentConfig?.sourcePath ? path.dirname(experimentConfig.sourcePath) : cwd; + if (!stepCwd) { + return base; + } + return path.isAbsolute(stepCwd) ? stepCwd : path.resolve(base, stepCwd); +} + +async function runExperimentCommand( + argv: readonly string[], + options: { + readonly cwd: string; + readonly env?: Record; + readonly timeoutMs?: number; + readonly label: string; + }, +): Promise { + if (argv.length === 0) { + throw new Error(`${options.label} command must not be empty.`); + } + + await new Promise((resolve, reject) => { + const cmd = argv[0]; + if (!cmd) { + reject(new Error(`${options.label} command must not be empty.`)); + return; + } + const args = argv.slice(1); + const child = spawn(cmd, args, { + cwd: options.cwd, + env: options.env ? { ...process.env, ...options.env } : process.env, + stdio: 'inherit', + }); + let completed = false; + const timeout = + options.timeoutMs !== undefined + ? setTimeout(() => { + if (!completed) { + completed = true; + child.kill('SIGKILL'); + reject(new Error(`${options.label} timed out after ${options.timeoutMs}ms`)); + } + }, options.timeoutMs) + : undefined; + + child.on('error', (error) => { + if (completed) { + return; + } + completed = true; + if (timeout !== undefined) { + clearTimeout(timeout); + } + reject(error); + }); + child.on('exit', (code) => { + if (completed) { + return; + } + completed = true; + if (timeout !== undefined) { + clearTimeout(timeout); + } + if (code === 0) { + resolve(); + } else { + reject(new Error(`${options.label} exited with code ${code ?? 'unknown'}`)); + } + }); + }); +} + type ProgressReporter = { readonly isInteractive: boolean; start(): void; @@ -699,6 +1029,24 @@ async function prepareFileMetadata(params: { const testIds = suite.tests.map((value) => value.id); const suiteTargets = suite.targets; + if (suite.tests.length === 0) { + return { + testIds, + testCases: suite.tests, + selections: [], + trialsConfig: options.experimentTrialsConfig, + suiteTargets, + yamlWorkers: suite.workers, + yamlCache: suite.cacheConfig?.enabled, + yamlCachePath: suite.cacheConfig?.cachePath, + budgetUsd: suite.budgetUsd, + failOnError: suite.failOnError, + threshold: suite.threshold, + tags: suite.metadata?.tags, + providerFactory: suite.providerFactory, + }; + } + let selections: { selection: TargetSelection; inlineTargetLabel: string }[]; if (options.transcript) { @@ -776,18 +1124,23 @@ async function prepareFileMetadata(params: { const cliTargets = options.cliTargets; const suiteTargets = suite.targets; const suiteTargetRefs = suite.targetRefs; + const experimentTargetRefs = options.experimentTargetRefs; - // Resolve which target names to use (precedence: CLI > suite YAML targets > default) + // Resolve which target names to use (precedence: CLI/experiment > suite YAML targets > default) let targetNames: readonly string[]; + let targetRefs: readonly EvalTargetRef[] | undefined; if (cliTargets.length > 0) { targetNames = cliTargets; + targetRefs = experimentTargetRefs; } else if (suiteTargets && suiteTargets.length > 0) { targetNames = suiteTargets; + targetRefs = suiteTargetRefs; } else { targetNames = []; + targetRefs = undefined; } - if (targetNames.length > 1) { + if (targetNames.length > 1 || (targetNames.length === 1 && targetRefs)) { // Matrix mode: multiple targets const multiSelections = await selectMultipleTargets({ testFilePath, @@ -800,7 +1153,7 @@ async function prepareFileMetadata(params: { dryRunDelayMax: options.dryRunDelayMax, env: process.env, targetNames, - targetRefs: suiteTargetRefs, + targetRefs, }); selections = multiSelections.map((sel) => ({ @@ -823,9 +1176,7 @@ async function prepareFileMetadata(params: { }); // Attach target hooks from eval file if available - const singleTargetHooks = suiteTargetRefs?.find( - (ref) => ref.name === selection.targetName, - )?.hooks; + const singleTargetHooks = targetRefs?.find((ref) => ref.name === selection.targetName)?.hooks; const augmentedSelection: TargetSelection = singleTargetHooks ? { ...selection, targetHooks: singleTargetHooks } : selection; @@ -846,7 +1197,7 @@ async function prepareFileMetadata(params: { testIds, testCases: suite.tests, selections, - trialsConfig: suite.trials, + trialsConfig: options.experimentTrialsConfig, suiteTargets, yamlWorkers: suite.workers, yamlCache: suite.cacheConfig?.enabled, @@ -1167,6 +1518,30 @@ export async function runEvalCommand( } let options = normalizeOptions(input.rawOptions, config, yamlConfig?.execution); + const resolvedExperiment = await resolveExperimentForRun({ + cwd, + explicitExperiment: options.experiment, + yamlDefaultExperiment: resolveDefaultExperimentReference(yamlConfig), + tsDefaultExperiment: normalizeTsDefaultExperiment(config), + }); + options = { + ...applyExperimentOptions(options, resolvedExperiment.config), + experiment: resolvedExperiment.name, + }; + + const evalPathInputs = + input.testFiles.length > 0 + ? [...input.testFiles] + : options.experimentConfig?.evals !== undefined + ? [...(yamlConfig?.eval_patterns ?? DEFAULT_EVAL_PATTERNS)] + : []; + if (evalPathInputs.length === 0 && process.stdin.isTTY) { + const { launchInteractiveWizard } = await import('./interactive.js'); + await launchInteractiveWizard(); + return undefined; + } + const resolvedTestFiles = await resolveEvalPaths(evalPathInputs, cwd); + if (!process.env.AGENTV_EXPERIMENT) { process.env.AGENTV_EXPERIMENT = normalizeExperimentName(options.experiment); } @@ -1382,8 +1757,13 @@ export async function runEvalCommand( console.log(`Artifact directory: ${runDir}`); + await runExperimentSetup({ + config: options.experimentConfig, + cwd, + runDir, + }); + // Log file export paths - const resolvedTestFiles = input.testFiles.map((file) => path.resolve(file)); if (options.otelFile) { console.log(`OTLP JSON file: ${path.resolve(options.otelFile)}`); } @@ -1628,6 +2008,7 @@ export async function runEvalCommand( evalFile, plannedTestCount: totalEvalCount, experiment: normalizeExperimentName(options.experiment), + experimentMetadata: options.experimentMetadata, }); } @@ -1883,7 +2264,11 @@ export async function runEvalCommand( }); const { benchmarkPath: workspaceBenchmarkPath, timingPath } = await aggregateRunDir( runDir, - { evalFile, experiment: normalizeExperimentName(options.experiment) }, + { + evalFile, + experiment: normalizeExperimentName(options.experiment), + experimentMetadata: options.experimentMetadata, + }, ); const indexPath = path.join(runDir, 'index.jsonl'); console.log(`Artifact workspace updated: ${runDir}`); @@ -1900,6 +2285,7 @@ export async function runEvalCommand( } = await writeArtifactsFromResults(allResults, runDir, { evalFile, experiment: normalizeExperimentName(options.experiment), + experimentMetadata: options.experimentMetadata, cwd, repoRoot, sourceTests, @@ -2002,6 +2388,13 @@ export async function runEvalCommand( await wipLoop.stopAndDeleteWipBranch(); } + await runExperimentSteps({ + label: 'script', + steps: options.experimentConfig?.scripts, + cwd, + experimentConfig: options.experimentConfig, + }); + return { executionErrorCount: summary.executionErrorCount, outputPath, diff --git a/apps/cli/src/commands/pipeline/bench.ts b/apps/cli/src/commands/pipeline/bench.ts index 7fe4db49f..f7466f0ed 100644 --- a/apps/cli/src/commands/pipeline/bench.ts +++ b/apps/cli/src/commands/pipeline/bench.ts @@ -132,7 +132,6 @@ export const evalBenchCommand = command({ const grading = { assertions: allAssertions, summary: { passed, failed, total: allAssertions.length, pass_rate: passRate }, - execution_metrics: { tool_calls: {}, total_tool_calls: 0, errors_encountered: 0 }, graders: evaluators.map((e) => ({ name: e.name, type: e.type, diff --git a/apps/cli/src/commands/results/combine-run.ts b/apps/cli/src/commands/results/combine-run.ts index b44c4ed39..18d7bbb99 100644 --- a/apps/cli/src/commands/results/combine-run.ts +++ b/apps/cli/src/commands/results/combine-run.ts @@ -363,6 +363,8 @@ function resolveCombinedExperiment( const MANIFEST_PATH_FIELDS = [ 'artifact_dir', + 'benchmark_path', + 'summary_path', 'grading_path', 'timing_path', 'input_path', diff --git a/apps/cli/src/commands/results/export.ts b/apps/cli/src/commands/results/export.ts index 0df4c7678..bfd4e51df 100644 --- a/apps/cli/src/commands/results/export.ts +++ b/apps/cli/src/commands/results/export.ts @@ -9,9 +9,8 @@ * / * grading.json — per-test grading artifact (assertions, graders) * timing.json — per-test timing artifact - * outputs/ - * response.md — human-readable agent response for this test - * input.md — human-readable input messages for this test + * outputs/answer.md — human-readable agent response for this test + * task/PROMPT.md — human-readable input messages for this test * * This module delegates artifact building to the shared artifact-writer so * that benchmark/grading/timing schemas stay aligned with `agentv eval`. diff --git a/apps/cli/src/commands/results/manifest.ts b/apps/cli/src/commands/results/manifest.ts index c9a272e5f..679cff1ef 100644 --- a/apps/cli/src/commands/results/manifest.ts +++ b/apps/cli/src/commands/results/manifest.ts @@ -39,6 +39,8 @@ export interface ResultManifestRecord { readonly reasoning?: number; }; readonly trace?: Record; + readonly benchmark_path?: string; + readonly summary_path?: string; readonly grading_path?: string; readonly timing_path?: string; readonly input_path?: string; diff --git a/apps/cli/src/commands/results/projection-bundle.ts b/apps/cli/src/commands/results/projection-bundle.ts index a81b96c15..e6b69c6d5 100644 --- a/apps/cli/src/commands/results/projection-bundle.ts +++ b/apps/cli/src/commands/results/projection-bundle.ts @@ -87,12 +87,13 @@ export type ProjectionBundleArtifactRefs = Partial< Pick< IndexArtifactEntry, | 'artifact_dir' + | 'benchmark_path' + | 'summary_path' | 'grading_path' | 'timing_path' | 'input_path' | 'output_path' | 'answer_path' - | 'response_path' | 'transcript_path' | 'metrics_path' | 'task_dir' @@ -172,11 +173,12 @@ function artifactRefs( return dropUndefined({ ...metadataRefs, artifact_dir: indexEntry.artifact_dir, + benchmark_path: indexEntry.benchmark_path, + summary_path: indexEntry.summary_path, grading_path: indexEntry.grading_path, input_path: indexEntry.input_path, output_path: indexEntry.output_path, answer_path: indexEntry.answer_path, - response_path: indexEntry.response_path, transcript_path: indexEntry.transcript_path, metrics_path: indexEntry.metrics_path, trace_path: tracePathFor(indexEntry), @@ -280,7 +282,6 @@ function buildEntry( artifacts: { trace_path: tracePathFor(indexRecord ?? plannedIndexEntry), answer_path: result.output.length > 0 ? 'outputs/answer.md' : undefined, - response_path: result.output.length > 0 ? 'outputs/response.md' : undefined, }, duplicatePolicy: options.duplicatePolicy, capture: captureOptions(includeRawContent), diff --git a/apps/cli/src/commands/results/serve.ts b/apps/cli/src/commands/results/serve.ts index 1029475d2..ef7766b93 100644 --- a/apps/cli/src/commands/results/serve.ts +++ b/apps/cli/src/commands/results/serve.ts @@ -742,6 +742,8 @@ function buildResultArtifactCatalog( addPointerArtifactCatalogEntry(entries, seen, trace, 'trace', options?.runPath); addPointerArtifactCatalogEntry(entries, seen, answer, 'answer', options?.runPath); + addDirectArtifactCatalogEntry(entries, seen, record.benchmark_path, 'artifact'); + addDirectArtifactCatalogEntry(entries, seen, record.summary_path, 'artifact'); addDirectArtifactCatalogEntry(entries, seen, record.grading_path, 'artifact'); addDirectArtifactCatalogEntry(entries, seen, record.timing_path, 'artifact'); addDirectArtifactCatalogEntry(entries, seen, record.input_path, 'artifact'); diff --git a/apps/cli/src/commands/results/validate.ts b/apps/cli/src/commands/results/validate.ts index c0ffb8dc0..680f5b0b4 100644 --- a/apps/cli/src/commands/results/validate.ts +++ b/apps/cli/src/commands/results/validate.ts @@ -6,7 +6,7 @@ * 1. Directory follows the `.agentv/results//` naming convention * 2. index.jsonl exists and each line has required fields * 3. Per-test grading.json exists for every entry in the index - * 4. Per-test timing.json exists (warning if missing) + * 4. Per-test timing.json exists for direct case rows (warning if missing) * 5. benchmark.json exists (warning if missing) * 6. Scores are within [0, 1] * 7. index.jsonl entries have `scores[]` array (warning if missing — dashboard needs it) @@ -34,6 +34,8 @@ interface IndexEntry { readonly target?: string; readonly scores?: unknown[]; readonly execution_status?: string; + readonly benchmark_path?: string; + readonly summary_path?: string; readonly grading_path?: string; readonly timing_path?: string; readonly [key: string]: unknown; @@ -139,10 +141,10 @@ function checkIndexJsonl(runDir: string): { diagnostics: Diagnostic[]; entries: }); } - if (!entry.grading_path) { + if (!entry.grading_path && !entry.benchmark_path) { diagnostics.push({ severity: 'warning', - message: `index.jsonl line ${i + 1} (${entry.test_id ?? '?'}): missing 'grading_path'`, + message: `index.jsonl line ${i + 1} (${entry.test_id ?? '?'}): missing 'grading_path' or 'benchmark_path'`, }); } @@ -203,6 +205,26 @@ function checkArtifactFiles(runDir: string, entries: IndexEntry[]): Diagnostic[] for (const entry of entries) { const testId = entry.test_id ?? '?'; + if (entry.summary_path) { + const summaryPath = path.join(runDir, entry.summary_path); + if (!existsSync(summaryPath)) { + diagnostics.push({ + severity: 'error', + message: `${testId}: summary.json not found at '${entry.summary_path}'`, + }); + } + } + + if (entry.benchmark_path) { + const benchmarkPath = path.join(runDir, entry.benchmark_path); + if (!existsSync(benchmarkPath)) { + diagnostics.push({ + severity: 'error', + message: `${testId}: benchmark.json not found at '${entry.benchmark_path}'`, + }); + } + } + // Check grading.json if (entry.grading_path) { const gradingPath = path.join(runDir, entry.grading_path); diff --git a/apps/cli/test/commands/eval/aggregate.test.ts b/apps/cli/test/commands/eval/aggregate.test.ts index 91200aa61..40b00d9a1 100644 --- a/apps/cli/test/commands/eval/aggregate.test.ts +++ b/apps/cli/test/commands/eval/aggregate.test.ts @@ -191,12 +191,12 @@ describe('writePerTestArtifacts', () => { expect(grading2.assertions).toHaveLength(1); }); - it('writes response.md for results with output', async () => { + it('writes outputs/answer.md for results with output', async () => { const results = [makeResult({ testId: 'test-1', output: 'hello' })]; await writePerTestArtifacts(results, tmpDir); - const response = readFileSync(path.join(tmpDir, 'test-1', 'outputs', 'response.md'), 'utf8'); - expect(response).toContain('hello'); + const answer = readFileSync(path.join(tmpDir, 'test-1', 'outputs', 'answer.md'), 'utf8'); + expect(answer).toContain('hello'); }); }); diff --git a/apps/cli/test/commands/eval/artifact-writer.test.ts b/apps/cli/test/commands/eval/artifact-writer.test.ts index 1d05c3e6f..3e2feac68 100644 --- a/apps/cli/test/commands/eval/artifact-writer.test.ts +++ b/apps/cli/test/commands/eval/artifact-writer.test.ts @@ -142,6 +142,56 @@ describe('buildGradingArtifact', () => { }); }); + it('preserves repeat trial metadata', () => { + const result = makeResult({ + trials: [ + { + attempt: 0, + score: 0.4, + verdict: 'fail', + executionStatus: 'quality_failure', + failureStage: 'evaluator', + failureReasonCode: 'threshold_not_met', + }, + { + attempt: 1, + score: 1, + verdict: 'pass', + costUsd: 0.03, + }, + ], + aggregation: { + strategy: 'pass_at_k', + passedAttempts: 1, + totalAttempts: 2, + }, + }); + + const grading = buildGradingArtifact(result); + + expect(grading.trials).toEqual([ + { + attempt: 0, + score: 0.4, + verdict: 'fail', + execution_status: 'quality_failure', + failure_stage: 'evaluator', + failure_reason_code: 'threshold_not_met', + }, + { + attempt: 1, + score: 1, + verdict: 'pass', + cost_usd: 0.03, + }, + ]); + expect(grading.aggregation).toEqual({ + strategy: 'pass_at_k', + passed_attempts: 1, + total_attempts: 2, + }); + }); + it('uses top-level assertions when no grader scores', () => { const result = makeResult({ assertions: [ @@ -176,10 +226,10 @@ describe('buildGradingArtifact', () => { expect(grading.graders?.[1].score).toBe(0.7); }); - it('records error as errors_encountered', () => { + it('keeps grading.json focused on grading evidence', () => { const result = makeResult({ error: 'Timeout exceeded' }); const grading = buildGradingArtifact(result); - expect(grading.execution_metrics.errors_encountered).toBe(1); + expect(grading).not.toHaveProperty('execution_metrics'); }); it('handles result with no assertions or scores', () => { @@ -528,8 +578,9 @@ describe('buildIndexArtifactEntry', () => { outputDir: '/tmp/artifacts', gradingPath: '/tmp/artifacts/alpha/grading.json', timingPath: '/tmp/artifacts/alpha/timing.json', - outputPath: '/tmp/artifacts/alpha/outputs/response.md', - inputPath: '/tmp/artifacts/alpha/input.md', + outputPath: '/tmp/artifacts/alpha/outputs/answer.md', + answerPath: '/tmp/artifacts/alpha/outputs/answer.md', + inputPath: '/tmp/artifacts/alpha/task/PROMPT.md', }, ); @@ -559,8 +610,43 @@ describe('buildIndexArtifactEntry', () => { error: 'model drift', grading_path: 'alpha/grading.json', timing_path: 'alpha/timing.json', - output_path: 'alpha/outputs/response.md', - input_path: 'alpha/input.md', + output_path: 'alpha/outputs/answer.md', + answer_path: 'alpha/outputs/answer.md', + input_path: 'alpha/task/PROMPT.md', + }); + }); + + it('includes repeat trial metadata', () => { + const entry = buildIndexArtifactEntry( + makeResult({ + testId: 'alpha', + trials: [ + { attempt: 0, score: 0.8, verdict: 'pass' }, + { attempt: 1, score: 0.6, verdict: 'fail', error: 'missing token' }, + ], + aggregation: { + strategy: 'mean', + mean: 0.7, + min: 0.6, + max: 0.8, + }, + }), + { + outputDir: '/tmp/artifacts', + gradingPath: '/tmp/artifacts/alpha/grading.json', + timingPath: '/tmp/artifacts/alpha/timing.json', + }, + ); + + expect(entry.trials).toEqual([ + { attempt: 0, score: 0.8, verdict: 'pass' }, + { attempt: 1, score: 0.6, verdict: 'fail', error: 'missing token' }, + ]); + expect(entry.aggregation).toEqual({ + strategy: 'mean', + mean: 0.7, + min: 0.6, + max: 0.8, }); }); }); @@ -791,7 +877,7 @@ describe('writeArtifactsFromResults', () => { await readFile(path.join(paths.testArtifactDir, 'alpha', 'grading.json'), 'utf8'), ); expect(alphaGrading.summary).toBeDefined(); - expect(alphaGrading.execution_metrics).toBeDefined(); + expect(alphaGrading).not.toHaveProperty('execution_metrics'); const alphaTiming: TimingArtifact = JSON.parse( await readFile(path.join(paths.testArtifactDir, 'alpha', 'timing.json'), 'utf8'), @@ -819,6 +905,191 @@ describe('writeArtifactsFromResults', () => { expect(indexLines[0]?.metrics_path).toBe('alpha/metrics.json'); }); + it('writes repeat runs in Vercel-compatible case and run folders', async () => { + const results = [ + makeResult({ + testId: 'repeat-case', + score: 1, + trials: [ + { + attempt: 0, + score: 0.25, + verdict: 'fail', + result: makeResult({ + testId: 'repeat-case', + score: 0.25, + output: 'first attempt', + durationMs: 2000, + executionStatus: 'quality_failure', + }), + }, + { + attempt: 1, + score: 1, + verdict: 'pass', + result: makeResult({ + testId: 'repeat-case', + score: 1, + output: 'second attempt', + durationMs: 4000, + }), + }, + ], + aggregation: { + strategy: 'confidence_interval', + mean: 0.625, + ci95Lower: 0.1, + ci95Upper: 1, + stddev: 0.53, + }, + }), + ]; + + const sourceTests = [ + { + id: 'repeat-case', + input: [{ role: 'user', content: 'Repeat this task prompt.' }], + expected_output: [], + reference_answer: '', + file_paths: [], + criteria: 'Repeats the task prompt', + evaluator: 'llm-grader', + assertions: [], + } as unknown as EvalTest, + ]; + + const paths = await writeArtifactsFromResults(results, testDir, { sourceTests }); + + const [indexEntry] = (await readFile(paths.indexPath, 'utf8')) + .trim() + .split('\n') + .map((line) => JSON.parse(line) as IndexArtifactEntry); + expect(indexEntry?.trials).toEqual([ + { attempt: 0, run_path: 'run-1', score: 0.25, verdict: 'fail' }, + { attempt: 1, run_path: 'run-2', score: 1, verdict: 'pass' }, + ]); + expect(indexEntry?.aggregation).toEqual({ + strategy: 'confidence_interval', + mean: 0.625, + ci95_lower: 0.1, + ci95_upper: 1, + stddev: 0.53, + }); + expect(indexEntry?.artifact_dir).toBe('repeat-case'); + expect(indexEntry?.summary_path).toBe('repeat-case/summary.json'); + expect(indexEntry?.task_dir).toBe('repeat-case/task'); + expect(indexEntry?.input_path).toBe('repeat-case/task/PROMPT.md'); + expect(indexEntry?.benchmark_path).toBeUndefined(); + expect(indexEntry?.grading_path).toBe('repeat-case/grading.json'); + expect(indexEntry?.timing_path).toBeUndefined(); + expect(indexEntry?.metrics_path).toBeUndefined(); + + const repeatEntries = await readdir(path.join(paths.testArtifactDir, 'repeat-case')); + expect(repeatEntries.sort()).toEqual([ + 'grading.json', + 'run-1', + 'run-2', + 'summary.json', + 'task', + ]); + + const prompt = await readFile( + path.join(paths.testArtifactDir, 'repeat-case', 'task', 'PROMPT.md'), + 'utf8', + ); + expect(prompt).toBe('@[user]:\nRepeat this task prompt.'); + + const caseSummary = JSON.parse( + await readFile(path.join(paths.testArtifactDir, 'repeat-case', 'summary.json'), 'utf8'), + ) as Record; + expect(caseSummary).toMatchObject({ + total_runs: 2, + passed_runs: 1, + pass_rate: '50%', + mean_duration_ms: 3000, + mean_duration_seconds: 3, + duration_ms: 6000, + total_duration_seconds: 6, + duration_stats: { + count: 2, + mean_ms: 3000, + mean_seconds: 3, + stddev_ms: 1000, + stddev_seconds: 1, + min_ms: 2000, + max_ms: 4000, + }, + total_tokens: 0, + cost_usd: null, + token_usage: { input: 0, output: 0, reasoning: 0 }, + usage_sources: { + token_usage: 'unavailable', + total_tokens: 'unavailable', + duration: 'aggregate', + cost: 'unavailable', + }, + }); + expect(typeof caseSummary.fingerprint).toBe('string'); + + const aggregateGrading: GradingArtifact = JSON.parse( + await readFile(path.join(paths.testArtifactDir, 'repeat-case', 'grading.json'), 'utf8'), + ); + expect(aggregateGrading.trials).toEqual(indexEntry?.trials); + expect(aggregateGrading.aggregation).toEqual(indexEntry?.aggregation); + + for (const runDir of ['run-1', 'run-2']) { + const runEntries = await readdir(path.join(paths.testArtifactDir, 'repeat-case', runDir)); + expect(runEntries.sort()).toEqual([ + 'grading.json', + 'outputs', + 'result.json', + 'transcript-raw.jsonl', + 'transcript.json', + ]); + } + + const runOneResult = JSON.parse( + await readFile( + path.join(paths.testArtifactDir, 'repeat-case', 'run-1', 'result.json'), + 'utf8', + ), + ) as Record; + expect(runOneResult).toMatchObject({ + status: 'failed', + duration_ms: 2000, + duration_seconds: 2, + model: 'test-target', + grading_path: './grading.json', + transcript_path: './transcript.json', + transcript_raw_path: './transcript-raw.jsonl', + output_paths: { answer: './outputs/answer.md' }, + timing: { + duration_ms: 2000, + }, + }); + + const runTwoAnswer = await readFile( + path.join(paths.testArtifactDir, 'repeat-case', 'run-2', 'outputs', 'answer.md'), + 'utf8', + ); + expect(runTwoAnswer).toBe('second attempt'); + + const runTwoResult = JSON.parse( + await readFile( + path.join(paths.testArtifactDir, 'repeat-case', 'run-2', 'result.json'), + 'utf8', + ), + ) as Record; + expect(runTwoResult).toMatchObject({ + grading_path: './grading.json', + transcript_path: './transcript.json', + transcript_raw_path: './transcript-raw.jsonl', + timing: { + duration_ms: 4000, + }, + }); + }); + it('handles empty results array', async () => { const paths = await writeArtifactsFromResults([], testDir); diff --git a/apps/cli/test/commands/results/export-e2e-providers.test.ts b/apps/cli/test/commands/results/export-e2e-providers.test.ts index 19c0e4be4..cb1b2eee6 100644 --- a/apps/cli/test/commands/results/export-e2e-providers.test.ts +++ b/apps/cli/test/commands/results/export-e2e-providers.test.ts @@ -466,10 +466,12 @@ describe('export e2e — multi-provider metrics verification', () => { expect(grading.summary.failed).toBe(0); expect(grading.summary.pass_rate).toBe(1.0); - // Tool calls from trace - expect(grading.execution_metrics.total_tool_calls).toBe(3); - expect(grading.execution_metrics.tool_calls.Read).toBe(2); - expect(grading.execution_metrics.tool_calls.Write).toBe(1); + const metrics = JSON.parse( + readFileSync(path.join(artifactDir(outputDir, CLAUDE_CLI_RESULT), 'metrics.json'), 'utf8'), + ); + expect(metrics.metrics.total_tool_calls).toBe(3); + expect(metrics.metrics.tool_call_counts.Read).toBe(2); + expect(metrics.metrics.tool_call_counts.Write).toBe(1); // Graders expect(grading.graders).toHaveLength(1); @@ -490,8 +492,10 @@ describe('export e2e — multi-provider metrics verification', () => { expect(grading.summary.failed).toBe(1); expect(grading.summary.pass_rate).toBe(0.5); - // No trace means no tool calls - expect(grading.execution_metrics.total_tool_calls).toBe(0); + const metrics = JSON.parse( + readFileSync(path.join(artifactDir(outputDir, COPILOT_RESULT), 'metrics.json'), 'utf8'), + ); + expect(metrics.metrics.total_tool_calls).toBe(0); }); it('should handle error result in grading', async () => { @@ -507,7 +511,10 @@ describe('export e2e — multi-provider metrics verification', () => { // Error result has empty assertions expect(grading.summary.total).toBe(0); expect(grading.summary.pass_rate).toBe(0); - expect(grading.execution_metrics.errors_encountered).toBe(1); + const metrics = JSON.parse( + readFileSync(path.join(artifactDir(outputDir, ERROR_RESULT), 'metrics.json'), 'utf8'), + ); + expect(metrics.metrics.errors_encountered).toBe(1); }); it('should produce grading files for all test IDs in multi-target run', async () => { diff --git a/apps/cli/test/commands/results/export.test.ts b/apps/cli/test/commands/results/export.test.ts index 2056c16d7..9bb6b283a 100644 --- a/apps/cli/test/commands/results/export.test.ts +++ b/apps/cli/test/commands/results/export.test.ts @@ -349,10 +349,9 @@ describe('results export', () => { }); expect(bundle.entries[0].artifact_refs).toMatchObject({ status: 'planned_export', - input_path: 'privacy/test-private/input.md', + input_path: 'privacy/test-private/task/PROMPT.md', output_path: 'privacy/test-private/outputs/answer.md', answer_path: 'privacy/test-private/outputs/answer.md', - response_path: 'privacy/test-private/outputs/response.md', trace_path: 'privacy/test-private/trace.json', }); expect(bundle.entries[0].trace.envelope_ref).toBe('privacy/test-private/trace.json'); @@ -417,9 +416,8 @@ describe('results export', () => { timing_path: 'demo/test-greeting/timing.json', output_path: 'demo/test-greeting/outputs/answer.md', answer_path: 'demo/test-greeting/outputs/answer.md', - response_path: 'demo/test-greeting/outputs/response.md', transcript_path: 'demo/test-greeting/transcript.jsonl', - input_path: 'demo/test-greeting/input.md', + input_path: 'demo/test-greeting/task/PROMPT.md', }); expect(entries[0].projection_identity).toMatchObject({ schema_version: 'agentv.projection_identity.v1', @@ -589,8 +587,8 @@ describe('results export', () => { expect(grading.summary).toHaveProperty('total'); expect(grading.summary).toHaveProperty('pass_rate'); - // Has execution_metrics - expect(grading.execution_metrics).toBeDefined(); + // Grading artifacts stay focused on assertion evidence; execution data lives in metrics.json. + expect(grading).not.toHaveProperty('execution_metrics'); // Has evaluators expect(grading.graders).toBeDefined(); @@ -613,8 +611,7 @@ describe('results export', () => { expect(readFileSync(answerPath, 'utf8')).toBe('Hello, Alice!'); const responsePath = path.join(artifactDir(outputDir, RESULT_FULL), 'outputs', 'response.md'); - expect(existsSync(responsePath)).toBe(true); - expect(readFileSync(responsePath, 'utf8')).toBe('Hello, Alice!'); + expect(existsSync(responsePath)).toBe(false); }); it('should group results by target in benchmark.json', async () => { @@ -715,7 +712,7 @@ describe('results export', () => { expect(grading.summary.total).toBe(0); }); - it('should write string input to /input.md', async () => { + it('should write string input to /task/PROMPT.md', async () => { const outputDir = path.join(tempDir, 'output'); const resultWithInput = { ...RESULT_FULL, @@ -725,12 +722,12 @@ describe('results export', () => { await exportResults('test.jsonl', content, outputDir); - const inputPath = path.join(artifactDir(outputDir, resultWithInput), 'input.md'); + const inputPath = path.join(artifactDir(outputDir, resultWithInput), 'task', 'PROMPT.md'); expect(existsSync(inputPath)).toBe(true); expect(readFileSync(inputPath, 'utf8')).toBe('What is the capital of France?'); }); - it('should write Message[] input to /input.md as markdown', async () => { + it('should write Message[] input to /task/PROMPT.md as markdown', async () => { const outputDir = path.join(tempDir, 'output'); const resultWithMessages = { ...RESULT_FULL, @@ -743,7 +740,7 @@ describe('results export', () => { await exportResults('test.jsonl', content, outputDir); - const inputPath = path.join(artifactDir(outputDir, resultWithMessages), 'input.md'); + const inputPath = path.join(artifactDir(outputDir, resultWithMessages), 'task', 'PROMPT.md'); expect(existsSync(inputPath)).toBe(true); expect(readFileSync(inputPath, 'utf8')).toBe('@[user]:\nHello\n\n@[assistant]:\nHi there!'); }); @@ -754,7 +751,7 @@ describe('results export', () => { await exportResults('test.jsonl', content, outputDir); - const inputPath = path.join(artifactDir(outputDir, RESULT_FULL), 'input.md'); + const inputPath = path.join(artifactDir(outputDir, RESULT_FULL), 'task', 'PROMPT.md'); expect(existsSync(inputPath)).toBe(false); }); diff --git a/apps/cli/test/commands/results/serve.test.ts b/apps/cli/test/commands/results/serve.test.ts index 439b10ab2..8a6c5ef62 100644 --- a/apps/cli/test/commands/results/serve.test.ts +++ b/apps/cli/test/commands/results/serve.test.ts @@ -3193,7 +3193,7 @@ describe('serve app', () => { const runId = 'with-transcript::2026-03-25T10-00-00-000Z'; const timestampDir = path.join(runsDir, '2026-03-25T10-00-00-000Z'); const transcriptArtifactPath = 'demo/test-greeting/transcript.jsonl'; - const answerArtifactPath = 'demo/test-greeting/outputs/answer.md'; + const answerArtifactPath = 'demo/test-greeting/answer.md'; const transcriptPath = path.join(timestampDir, transcriptArtifactPath); const answerPath = path.join(timestampDir, answerArtifactPath); const transcriptJsonl = `${JSON.stringify({ @@ -3558,7 +3558,7 @@ describe('serve app', () => { const runId = 'escaped-answer::2026-03-25T13-45-00-000Z'; const timestampDir = path.join(runsDir, '2026-03-25T13-45-00-000Z'); const transcriptArtifactPath = 'demo/test-greeting/transcript.jsonl'; - const answerArtifactPath = 'demo/test-greeting/outputs/answer.md'; + const answerArtifactPath = 'demo/test-greeting/answer.md'; const transcriptPath = path.join(timestampDir, transcriptArtifactPath); const answerPath = path.join(timestampDir, answerArtifactPath); const transcriptJsonl = `${JSON.stringify({ diff --git a/apps/cli/test/commands/runs/rerun.test.ts b/apps/cli/test/commands/runs/rerun.test.ts index 9a6e39292..1ca5bdae6 100644 --- a/apps/cli/test/commands/runs/rerun.test.ts +++ b/apps/cli/test/commands/runs/rerun.test.ts @@ -61,7 +61,7 @@ tests: await writeFile(path.join(taskDir, 'targets.yaml'), options.targetsYaml, 'utf8'); await writeFile(path.join(artifactDir, 'grading.json'), '{"assertions":[]}\n', 'utf8'); await writeFile(path.join(artifactDir, 'timing.json'), '{"duration_ms":1}\n', 'utf8'); - await writeFile(path.join(outputsDir, 'response.md'), '@[assistant]:\nCaptured answer\n', 'utf8'); + await writeFile(path.join(outputsDir, 'answer.md'), '@[assistant]:\nCaptured answer\n', 'utf8'); return { timestamp: '2024-01-01T00:00:00.000Z', @@ -71,8 +71,8 @@ tests: artifact_dir: options.testId, grading_path: `${options.testId}/grading.json`, timing_path: `${options.testId}/timing.json`, - output_path: `${options.testId}/outputs/response.md`, - response_path: `${options.testId}/outputs/response.md`, + output_path: `${options.testId}/outputs/answer.md`, + answer_path: `${options.testId}/outputs/answer.md`, task_dir: `${options.testId}/task`, eval_path: `${options.testId}/task/EVAL.yaml`, targets_path: `${options.testId}/task/targets.yaml`, @@ -197,10 +197,10 @@ describe('agentv runs rerun', () => { }, }); - const responsePath = path.join(created.outputDir, String(rows[0].response_path)); - const response = await readFile(responsePath, 'utf8'); - expect(response).toContain('Alpha answer'); - expect(response).not.toContain('Captured answer'); + const answerPath = path.join(created.outputDir, String(rows[0].answer_path)); + const answer = await readFile(answerPath, 'utf8'); + expect(answer).toContain('Alpha answer'); + expect(answer).not.toContain('Captured answer'); }, 30_000); it('fails clearly for missing env and accepts an explicit env file', async () => { diff --git a/apps/cli/test/eval.integration.test.ts b/apps/cli/test/eval.integration.test.ts index 29830509b..083819931 100644 --- a/apps/cli/test/eval.integration.test.ts +++ b/apps/cli/test/eval.integration.test.ts @@ -518,6 +518,105 @@ describe('agentv eval CLI', () => { } }, 30_000); + it('runs a native experiment file with eval selection and run knobs', async () => { + const fixture = await createFixture(); + try { + const experimentsDir = path.join(fixture.suiteDir, 'experiments'); + await mkdir(experimentsDir, { recursive: true }); + await writeFile( + path.join(fixture.suiteDir, '.agentv', 'config.yaml'), + 'eval_patterns:\n - sample.test.yaml\n - unused.test.yaml\n', + 'utf8', + ); + await writeFile( + path.join(fixture.suiteDir, 'unused.test.yaml'), + [ + 'description: unmatched eval file should not resolve targets', + 'target: missing-target', + 'tests:', + ' - id: case-unused', + ' criteria: System responds with unused', + ' input: unused', + ' expected_output: unused', + '', + ].join('\n'), + 'utf8', + ); + const experimentPath = path.join(experimentsDir, 'default.yaml'); + await writeFile( + experimentPath, + [ + 'name: native-exp', + 'target: cli-target', + 'evals: case-alpha', + 'timeout_seconds: 12', + 'workers: 4', + 'repeat:', + ' count: 2', + ' strategy: mean', + ' cost_limit_usd: 1.25', + 'early_exit: false', + 'setup:', + ' - script: "printf setup > ../experiment-setup.txt"', + 'scripts:', + ' - script: "printf script > ../experiment-script.txt"', + '', + ].join('\n'), + 'utf8', + ); + + const { stdout, exitCode } = await runCli(fixture, [ + 'eval', + '--experiment', + 'experiments/default.yaml', + ]); + + expect(exitCode).toBe(0); + const outputPath = extractOutputPath(stdout); + expect(outputPath).toContain(`${path.sep}native-exp${path.sep}`); + + const diagnostics = await readDiagnostics(fixture); + expect(diagnostics).toMatchObject({ + target: 'cli-target', + agentTimeoutMs: 12000, + maxConcurrency: 4, + trials: { + count: 2, + strategy: 'mean', + costLimitUsd: 1.25, + earlyExit: false, + }, + }); + + await expectFileExists(path.join(fixture.suiteDir, 'experiment-setup.txt')); + await expectFileExists(path.join(fixture.suiteDir, 'experiment-script.txt')); + + const benchmark = JSON.parse( + await readFile(path.join(path.dirname(outputPath), 'benchmark.json'), 'utf8'), + ) as { metadata?: Record }; + expect(benchmark.metadata?.experiment).toBe('native-exp'); + expect(benchmark.metadata?.experiment_config).toMatchObject({ + name: 'native-exp', + source_path: experimentPath, + target: 'cli-target', + evals: 'case-alpha', + repeat: { + count: 2, + strategy: 'mean', + cost_limit_usd: 1.25, + }, + early_exit: false, + timeout_seconds: 12, + workers: 4, + }); + expect( + (benchmark.metadata?.experiment_config as Record).fingerprint, + ).toMatch(/^[a-f0-9]{64}$/); + } finally { + await rm(fixture.baseDir, { recursive: true, force: true }); + } + }, 30_000); + it('honors agentv.config.ts cache.path when response cache is enabled there', async () => { const fixture = await createFixture(); try { diff --git a/apps/cli/test/fixtures/mock-run-evaluation.ts b/apps/cli/test/fixtures/mock-run-evaluation.ts index c0957ca5d..32162888e 100644 --- a/apps/cli/test/fixtures/mock-run-evaluation.ts +++ b/apps/cli/test/fixtures/mock-run-evaluation.ts @@ -18,6 +18,13 @@ interface RunEvaluationOptionsLike { readonly filter?: string | readonly string[]; readonly evalCases?: ReadonlyArray; readonly verbose?: boolean; + readonly maxConcurrency?: number; + readonly trials?: { + readonly count: number; + readonly strategy: string; + readonly costLimitUsd?: number; + readonly earlyExit?: boolean; + }; readonly budgetUsd?: number; readonly runBudgetTracker?: { readonly budgetCapUsd?: number; @@ -172,6 +179,8 @@ async function maybeWriteDiagnostics( envRootOnly: process.env.CLI_ENV_ROOT_ONLY ?? null, envLocalOnly: process.env.CLI_ENV_LOCAL_ONLY ?? null, budgetUsd: options.budgetUsd ?? null, + maxConcurrency: options.maxConcurrency ?? null, + trials: options.trials ?? null, hasRunBudgetTracker: options.runBudgetTracker !== undefined, runBudgetCapUsd: options.runBudgetTracker?.budgetCapUsd ?? null, replayRecording: options.replayRecording ?? null, diff --git a/apps/dashboard/src/components/__fixtures__/structured-transcript.ts b/apps/dashboard/src/components/__fixtures__/structured-transcript.ts index 93f37e21b..0a51654de 100644 --- a/apps/dashboard/src/components/__fixtures__/structured-transcript.ts +++ b/apps/dashboard/src/components/__fixtures__/structured-transcript.ts @@ -86,7 +86,7 @@ export const structuredTranscriptFiles: FileNode[] = [ type: 'dir', children: [ { name: 'grading.json', path: 'final-json-answer__codex/grading.json', type: 'file' }, - { name: 'answer.md', path: 'final-json-answer__codex/answer.md', type: 'file' }, + { name: 'answer.md', path: 'final-json-answer__codex/outputs/answer.md', type: 'file' }, { name: 'outputs', path: 'final-json-answer__codex/outputs', diff --git a/apps/web/src/content/docs/docs/evaluation/eval-files.mdx b/apps/web/src/content/docs/docs/evaluation/eval-files.mdx index 4b123162b..e4d932da0 100644 --- a/apps/web/src/content/docs/docs/evaluation/eval-files.mdx +++ b/apps/web/src/content/docs/docs/evaluation/eval-files.mdx @@ -5,13 +5,13 @@ sidebar: order: 1 --- -Evaluation files define the test cases, targets, and graders for an evaluation run. AgentV supports two formats: YAML and JSONL. +Evaluation files define the test cases and graders for an evaluation run. Runtime choices such as target matrices, setup, scripts, and repeat runs belong in [experiments](/docs/evaluation/experiments/). AgentV supports two eval formats: YAML and JSONL. YAML is the canonical portable model. TypeScript helpers, generated fixtures, and Python scripts should lower to the same YAML/JSONL shapes rather than inventing a separate eval contract. ## Suites -An eval file is a **suite**: it binds test cases to execution context (workspace, hooks, targets, trials). Test cases can be inline or loaded from an external file via `tests: ./cases.yaml` for reuse across suites. +An eval file is a **suite**: it binds test cases to task context, assertions, and reusable fixtures. Runtime choices such as target matrices, setup, and run counts belong in experiments. Test cases can be inline or loaded from an external file via `tests: ./cases.yaml` for reuse across suites. ## YAML Format @@ -186,6 +186,35 @@ Each test's effective input becomes a single user message with `[file blocks..., Per-test `input_files` overrides the suite-level value (it does not merge). To opt out, set `execution.skip_defaults: true` on the test. +### PROMPT.md Fallback + +For Vercel-style eval directories, a test may omit `input` and keep the task +prompt in Markdown instead. AgentV resolves the prompt in this order: + +1. If the effective `input_files` contains a file named exactly `PROMPT.md`, that file becomes the test prompt. +2. Otherwise, if a `PROMPT.md` exists beside the `EVAL.yaml`, that file becomes the test prompt. +3. Other `input_files` remain attachments. `PROMPT.md` is removed from the attachment list so the prompt is not duplicated. + +```text +agent-001-fix-bug/ + EVAL.yaml + PROMPT.md + fixtures/ + failing-test.log +``` + +```yaml +tests: + - id: fix-bug + criteria: Fixes the regression described in the prompt + input_files: + - ./fixtures/failing-test.log +``` + +Use explicit `input` when the prompt is short or generated from YAML variables. +Use `PROMPT.md` when the task text is long enough that duplicating it inside +YAML would make the eval hard to review. + ### Tests as String Path Instead of inlining tests in the same file, you can point `tests` to an external YAML or JSONL file. This is the inverse of the sidecar pattern — the metadata file references the test data: diff --git a/apps/web/src/content/docs/docs/evaluation/examples.mdx b/apps/web/src/content/docs/docs/evaluation/examples.mdx index 11a29c8fb..7a848930e 100644 --- a/apps/web/src/content/docs/docs/evaluation/examples.mdx +++ b/apps/web/src/content/docs/docs/evaluation/examples.mdx @@ -422,5 +422,5 @@ See the [suite-level-input example](https://github.com/EntityProcess/agentv/tree For complete end-to-end workflows that combine multiple features, see the showcases in [`examples/showcase/`](https://github.com/EntityProcess/agentv/tree/main/examples/showcase): -- **[Multi-Model Benchmark](https://github.com/EntityProcess/agentv/tree/main/examples/showcase/multi-model-benchmark)** — targets matrix × weighted metrics × trials × compare workflow. Runs the same tests against multiple models, scores with weighted graders, measures variability, and compares results side-by-side. +- **[Multi-Model Benchmark](https://github.com/EntityProcess/agentv/tree/main/examples/showcase/multi-model-benchmark)** — experiment target matrix × weighted metrics × repeated runs × compare workflow. Runs the same tests against multiple models, scores with weighted graders, measures variability, and compares results side-by-side. - **[Export Screening](https://github.com/EntityProcess/agentv/tree/main/examples/showcase/export-screening)** — classification eval with confusion matrix metrics and CI gating. diff --git a/apps/web/src/content/docs/docs/evaluation/experiments.mdx b/apps/web/src/content/docs/docs/evaluation/experiments.mdx new file mode 100644 index 000000000..5e6f6a63b --- /dev/null +++ b/apps/web/src/content/docs/docs/evaluation/experiments.mdx @@ -0,0 +1,147 @@ +--- +title: Experiments +description: Configure how AgentV evals run +sidebar: + order: 2 +--- + +Experiments define **how** eval cases run: target or target matrix, setup, +scripts, timeout, sandbox, case filters, and repeat-run policy. Eval files stay +focused on **what** is tested: prompts, datasets, assertions, and task fixtures. + +## Experiment YAML + +Committed experiments conventionally live under `experiments/`: + +```yaml +name: baseline +target: codex-gpt5 +evals: "agent-*" +timeout_seconds: 720 +repeat: + count: 4 + strategy: pass_at_k + cost_limit_usd: 2.00 +setup: + - script: bun install +scripts: + - build +``` + +Wire fields use `snake_case`. AgentV translates to internal `camelCase` when it +loads the file. + +## Repeat runs + +`repeat` is the full AgentV replacement for the old eval-level +`execution.trials` shape. It supports the same core strategies: + +```yaml +repeat: + count: 3 + strategy: mean + cost_limit_usd: 1.50 +``` + +Supported strategies: + +| Strategy | Behavior | +| --- | --- | +| `pass_at_k` | Uses the best passing attempt; early-exits by default unless the experiment sets `early_exit: false` | +| `mean` | Aggregates repeated attempt scores by mean | +| `confidence_interval` | Uses the lower bound of a 95% confidence interval as the conservative score | + +`repeat.cost_limit_usd` caps repeat-run spend. `repeat.costLimitUsd` is also +accepted for prerelease trial-schema parity, but new YAML should use +`cost_limit_usd`. + +## Vercel-compatible shorthand + +AgentV also accepts Vercel-style top-level `runs` and `early_exit`: + +```yaml +runs: 4 +early_exit: true +``` + +This is shorthand for a `pass_at_k` repeat run. Use `repeat` when you need +AgentV-specific strategy or cost-limit fields. + +Do not set both `repeat` and `runs` in the same experiment. `repeat` is the +canonical AgentV shape; `runs` exists only for Vercel-compatible shorthand. + +Vercel defines the requested run count at the experiment level. Some result +summaries show fewer actual runs for a case because `earlyExit: true` stops +remaining attempts after the first pass; smoke runs can also force one run. +AgentV follows the same experiment-level placement while keeping the richer +`repeat` block for AgentV strategies. + +Repeat-enabled cases use a Vercel-style physical layout with AgentV aggregate +provenance: + +```text +/index.jsonl +/benchmark.json +///summary.json +///grading.json +///run-1/result.json +///run-1/transcript.json +///run-1/transcript-raw.jsonl +///run-1/outputs/answer.md +///run-1/grading.json +``` + +The repeated case aggregate folder uses `summary.json` for run-count, pass-rate, +fingerprint, and flattened snake_case timing fields such as +`mean_duration_ms`, and `grading.json` for compact trial/aggregation verdicts. +Each `run-N/result.json` is the per-attempt manifest and includes +`grading_path`, transcript/output paths, and embedded timing/o11y metrics. +Root `index.jsonl` and root `benchmark.json` remain stable for existing CI +summary scripts and uploaded artifact consumers. + +## Targets and setup + +Experiments reuse targets from `.agentv/targets.yaml`; they do not define a new +provider registry. + +```yaml +targets: + - copilot + - claude + - name: gemini-with-hooks + use_target: gemini +``` + +Setup and scripts belong on the experiment because they are often the A/B +variable: + +```yaml +setup: + - script: cp skills/with-docs/AGENTS.md AGENTS.md +scripts: + - script: bun test + timeout_seconds: 120 +``` + +## Running experiments + +Run a specific experiment: + +```bash +bun agentv eval evals/suite.eval.yaml --experiment experiments/default.yaml +``` + +If no experiment is passed, AgentV checks `.agentv/config.yaml` for a default: + +```yaml +experiments: + default: experiments/default.yaml +``` + +If no default is configured, AgentV keeps the old behavior and uses the +`default` experiment label. + +## Schema + +The generated JSON Schema is available at +`skills-data/agentv-eval-writer/references/experiment-schema.json`. diff --git a/apps/web/src/content/docs/docs/evaluation/running-evals.mdx b/apps/web/src/content/docs/docs/evaluation/running-evals.mdx index 22efc61e1..460a43f7b 100644 --- a/apps/web/src/content/docs/docs/evaluation/running-evals.mdx +++ b/apps/web/src/content/docs/docs/evaluation/running-evals.mdx @@ -114,8 +114,8 @@ my-results/ / grading.json timing.json - input.md - outputs/response.md + task/PROMPT.md + outputs/answer.md task/ EVAL.yaml targets.yaml @@ -132,6 +132,11 @@ case directories are still useful for organizing bulky prompts, fixtures, or tests while authoring an eval, but they are optional input organization rather than a separate artifact schema. +If the source eval uses the `PROMPT.md` fallback instead of inline `input`, +AgentV still writes the resolved task prompt to the generated +`task/PROMPT.md`. This keeps the run artifact self-contained without requiring +the same long prompt to be duplicated in both Markdown and YAML. + ### Manual or External-Agent Attempts Use `agentv prepare` when you want AgentV to set up one eval case but a human, diff --git a/apps/web/src/content/docs/docs/tools/results.mdx b/apps/web/src/content/docs/docs/tools/results.mdx index e5f653ed8..63eee4867 100644 --- a/apps/web/src/content/docs/docs/tools/results.mdx +++ b/apps/web/src/content/docs/docs/tools/results.mdx @@ -112,13 +112,21 @@ Duplicate policy is explicit: ### Metrics sidecar -Each per-case artifact directory includes `metrics.json` +Each direct per-case artifact directory includes `metrics.json` (`schema_version: "agentv.metrics.v1"`). This is an AgentV-owned derived projection over `trace.json`, the result row, and `grading.json`. It is the compact executor behavior summary for dashboards, comparison exports, and metric-style graders; it is not canonical trace storage and does not carry token/cost usage. +Repeat-enabled cases use aggregate `summary.json` with flattened snake_case +timing fields plus aggregate `grading.json`, then store attempt details under +`run-N/`. Each `run-N/` contains a compact per-attempt manifest +`result.json`, `transcript.json`, `transcript-raw.jsonl`, and `outputs/answer.md`, +plus AgentV `grading.json`. The `result.json` file carries `grading_path`, +transcript/output paths, and embedded timing/o11y metrics; repeat attempts do +not write a separate `metrics.json` sidecar. + `transcript.jsonl` remains the ordered conversational/log compatibility projection. Full trace detail stays in `trace.json` (`agentv.trace.v1`). `benchmark.json` remains the run-level aggregate summary, @@ -143,11 +151,11 @@ Vercel `@vercel/agent-eval` `results.o11y` maps into AgentV like this: | `shellCommands` | `metrics.shell_commands` | `metrics.json` | | `filesRead` | `metrics.files_read` | `metrics.json` | | `filesModified` | `metrics.files_modified` | `metrics.json` | -| `toolCalls` | `metrics.tool_call_events`, `metrics.tool_calls`, and `metrics.tool_call_counts` | `metrics.json`; compact counts can also appear in `grading.json.execution_metrics` and `benchmark.json.run_summary[*].tool_calls` | +| `toolCalls` | `metrics.tool_call_events`, `metrics.tool_calls`, and `metrics.tool_call_counts` | `metrics.json`; compact counts can also appear in `benchmark.json.run_summary[*].tool_calls` | | `totalToolCalls` | `metrics.total_tool_calls` | `metrics.json` | | `webFetches` | `metrics.web_fetches` | `metrics.json` | | `totalTurns` | `metrics.total_turns` | `metrics.json`; conversational rows remain in `transcript.jsonl` | -| `errors` | `metrics.errors` | `metrics.json`; compact count remains in `grading.json.execution_metrics.errors_encountered` | +| `errors` | `metrics.errors` | `metrics.json` | | `thinkingBlocks` | `metrics.reasoning_blocks` and `thinking_blocks` | `metrics.json` | Agent Skills eval artifacts map into AgentV like this: @@ -155,7 +163,7 @@ Agent Skills eval artifacts map into AgentV like this: | Agent Skills pattern | AgentV field | Artifact location | |----------------------|--------------|-------------------| | Authored `evals/evals.json` cases | AgentV eval cases and task bundle paths | Eval source plus optional `task_dir`, `eval_path`, `targets_path`, `files_path`, and `graders_path` in `index.jsonl` | -| Per-case `outputs/` | Generated target output artifacts | `outputs/answer.md`, `outputs/response.md` | +| Per-case answer | Generated target output artifact | `outputs/answer.md` | | Per-case sidecars | Trace, transcript, metrics, and raw provider evidence | `trace.json`, `transcript.jsonl`, `metrics.json`, `provider.log` | | Per-case `timing.json` | Duration, token totals, cost, and usage source labels | `timing.json` | | Per-case `grading.json` | Assertions, graders, execution metrics, workspace changes | `grading.json`; summary fields can reference the same trace/result facts | @@ -198,7 +206,7 @@ export `index.jsonl` and use `artifact_refs.status: "emitted"`. Raw prompt text, final output, and tool arguments/results are excluded by default, and raw-bearing artifact refs such as `grading_path`, `input_path`, -`answer_path`, `response_path`, `transcript_path`, and `trace_path` are omitted from +`answer_path`, `transcript_path`, and `trace_path` are omitted from metadata-only bundles. To include raw payloads and raw-bearing refs in the bundle, opt in explicitly: diff --git a/docs/adr/2026-06-23-experiments-vs-eval-separation.md b/docs/adr/2026-06-23-experiments-vs-eval-separation.md new file mode 100644 index 000000000..aae9955f8 --- /dev/null +++ b/docs/adr/2026-06-23-experiments-vs-eval-separation.md @@ -0,0 +1,155 @@ +# ADR: Separate experiments from eval definitions + +Date: 2026-06-23 + +Status: Proposed + +## Context + +AgentV currently treats an experiment as a run label. The label is threaded +through evaluation config and recorded in run artifacts, but the agent, model, +harness, repeat count, timeout, sandbox, and setup choices mostly remain in +`eval.yaml` `execution` fields or CLI flags. + +That conflates two concerns: + +- The eval definition is the task contract: prompt, dataset, workspace fixture + required by the task, and assertions or graders. +- The experiment is the run contract: which agent or target is under test, which + model and harness are used, how many runs to execute, what setup is injected, + and which eval cases are selected. + +The public Vercel `agent-eval` ecosystem is a useful reference point. The +`vercel/next.js` evals keep task fixtures under `evals/`, while experiments are +generated or committed separately. `vercel/next-evals-oss` commits many +`experiments/*.ts` variants and stores results under experiment-specific +directories. The useful pattern is the vocabulary and ownership split, not a +requirement that AgentV adopt Vercel's package as its core runtime. + +This decision must also preserve AgentV's existing product boundary: + +- AgentV stays repo-native and zero-infra by default. +- Portable run artifacts remain the source of truth. +- Core primitives should stay small and composable. +- Public wire formats use `snake_case`; TypeScript internals use `camelCase`. +- `project` means the run, trace, and experiment container; `benchmark` means a + curated eval suite. + +## Vocabulary + +An eval is a frozen task definition. It includes the prompt or dataset, expected +behavior, task-owned workspace fixtures, and assertions. AgentV's LLM-judge, +code-grader, deterministic assertions, and hidden or explicit evaluation +criteria belong here. + +An experiment is a committed or generated run definition. It declares which +agent, target, provider, model, harness options, setup steps, run count, timeout, +sandbox, and case selector are used. Setup that changes the system under test, +such as installing dependencies or dropping an `AGENTS.md` or skill file, belongs +here because it is an A/B variable. + +An execution compatibility block is the legacy `eval.yaml` location for runner +selection and runtime controls. It remains supported during migration but should +stop being the canonical home for experiment-level choices. + +## Decision + +AgentV will make experiments first-class configuration units separate from +eval definitions. + +Eval files remain YAML-authored by default. They should describe what is tested: +task inputs, datasets, assertions, and task fixtures. They should not be the +canonical place for which agent, model, harness, setup injection, sandbox, or run +matrix executes the task. + +Experiment files will live under `experiments/` by convention. AgentV will +support YAML as the canonical authoring path for the abstraction story and TypeScript +as the power-user escape hatch: + +```yaml +name: copilot-gpt55-withskill +target: copilot-gpt55 +model: openai/gpt-5.5 +evals: "agent-042-*" +scripts: + - build +repeat: + count: 3 + strategy: pass_at_k + cost_limit_usd: 2.00 +early_exit: false +timeout_seconds: 900 +sandbox: auto +setup: + - script: bun install + - script: cp skills/copilot/AGENTS.md AGENTS.md +``` + +`config.yaml` will gain a default experiment pointer so existing `agentv eval` +usage keeps working: + +```yaml +experiments: + default: experiments/default.yaml +``` + +If no default experiment is configured, AgentV keeps the current behavior and +uses the `default` experiment label. Existing `eval.yaml`-only repositories +remain valid. + +Legacy `eval.yaml execution` fields that select targets, targets matrices, +workers, cache, budget, thresholds, and workspace runtime behavior will +continue to parse as a compatibility shim until docs and examples have moved. +The prerelease `execution.trials` surface is hard-removed with no alias: run +counts live on the experiment as canonical `repeat` config, with Vercel-style +`runs`/`early_exit` accepted as shorthand for `pass_at_k`. + +AgentV should adopt Vercel's structure and lowest-common-denominator contract +ideas, not depend on `@vercel/agent-eval` as core infrastructure in this phase. +The package's `ExperimentConfig` shape is a strong public reference for +experiment vocabulary: agent, model, agent options, case filter, scripts, runs, +early exit, timeout, sandbox, and setup. A direct dependency would force AgentV +to absorb Vercel's fixture model, sandbox assumptions, result caching semantics, +and TypeScript-first authoring story before those boundaries are stable for +AgentV. + +## Consequences + +Positive: + +- Evals become portable task definitions that can be run against multiple agents + without editing the task file. +- A/B setup variants such as baseline versus skill injection become reviewable, + committed experiment files. +- Existing artifact paths already use experiment labels, so this decision extends + an established storage axis instead of introducing a parallel result concept. +- The default experiment pointer gives old repos a non-breaking migration path. +- AgentV can align with Vercel conventions while preserving YAML authoring, + LLM-judge assertions, workspace fixtures, and Git-backed artifacts. + +Negative: + +- The migration creates two valid locations for some runtime controls until + deprecation completes. +- The CLI must resolve explicit experiments, configured defaults, and legacy + label-only runs without surprising users. +- Artifact readers need a richer experiment fingerprint and provenance model + beyond the current string label. + +## Non-Goals + +- Do not replace AgentV's evaluator engine with `@vercel/agent-eval` in the + initial migration. +- Do not convert AgentV eval YAML into Vercel `PROMPT.md` plus `EVAL.ts`. +- Do not move LLM-judge assertions out of eval definitions. +- Do not make Phoenix, Harbor, Opik, Vercel Sandbox, or another external system + required for local execution. +- Do not break existing `eval.yaml` files or current result artifacts. + +## References + +- Vercel `agent-eval`: https://github.com/vercel-labs/agent-eval +- Vercel Next.js eval results: https://github.com/vercel/next-evals-oss +- Anthropic Skills schema vocabulary: https://github.com/anthropics/skills/blob/main/skills/skill-creator/references/schemas.md +- Hugging Face Datasets vocabulary: https://huggingface.co/docs/datasets/en/package_reference/main_classes +- OpenInference trace vocabulary: https://arize-ai.github.io/openinference/spec/ diff --git a/docs/plans/2026-06-23-001-feat-repeat-runs-flaky-evals-plan.md b/docs/plans/2026-06-23-001-feat-repeat-runs-flaky-evals-plan.md index dd6759281..8aafcc5a6 100644 --- a/docs/plans/2026-06-23-001-feat-repeat-runs-flaky-evals-plan.md +++ b/docs/plans/2026-06-23-001-feat-repeat-runs-flaky-evals-plan.md @@ -87,8 +87,8 @@ Repeat-run reliability uses a separate metric: `attempt_success_rate`. It means Suggested wire shape: - `run-N/result.json.result.pass_rate`, or the equivalent current `run-N/grading.json` summary field, is assertion-level pass rate for that attempt and is omitted when the verifier has no assertion counts. -- `/summary.json.pass_rate` is optional aggregate assertion pass-rate stats, not binary attempt success frequency. -- `/summary.json.attempts.success_rate` and the index convenience field `attempt_success_rate` represent successful counted attempts divided by counted attempts. +- `/grading.json.summary.pass_rate` is aggregate assertion pass-rate stats, not binary attempt success frequency. +- `/summary.json`, the index `trials[]`/`aggregation`, and future explicit repeat fields represent successful counted attempts divided by counted attempts. - Binary-only harnesses write `passed: true` or `passed: false` per attempt and derive only `attempt_success_rate` across repeated attempts. --- @@ -100,7 +100,7 @@ Before introducing AgentV-specific contract shapes, implementation should check | Reference | Lowest common denominator to reuse | Intentional AgentV divergence | | --- | --- | --- | | Claude Skills schema | Use assertion, expectation, grading, `passed`, `failed`, `total`, and assertion-level `pass_rate` vocabulary for graders that expose assertion counts. | Do not copy the full skill-eval artifact shape. AgentV keeps `.agentv/results///...` as the portable run bundle and uses `attempt_success_rate` for repeat-run reliability. | -| Vercel `agent-eval` | Reuse fixture-driven hidden verifier ergonomics, durable `run-1`, `run-2` attempt directories, and aggregate case summaries. | Rename Vercel `passRate` to `attempt_success_rate` in AgentV, because Vercel's metric is passed attempts divided by total attempts. Do not inherit ambiguous CLI gating semantics. | +| Vercel `agent-eval` | Reuse fixture-driven hidden verifier ergonomics, case-level `summary.json`, and durable `run-1`, `run-2` attempt directories. | Keep AgentV root `benchmark.json` for current run-level compatibility, but do not write per-attempt `benchmark.json`. Rename Vercel `passRate` to `attempt_success_rate` where attempt-frequency stats are exposed in AgentV-specific artifacts. Do not inherit ambiguous CLI gating semantics. | | Hugging Face Datasets | Keep dataset, split, record, features, and row-oriented corpus vocabulary for eval inputs and benchmark corpora. Treat an AgentV case as a record-like unit when mapping to external datasets. | Do not require Arrow, the Hub, DatasetDict, or HF storage layout. AgentV cases remain repo files or generated case records inside benchmark/project artifacts. | | OpenInference | Preserve trace/span/tool-call/model-observability semantics when naming trace metadata and external trace correlation fields. | Do not require OpenTelemetry collection, Phoenix, or OpenInference export as core runtime infrastructure. AgentV stores portable traces/transcripts as artifacts and supports link-out correlation through `external_trace` metadata. | @@ -148,12 +148,12 @@ Public docs and implementation notes must not reference non-public sources. If a ## Key Technical Decisions -- KTD1. Prefer `execution.repeat` as the durable public config surface. The existing `execution.trials` code path mixes attempt aggregation and gate behavior; implementation should either hard-correct it before stable release or keep it as a compatibility alias with explicit mapping and warnings. +- KTD1. The repeat config attaches to the **experiment** surface, not to `eval.yaml` `execution`, per the experiments-separation decision (epic `av-991`, recorded on `av-991.1`). This aligns with Vercel agent-eval, where `runs`/`earlyExit` are experiment-level. This epic (`av-i0l`) owns the repeat **mechanics** (schema shape, gate policies, attempt aggregation, flake classification, and the run-N artifact layout); `av-991` owns **placement** (the experiment contract the repeat block lives on). The existing `execution.trials` code path is **hard-removed** (no compatibility alias) because usage is rare; its behavior is replaced by the experiment-level repeat block. Because the experiment surface is delivered by `av-991`, the schema work in `av-i0l.1` depends on that contract landing. - KTD2. Keep one-run CI as the default. Repeat runs are for reliability evidence unless `repeat.gate` says they are a CI gate. -- KTD3. Store aggregate rows in the top-level `index.jsonl`, not one row per attempt. Attempt details live in case-local `summary.json` and `run-N/` directories so existing aggregate consumers do not inflate case counts. +- KTD3. Store aggregate rows in the top-level `index.jsonl`, not one row per attempt. Attempt details live in case-local `summary.json`, `grading.json`, and `run-N/` directories so existing aggregate consumers do not inflate case counts. - KTD4. Single-run cases keep direct case-local files instead of always nesting under `run-1`. This preserves the simple default artifact shape and makes `.agentv/results////grading.json` easy to inspect. Repeat-enabled cases use `run-1/`, `run-2/`, and so on under the case directory. -- KTD5. Write `summary.json` for every case after the artifact-layout migration. In a single-run case it summarizes the direct files; in a repeat-run case it summarizes the `run-N/` children. -- KTD6. Default repeat reporting uses full sampling with no early exit. Early exit saves cost but biases reliability statistics, so it should be explicit and recorded. +- KTD5. Root run aggregates keep the existing AgentV `benchmark.json` for compatibility. Repeat case aggregates use `summary.json` with flattened snake_case timing fields plus AgentV aggregate `grading.json`; repeat attempts use `run-N/` children. +- KTD6. `pass_at_k` keeps the existing AgentV/Vercel ergonomics: early exit is enabled unless explicitly disabled. Full reliability sampling requires `early_exit: false` on the experiment and should be recorded because it changes cost and statistics. - KTD7. Do not inherit Vercel's implicit CI ambiguity. All policies that can make one failed plus one passed attempt count as passing must be visible in config and artifacts. - KTD8. Reuse current failure classification fields before adding new enums. Add aggregate classification fields only after mapping from `execution_status`, `failure_stage`, and `failure_reason_code` proves insufficient. - KTD9. Do not reuse `pass_rate` for attempt success frequency. AgentV uses `attempt_success_rate` for repeat-run reliability and reserves `pass_rate` for assertion or expectation pass rate. @@ -165,7 +165,7 @@ Public docs and implementation notes must not reference non-public sources. If a ```mermaid flowchart TB - Config[Eval config execution.repeat] --> Runner[Eval runner] + Config[Experiment repeat config] --> Runner[Eval runner] Runner --> Attempt1[Attempt run-1] Runner --> Attempt2[Attempt run-2] Runner --> AttemptN[Attempt run-N] @@ -187,42 +187,45 @@ The runner executes the configured attempts, writes attempt artifacts, computes Preferred v1 shape: ```yaml -execution: - repeat: - runs: 3 - seed: 1234 - max_parallel_attempts: 1 - timeout_ms: 300000 - budget_usd: 5 - early_exit: never - retry: - max_attempts: 1 - on: - - verifier_error - - infrastructure_error - - timeout - gate: - policy: attempt_success_rate_at_least - threshold: 0.8 +repeat: + count: 3 + strategy: pass_at_k + cost_limit_usd: 5 + seed: 1234 + max_parallel_attempts: 1 + timeout_ms: 300000 + early_exit: never + retry: + max_attempts: 1 + on: + - verifier_error + - infrastructure_error + - timeout + gate: + policy: attempt_success_rate_at_least + threshold: 0.8 ``` Field notes: -- `runs` is the planned reliability sample count. Missing or `1` means normal single-run behavior. -- `max_attempts` belongs to retry handling, not reliability sampling. For example, `runs: 3` with `retry.max_attempts: 1` may write up to six physical attempts, but only the three counted attempts feed reliability stats. -- `early_exit` starts with `never`, `on_gate_satisfied`, and `on_gate_failed`. Reliability reports should default to `never`. +- `count` is the planned reliability sample count. Missing repeat config or `count: 1` means normal single-run behavior. +- `strategy` starts with the existing AgentV aggregation strategies: `pass_at_k`, `mean`, and `confidence_interval`. +- `max_attempts` belongs to retry handling, not reliability sampling. For example, `count: 3` with `retry.max_attempts: 1` may write up to six physical attempts, but only the three counted attempts feed reliability stats. +- `early_exit` should be represented by the experiment-level boolean in the native experiments branch. Future gate-aware modes such as `never`, `on_gate_satisfied`, and `on_gate_failed` remain non-goals until gate policies are implemented. - `seed` is best effort. Providers that support deterministic seeds receive a per-attempt seed derived from the base seed and attempt number; providers that do not support seeds record `seed_unsupported`. -- `budget_usd` composes with existing run-level budget controls and stops new attempts when the budget is exhausted. +- `cost_limit_usd` composes with existing run-level budget controls and stops new attempts when the repeat budget is exhausted. -Compatibility with `execution.trials`: +Migration from `execution.trials` (hard removal, no alias): -- If `trials` is still prerelease-only, replace it with `repeat` and migrate tests/docs. -- If `trials` has external consumers, keep it as an alias: - - `trials.count` maps to `repeat.runs`. - - `trials.cost_limit_usd` maps to `repeat.budget_usd`. - - `trials.strategy: pass_at_k` maps to `repeat.gate.policy: any_attempt_successful`. - - `trials.strategy: mean` maps to `mean_score_at_least` only when a threshold is present; otherwise it is report-only aggregation. - - `trials.strategy: confidence_interval` remains report-only until AgentV has a reviewed CI policy for confidence bounds. +`execution.trials` is removed from `eval.yaml` outright; the repeat block lives on the experiment instead. Existing semantics map as follows so any prerelease evals can be ported by hand: + +- `trials.count` maps to `repeat.count`. +- `trials.cost_limit_usd` maps to `repeat.cost_limit_usd`. +- `trials.costLimitUsd` is accepted only as `repeat.costLimitUsd` for prerelease parity; new YAML should use `cost_limit_usd`. +- `trials.strategy: pass_at_k` maps to `repeat.strategy: pass_at_k`. +- `trials.strategy: mean` maps to `repeat.strategy: mean`. +- `trials.strategy: confidence_interval` maps to `repeat.strategy: confidence_interval`. +- Existing gate-policy ideas remain future work; do not overload `strategy` to imply CI policy. --- @@ -241,16 +244,15 @@ Single-run cases keep direct case-local files: ```text .agentv/results///index.jsonl .agentv/results///benchmark.json -.agentv/results////summary.json .agentv/results////grading.json .agentv/results////timing.json -.agentv/results////input.md +.agentv/results////task/PROMPT.md .agentv/results////outputs/trace.json .agentv/results////outputs/transcript.jsonl -.agentv/results////outputs/response.md +.agentv/results////outputs/answer.md ``` -Rationale: the common path stays readable, old mental models stay close, and no user pays a `run-1/` nesting tax for default CI. `summary.json` gives readers a uniform aggregate entry point without moving direct single-run sidecars. +Rationale: the common path stays readable, old mental models stay close, and no user pays a `run-1/` nesting tax for default CI. ### Repeat-Run Case @@ -258,23 +260,31 @@ Repeat-run cases use attempt directories: ```text .agentv/results////summary.json +.agentv/results////grading.json +.agentv/results////run-1/result.json .agentv/results////run-1/grading.json -.agentv/results////run-1/timing.json -.agentv/results////run-1/outputs/trace.json -.agentv/results////run-1/outputs/transcript.jsonl +.agentv/results////run-1/transcript.json +.agentv/results////run-1/transcript-raw.jsonl +.agentv/results////run-1/outputs/answer.md +.agentv/results////run-2/result.json .agentv/results////run-2/grading.json -.agentv/results////run-2/timing.json -.agentv/results////run-2/outputs/trace.json -.agentv/results////run-2/outputs/transcript.jsonl +.agentv/results////run-2/transcript.json +.agentv/results////run-2/transcript-raw.jsonl +.agentv/results////run-2/outputs/answer.md ``` +Each `run-N/result.json` is the per-attempt manifest. It carries paths such as +`grading_path`, `transcript_path`, `transcript_raw_path`, `output_paths.answer`, +plus embedded timing/o11y metrics so repeat attempts do not need a separate +`metrics.json` sidecar. + `` should reuse the sanitized artifact key produced by the current artifact writer, including suite disambiguation where needed. Attempt directories are one-indexed because users naturally inspect `run-1`, `run-2`, and this matches the Vercel comparison. --- ## Aggregation Semantics -Each case summary should expose: +Each repeat case aggregate should expose: | Field | Meaning | | --- | --- | @@ -366,7 +376,7 @@ Dashboard should present repeat-run cases aggregate-first: - Attempt drill-down lists `run-1`, `run-2`, and so on with score, status, duration, cost, failure reason, and retry/exclusion reason. - Selecting an attempt opens the same Checks, Transcript, Source, Files, and Feedback affordances as a normal single-run result. - Dashboard must not hide individual traces, transcripts, raw provider logs, or grader output behind the aggregate. -- Historical single-run rows render as they do today, with `summary.json` absent or minimal. +- Historical single-run rows render as they do today, with `summary.json` absent. For trend and compare views, repeat aggregates should be the default unit. Attempt-level views can be added as a filter later, but they must not silently change run-level counts. @@ -432,13 +442,13 @@ Search behavior: Repeat runs can multiply provider spend. V1 should ship with conservative controls: -- `runs` and retry `max_attempts` must be bounded by validation. +- `repeat.count` and retry `max_attempts` must be bounded by validation. - `max_parallel_attempts` limits per-case concurrent attempts; it composes with existing eval workers and provider-specific concurrency guidance. - Agent-provider targets should keep the existing "limit concurrency to 3 targets" operational guidance. -- `budget_usd` stops scheduling new attempts when exceeded and records `budget_exceeded`. +- `repeat.cost_limit_usd` stops scheduling new attempts when exceeded and records `budget_exceeded`. - `timeout_ms` applies per attempt; existing top-level agent timeout remains the default if no repeat timeout is set. -- `early_exit` must be explicit and recorded in `summary.json`. -- Cache should be disabled or scoped by attempt when repeat runs measure stochastic behavior. Existing code already disables cache for `trials.count > 1`; keep that principle. +- `early_exit` must be recorded in `summary.json` when it differs from the default. +- Cache should be disabled or scoped by attempt when repeat runs measure stochastic behavior. Existing code already disables cache for repeated attempts; keep that principle. - Provider seed support is best effort and must be recorded per attempt so deterministic and stochastic runs are distinguishable. --- @@ -453,17 +463,17 @@ Repeat runs can multiply provider spend. V1 should ship with conservative contro **Bead:** `av-i0l.1`. -**Files:** `packages/core/src/evaluation/types.ts`, `packages/core/src/evaluation/loaders/config-loader.ts`, `packages/core/src/evaluation/validation/eval-file.schema.ts`, `packages/core/scripts/generate-eval-schema.ts`, `packages/core/test/evaluation/loaders/config-loader.test.ts`, `packages/core/test/evaluation/validation/eval-file-schema.test.ts`. +**Files:** `packages/core/src/evaluation/types.ts`, `packages/core/src/evaluation/experiment.ts`, `packages/core/src/evaluation/validation/experiment-file.schema.ts`, `packages/core/scripts/generate-eval-schema.ts`, `packages/core/test/evaluation/experiment.test.ts`, `packages/core/test/evaluation/validation/eval-schema-sync.test.ts`. -**Approach:** Add `execution.repeat` with a narrow schema and normalize it into internal camelCase. Decide whether `execution.trials` is removed as prerelease cleanup or retained as a legacy alias with explicit mapping. +**Approach:** Add the repeat block as a narrow schema on the **experiment** surface (delivered by `av-991`) and normalize it into internal camelCase. **Hard-remove** `execution.trials` from `eval.yaml` as prerelease cleanup (no legacy alias); port any existing usage by hand using the mapping above. **Test Scenarios:** -- Valid `repeat.runs: 3` parses into internal repeat config with no gate policy. -- Invalid `runs`, `threshold`, `max_parallel_attempts`, and retry values are rejected or warned consistently with existing config parsing. +- Valid `repeat.count: 3` parses into internal repeat config with no gate policy. +- Invalid `repeat.count`, `threshold`, `max_parallel_attempts`, and retry values are rejected or warned consistently with existing config parsing. - `all_attempts_successful`, `any_attempt_successful`, `attempt_success_rate_at_least`, and `mean_pass_rate_at_least` validate; `mean_score_at_least` validates only if included in v1. -- Legacy `trials` input maps or fails according to the chosen compatibility decision. -- Generated eval schema stays in sync. +- Legacy eval-level `trials` input fails; the compatibility path is explicit hand migration to experiment `repeat`. +- Generated experiment schema stays in sync. **Verification:** Schema tests pass and docs/examples can reference the accepted YAML shape. @@ -477,11 +487,11 @@ Repeat runs can multiply provider spend. V1 should ship with conservative contro **Files:** `packages/core/src/evaluation/run-artifacts.ts`, `packages/core/src/evaluation/result-row-schema.ts`, `apps/cli/src/commands/eval/artifact-writer.ts`, `apps/cli/src/commands/eval/result-layout.ts`, `apps/cli/test/commands/eval/artifact-writer.test.ts`, `apps/cli/test/commands/eval/aggregate.test.ts`, `apps/cli/test/commands/results/validate.test.ts`. -**Approach:** Extend the artifact writer to understand aggregate results with attempt children. Keep single-run case sidecars direct, add `summary.json`, and add optional repeat fields to index rows. Avoid putting full attempt payloads in `index.jsonl`. +**Approach:** Extend the artifact writer to understand aggregate results with attempt children. Keep single-run case sidecars direct, add case-local `summary.json` for repeat aggregates, and add optional repeat fields to index rows. Avoid putting full attempt payloads in `index.jsonl`. **Test Scenarios:** -- Single-run output writes direct case-local sidecars plus `summary.json` and remains readable by existing manifest hydration. +- Single-run output writes direct case-local sidecars and remains readable by existing manifest hydration. - Repeat-run output writes `summary.json` and `run-1/`, `run-2/` sidecars with correct relative paths. - `index.jsonl` has one aggregate row per case/target and compact attempt references. - Historical rows without repeat fields parse successfully. @@ -564,8 +574,8 @@ Repeat runs can multiply provider spend. V1 should ship with conservative contro | --- | --- | | Hidden CI behavior diverges across commands | Route every repeat-run gate through one policy evaluator and test CLI command paths against the same cases | | Repeat attempts inflate run counts in trend/compare views | Keep one aggregate row per case/target in top-level `index.jsonl` | -| Early exit biases reliability reports | Default to `early_exit: never`, record early-exit mode, and label incomplete samples | -| Existing `trials` behavior conflicts with the new contract | Decide compatibility first in `av-i0l.1`; do not keep two public names with different semantics | +| Early exit biases reliability reports | Keep pass-at-k early exit compatible by default, require `early_exit: false` for full sampling, and label incomplete samples | +| Existing `trials` behavior conflicts with the new contract | Hard-remove eval-level `execution.trials`; preserve its strategies and cost cap only through experiment `repeat` | | Attempt artifacts make rows too large | Keep only compact attempt references in `index.jsonl`; move large attempt indexes to case-local sidecars | | Artifact-layout migration lands concurrently | Depend on shared layout helpers and do not edit the `artifact-results-layout` branch from this worktree | diff --git a/docs/plans/2026-06-23-002-experiments-separation-plan.md b/docs/plans/2026-06-23-002-experiments-separation-plan.md new file mode 100644 index 000000000..7c626ecd1 --- /dev/null +++ b/docs/plans/2026-06-23-002-experiments-separation-plan.md @@ -0,0 +1,407 @@ +--- +title: "feat: Separate experiments from eval definitions" +type: feat +date: 2026-06-23 +origin: docs/adr/2026-06-23-experiments-vs-eval-separation.md +--- + +# feat: Separate experiments from eval definitions + +## Summary + +AgentV should separate eval task definitions from experiment run definitions. +Eval YAML stays the canonical authoring layer for prompts, datasets, assertions, +and task fixtures. Experiments become first-class committed files that select the +agent or target under test, model, harness options, setup injection, run knobs, +and case filter. + +This should ship in phases. Phase 1 adds the non-breaking foundation: +experiment contract types, default experiment resolution, and artifact +attribution by resolved experiment name. Later phases move runtime controls out +of `eval.yaml execution`, teach the CLI to run experiment matrices, and record +full experiment provenance and fingerprints in run bundles. + +## Problem Frame + +Today `experiment` is a string label passed through +`packages/core/src/evaluation/evaluate.ts`, `packages/core/src/evaluation/run-artifacts.ts`, +`packages/core/src/evaluation/results-repo.ts`, and +`packages/core/src/evaluation/trace-envelope.ts`. Runtime choices are still +scattered across CLI flags, TypeScript config, `.agentv/config.yaml`, and +`eval.yaml execution`. + +That makes it hard to review A/B variants such as `baseline` versus +`agents-md`, because the variable under test can be hidden inside the eval +definition. The desired model is: + +- Eval equals what is tested. +- Experiment equals how and with what it is run. +- Setup that changes the agent's environment belongs to the experiment. +- Existing eval-only repositories keep working through a default experiment + fallback. + +## Requirements + +- R1. Existing `eval.yaml` files validate and run without modification. +- R2. Experiment wire config uses `snake_case`; TypeScript types use + `camelCase`. +- R3. `config.yaml` can point at a default experiment, with no pointer falling + back to the current `default` experiment label. +- R4. `agentv eval --experiment