Generated: 2026-06-13T21:20:01.588Z
Source pattern: https://developers.openai.com/cookbook/examples/agents_sdk/agent_improvement_loop
NodeRoom adapts the cookbook loop as: traces -> human/model feedback -> reusable evals -> gate -> Codex handoff -> next harness change.
Latest run artifact: docs/eval/agent-improvement-loop/20260613T211105Z.json
Summary: 44 pass, 2 blocked, 0 fail, 9 skip.
| Step | Lane | Status | Duration | Command |
|---|---|---|---|---|
| Professional workflow catalog shape | deterministic | PASS | 2.1s | npm run eval:professional |
| Professional catalog proof gate | deterministic | PASS | 0.9s | npm run eval:professional:catalog-proofs |
| Professional proof ledger | deterministic | PASS | 1.1s | npm run eval:professional:proofs |
| GTM/finance workflow evals | deterministic | PASS | 2.9s | npx vitest run tests/workflowEvals.test.ts |
| Algorithm artifact runner smoke | deterministic | PASS | 0.9s | npm run algorithm-artifact:smoke |
| HALO context/path self-improvement smoke | deterministic | PASS | 1.0s | npm run halo:self-improve:smoke |
| HALO harness variant selection | deterministic | PASS | 1.1s | npm run halo:variant:select |
| HALO Convex job-context telemetry | deterministic | PASS | 1.1s | npm run halo:convex-context:smoke |
| Collaboration ladder L1-L6 | deterministic | PASS | 1.6s | npm run ladder -- --record |
| MM-banking credit decision evals | deterministic | PASS | 1.1s | npm run eval:credit -- --record |
| Official benchmark readiness report | deterministic | PASS | 0.8s | npm run benchmark:official:readiness |
| OpenRouter-on-Convex benchmark contract | deterministic | PASS | 0.9s | npm run benchmark:openrouter-convex -- --strict |
| Official benchmark promotion gate | deterministic | BLOCKED | 0.8s | npm run benchmark:official:readiness -- --strict |
| Official benchmark contamination fixture | deterministic | PASS | 2.8s | npx vitest run tests/benchmarkContamination.test.ts |
| BankerToolBench official ingest fixture | deterministic | PASS | 2.5s | npx vitest run tests/bankerToolBenchAdapter.test.ts |
| BankerToolBench sandbox stage fixture | deterministic | PASS | 2.9s | npx vitest run tests/bankerToolBenchStage.test.ts |
| BankerToolBench manifest lock fixture | deterministic | PASS | 1.0s | npm run benchmark:bankertoolbench:manifest-lock -- --root .tmp/official-benchmarks/btb-fixture --json-out docs/eval/bankertoolbench-manifest-lock-smoke.json |
| BankerToolBench staged runner fixture | deterministic | PASS | 3.5s | npx vitest run tests/bankerToolBenchRunner.test.ts |
| BankerToolBench local harness proof gate | deterministic | PASS | 0.9s | npm run benchmark:bankertoolbench:proof |
| BankerToolBench official execution contract | deterministic | BLOCKED | 0.9s | npm run benchmark:bankertoolbench:official-contract -- --strict |
| SpreadsheetBench official ingest fixture | deterministic | PASS | 2.6s | npx vitest run tests/spreadsheetBenchAdapter.test.ts |
| SpreadsheetBench sandbox stage fixture | deterministic | PASS | 2.6s | npx vitest run tests/spreadsheetBenchStage.test.ts |
| SpreadsheetBench V1 full-stage isolation proof | deterministic | PASS | 1.2s | npm run benchmark:spreadsheetbench:stage-proof -- --report docs/eval/spreadsheetbench-v1-full-stage-smoke.json --stage-root .tmp/official-benchmarks/staged-v1-full --track spreadsheetbench-v1 --min-tasks 400 |
| SpreadsheetBench V2 public-example stage proof | deterministic | PASS | 0.9s | npm run benchmark:spreadsheetbench:stage-proof -- --report docs/eval/spreadsheetbench-v2-stage-smoke.json --stage-root .tmp/official-benchmarks/staged-v2 --track spreadsheetbench-v2 --min-tasks 3 |
| SpreadsheetBench V1 route selection report | deterministic | PASS | 0.8s | npm run benchmark:spreadsheetbench:routes -- --stage-root .tmp/official-benchmarks/staged-v1-full --json-out docs/eval/spreadsheetbench-v1-route-selection.json |
| SpreadsheetBench V2 route selection report | deterministic | PASS | 0.9s | npm run benchmark:spreadsheetbench:routes -- --stage-root .tmp/official-benchmarks/staged-v2 --json-out docs/eval/spreadsheetbench-v2-route-selection.json |
| SpreadsheetBench V1 full-bundle copy-input baseline | deterministic | PASS | 437.8s | npm run benchmark:spreadsheetbench:run-chunked -- --stage-root .tmp/official-benchmarks/staged-v1-full --output-root .tmp/official-benchmarks/run-v1-copy-full --mode copy-input-baseline --chunk-size 25 --json-out docs/eval/spreadsheetbench-v1-copy-input-full-smoke.json --max-mismatches 5 |
| SpreadsheetBench workbook score fixture | deterministic | PASS | 2.8s | npx vitest run tests/spreadsheetBenchScorer.test.ts |
| SpreadsheetBench chart package score fixture | deterministic | PASS | 2.1s | npx vitest run tests/spreadsheetBenchChartScorer.test.ts |
| SpreadsheetBench rendered/VLM chart visual probe | deterministic | PASS | 31.0s | npm run benchmark:spreadsheetbench:chart-visual:probe -- --strict |
| SpreadsheetBench staged runner fixture | deterministic | PASS | 6.3s | npx vitest run tests/spreadsheetBenchRunner.test.ts |
| Agent workspace process sandbox | deterministic | PASS | 1.0s | npm run benchmark:agent-sandbox -- --json-out docs/eval/agent-workspace-sandbox-smoke.json |
| Docker/Harbor availability probe | deterministic | PASS | 2.4s | npm run benchmark:docker-sandbox:probe -- --require-pass |
| SpreadsheetBench staged artifact contamination | deterministic | PASS | 0.9s | npm run benchmark:contamination -- --root .tmp/official-benchmarks/staged-v1 --strict |
| SpreadsheetBench V1 full-stage contamination | deterministic | PASS | 1.1s | npm run benchmark:contamination -- --root .tmp/official-benchmarks/staged-v1-full --strict |
| SpreadsheetBench N5 run artifact contamination | deterministic | PASS | 0.8s | npm run benchmark:contamination -- --root .tmp/official-benchmarks/run-v1-model-edit-n5 --strict |
| SpreadsheetBench 3-task N5 run artifact contamination | deterministic | PASS | 0.9s | npm run benchmark:contamination -- --root .tmp/official-benchmarks/run-v1-model-edit-3task-n5 --strict |
| SpreadsheetBench 3-task N5 proof gate | deterministic | PASS | 0.9s | npm run benchmark:spreadsheetbench:proof -- --require-sidecar-files |
| SpreadsheetBench retry run artifact contamination | deterministic | PASS | 1.0s | npm run benchmark:contamination -- --root .tmp/official-benchmarks/run-v1-model-edit-retry --strict |
| SpreadsheetBench V2 staged artifact contamination | deterministic | PASS | 0.9s | npm run benchmark:contamination -- --root .tmp/official-benchmarks/staged-v2 --strict |
| SpreadsheetBench V2 run artifact contamination | deterministic | PASS | 0.9s | npm run benchmark:contamination -- --root .tmp/official-benchmarks/run-v2 --strict |
| BankerToolBench staged artifact contamination | deterministic | PASS | 0.9s | npm run benchmark:contamination -- --root .tmp/official-benchmarks/staged-btb --strict |
| BankerToolBench run artifact contamination | deterministic | PASS | 1.0s | npm run benchmark:contamination -- --root .tmp/official-benchmarks/run-btb --strict |
| Eval regression diff | deterministic | PASS | 1.0s | npm run eval:diff |
| Convex query/action/mutation boundaries | deterministic | PASS | 1.8s | npm run convex:boundaries |
| Architecture budget review | deterministic | PASS | 0.9s | npm run architecture:budget -- --human-approved |
| OpenRouter free-auto discovery | live | SKIP | 0.0s | npm run openrouter:free -- --limit=5 |
| HALO live provider path calibration | live | SKIP | 0.0s | npm run halo:live-path:calibrate -- --real deepseek/deepseek-v4-flash --repeats 5 |
| Professional live-provider catalog champion | live | SKIP | 0.0s | npm run eval:professional:live-catalog -- --real deepseek/deepseek-v4-flash --require-full --retry-failed 2 --json-out docs/eval/professional-live-catalog.json |
| Chat-first GTM live runtime | live | SKIP | 0.0s | npm run eval:chat-intake:live -- --json-out docs/eval/chat-intake-live.json --timeout-ms 240000 |
| Provider parser live smoke | live | SKIP | 0.0s | npm run provider-parser:smoke |
| Convex /free job smoke | full-live | SKIP | 0.0s | npm run free-job:smoke |
| V2 multi-model benchmark | full-live | SKIP | 0.0s | npm run benchmark -- --model-timeout-ms=180000 --model-reserve-ms=15000 --row-hard-timeout-ms=210000 |
| Free-auto router ladder | full-live | SKIP | 0.0s | npm run ladder:free |
| Gemini UI media review | ui | SKIP | 0.0s | npx tsx scripts/gemini-ui-review.ts |
Reviewed file profile: 70 files (23 CSV, 47 XLSX).
| Category | Cases |
|---|---|
| gtm_company_research | 11 |
| finance_ops | 5 |
| eval_harness | 2 |
| analytics_optimization | 2 |
| legacy_agent_outputs | 1 |
Research must describe the workflow, architecture gap, and existing capability fit before it proposes evals or code.
| Eval trust level | Meaning |
|---|---|
| candidate | Generated from traces or research; useful for discussion and advisory runs, not a merge gate. |
| research_validated | Backed by captured sources or online consensus; can create scoped handoffs, still not blocking by default. |
| contested | Credible sources disagree; advisory only, and the eval should check that disagreement is surfaced. |
| human_verified | Reviewed or accepted for critical use; may become blocking when deterministic and safety-safe. |
Gate modes: none, advisory, blocking.
Architecture fit checks before adding code:
- Can existing tools, prompts, context builders, and Convex mutations already handle the case?
- If not, is the missing piece a query, action, mutation, tool schema, validator, or UI review affordance?
- What is the smallest implementation that proves the workflow without a new subsystem?
- What old or proposed layer can be avoided because the existing artifact/job/lock path is enough?
Root-cause labels used for HALO diagnosis:
stale_context: The agent acted on old state or missing refreshed context.wrong_tool: The model chose a tool that could not satisfy the workflow contract.missing_read_before_write: A write occurred without a current source read and version.bad_mutation_contract: The server mutation allowed an unsafe or under-specified state change.weak_source_evidence: The output lacked source-backed evidence or cited the wrong evidence.bad_prompt_or_context: Instructions or context did not define the workflow sharply enough.permission_or_visibility: Scope, privacy, or role gating was wrong or underspecified.model_routing_or_budget: The chosen model, budget, or slice policy was unfit for the task.ui_review_friction: The human review or approval surface obscured the right decision.eval_measures_wrong_behavior: The eval target itself is suspect and needs research/calibration.
| Eval candidate | Trust | Gate | Architecture fit | Handoff decision |
|---|---|---|---|---|
| candidate-gtm-pitchbook-match | candidate | advisory | existing_capability | more_research: missing research packet evidence; candidate evals are advisory only |
| research-validated-finance-reconcile | research_validated | advisory | small_gap | implementation |
| contested-eval-harness-expansion | contested | advisory | existing_capability | eval_fixture: contested claims must stay advisory until resolved or explicitly modeled |
Allowed scope: Only files or modules named by the failing trace, failing eval, or handoff evidence.
Default allowed areas:
- src/nodeagent runtime, tools, context, and compaction
- Convex job/tool adapters that already participate in the failing flow
- eval fixtures and deterministic assertions for the affected workflow
Forbidden without human approval:
- new database tables
- new services or framework layers
- new UI surfaces
- graph/wiki/embedding expansion without a failing workflow eval
- weakened CAS, lock, draft, auth, privacy, or eval gates
- Unblock benchmark promotion step official-benchmark-promotion-gate: official BankerToolBench/SpreadsheetBench readiness remains blocked by external benchmark prerequisites.
- Unblock benchmark promotion step bankertoolbench-official-contract: BTB official contract is missing external Docker/MCP/Gandalf/provenance evidence.
- Implement scoped handoff for eval candidate research-validated-finance-reconcile.
- Run skipped free-route-discovery once prerequisites are present: pass --live and set OPENROUTER_API_KEY to discover current free-auto candidates.
- Run skipped halo-live-path-calibration once prerequisites are present: pass --live and set OPENROUTER_API_KEY to calibrate N=5 live path fingerprints.
- Run skipped professional-live-catalog once prerequisites are present: pass --live and set OPENROUTER_API_KEY to prove the professional catalog with the cheap champion route.
- Run skipped chat-intake-live-runtime once prerequisites are present: pass --live and set OPENROUTER_API_KEY to run the chat-intake room runtime against a real route.
- Persist each new live trace into a durable eval fixture before promoting README charts.
npm run agent:improve -- --livenpm run agent:improve -- --full-livenpm run agent:improve -- --ui-media=docs/eval/ui-recordings/<recording-or-screenshot>npm run benchmark:charts