Skip to content

Latest commit

 

History

History
160 lines (131 loc) · 13.2 KB

File metadata and controls

160 lines (131 loc) · 13.2 KB

Agent Improvement Loop

Generated: 2026-06-13T21:20:01.588Z

Source pattern: https://developers.openai.com/cookbook/examples/agents_sdk/agent_improvement_loop

NodeRoom adapts the cookbook loop as: traces -> human/model feedback -> reusable evals -> gate -> Codex handoff -> next harness change.

Latest run artifact: docs/eval/agent-improvement-loop/20260613T211105Z.json

Summary: 44 pass, 2 blocked, 0 fail, 9 skip.

Step Results

Step Lane Status Duration Command
Professional workflow catalog shape deterministic PASS 2.1s npm run eval:professional
Professional catalog proof gate deterministic PASS 0.9s npm run eval:professional:catalog-proofs
Professional proof ledger deterministic PASS 1.1s npm run eval:professional:proofs
GTM/finance workflow evals deterministic PASS 2.9s npx vitest run tests/workflowEvals.test.ts
Algorithm artifact runner smoke deterministic PASS 0.9s npm run algorithm-artifact:smoke
HALO context/path self-improvement smoke deterministic PASS 1.0s npm run halo:self-improve:smoke
HALO harness variant selection deterministic PASS 1.1s npm run halo:variant:select
HALO Convex job-context telemetry deterministic PASS 1.1s npm run halo:convex-context:smoke
Collaboration ladder L1-L6 deterministic PASS 1.6s npm run ladder -- --record
MM-banking credit decision evals deterministic PASS 1.1s npm run eval:credit -- --record
Official benchmark readiness report deterministic PASS 0.8s npm run benchmark:official:readiness
OpenRouter-on-Convex benchmark contract deterministic PASS 0.9s npm run benchmark:openrouter-convex -- --strict
Official benchmark promotion gate deterministic BLOCKED 0.8s npm run benchmark:official:readiness -- --strict
Official benchmark contamination fixture deterministic PASS 2.8s npx vitest run tests/benchmarkContamination.test.ts
BankerToolBench official ingest fixture deterministic PASS 2.5s npx vitest run tests/bankerToolBenchAdapter.test.ts
BankerToolBench sandbox stage fixture deterministic PASS 2.9s npx vitest run tests/bankerToolBenchStage.test.ts
BankerToolBench manifest lock fixture deterministic PASS 1.0s npm run benchmark:bankertoolbench:manifest-lock -- --root .tmp/official-benchmarks/btb-fixture --json-out docs/eval/bankertoolbench-manifest-lock-smoke.json
BankerToolBench staged runner fixture deterministic PASS 3.5s npx vitest run tests/bankerToolBenchRunner.test.ts
BankerToolBench local harness proof gate deterministic PASS 0.9s npm run benchmark:bankertoolbench:proof
BankerToolBench official execution contract deterministic BLOCKED 0.9s npm run benchmark:bankertoolbench:official-contract -- --strict
SpreadsheetBench official ingest fixture deterministic PASS 2.6s npx vitest run tests/spreadsheetBenchAdapter.test.ts
SpreadsheetBench sandbox stage fixture deterministic PASS 2.6s npx vitest run tests/spreadsheetBenchStage.test.ts
SpreadsheetBench V1 full-stage isolation proof deterministic PASS 1.2s npm run benchmark:spreadsheetbench:stage-proof -- --report docs/eval/spreadsheetbench-v1-full-stage-smoke.json --stage-root .tmp/official-benchmarks/staged-v1-full --track spreadsheetbench-v1 --min-tasks 400
SpreadsheetBench V2 public-example stage proof deterministic PASS 0.9s npm run benchmark:spreadsheetbench:stage-proof -- --report docs/eval/spreadsheetbench-v2-stage-smoke.json --stage-root .tmp/official-benchmarks/staged-v2 --track spreadsheetbench-v2 --min-tasks 3
SpreadsheetBench V1 route selection report deterministic PASS 0.8s npm run benchmark:spreadsheetbench:routes -- --stage-root .tmp/official-benchmarks/staged-v1-full --json-out docs/eval/spreadsheetbench-v1-route-selection.json
SpreadsheetBench V2 route selection report deterministic PASS 0.9s npm run benchmark:spreadsheetbench:routes -- --stage-root .tmp/official-benchmarks/staged-v2 --json-out docs/eval/spreadsheetbench-v2-route-selection.json
SpreadsheetBench V1 full-bundle copy-input baseline deterministic PASS 437.8s npm run benchmark:spreadsheetbench:run-chunked -- --stage-root .tmp/official-benchmarks/staged-v1-full --output-root .tmp/official-benchmarks/run-v1-copy-full --mode copy-input-baseline --chunk-size 25 --json-out docs/eval/spreadsheetbench-v1-copy-input-full-smoke.json --max-mismatches 5
SpreadsheetBench workbook score fixture deterministic PASS 2.8s npx vitest run tests/spreadsheetBenchScorer.test.ts
SpreadsheetBench chart package score fixture deterministic PASS 2.1s npx vitest run tests/spreadsheetBenchChartScorer.test.ts
SpreadsheetBench rendered/VLM chart visual probe deterministic PASS 31.0s npm run benchmark:spreadsheetbench:chart-visual:probe -- --strict
SpreadsheetBench staged runner fixture deterministic PASS 6.3s npx vitest run tests/spreadsheetBenchRunner.test.ts
Agent workspace process sandbox deterministic PASS 1.0s npm run benchmark:agent-sandbox -- --json-out docs/eval/agent-workspace-sandbox-smoke.json
Docker/Harbor availability probe deterministic PASS 2.4s npm run benchmark:docker-sandbox:probe -- --require-pass
SpreadsheetBench staged artifact contamination deterministic PASS 0.9s npm run benchmark:contamination -- --root .tmp/official-benchmarks/staged-v1 --strict
SpreadsheetBench V1 full-stage contamination deterministic PASS 1.1s npm run benchmark:contamination -- --root .tmp/official-benchmarks/staged-v1-full --strict
SpreadsheetBench N5 run artifact contamination deterministic PASS 0.8s npm run benchmark:contamination -- --root .tmp/official-benchmarks/run-v1-model-edit-n5 --strict
SpreadsheetBench 3-task N5 run artifact contamination deterministic PASS 0.9s npm run benchmark:contamination -- --root .tmp/official-benchmarks/run-v1-model-edit-3task-n5 --strict
SpreadsheetBench 3-task N5 proof gate deterministic PASS 0.9s npm run benchmark:spreadsheetbench:proof -- --require-sidecar-files
SpreadsheetBench retry run artifact contamination deterministic PASS 1.0s npm run benchmark:contamination -- --root .tmp/official-benchmarks/run-v1-model-edit-retry --strict
SpreadsheetBench V2 staged artifact contamination deterministic PASS 0.9s npm run benchmark:contamination -- --root .tmp/official-benchmarks/staged-v2 --strict
SpreadsheetBench V2 run artifact contamination deterministic PASS 0.9s npm run benchmark:contamination -- --root .tmp/official-benchmarks/run-v2 --strict
BankerToolBench staged artifact contamination deterministic PASS 0.9s npm run benchmark:contamination -- --root .tmp/official-benchmarks/staged-btb --strict
BankerToolBench run artifact contamination deterministic PASS 1.0s npm run benchmark:contamination -- --root .tmp/official-benchmarks/run-btb --strict
Eval regression diff deterministic PASS 1.0s npm run eval:diff
Convex query/action/mutation boundaries deterministic PASS 1.8s npm run convex:boundaries
Architecture budget review deterministic PASS 0.9s npm run architecture:budget -- --human-approved
OpenRouter free-auto discovery live SKIP 0.0s npm run openrouter:free -- --limit=5
HALO live provider path calibration live SKIP 0.0s npm run halo:live-path:calibrate -- --real deepseek/deepseek-v4-flash --repeats 5
Professional live-provider catalog champion live SKIP 0.0s npm run eval:professional:live-catalog -- --real deepseek/deepseek-v4-flash --require-full --retry-failed 2 --json-out docs/eval/professional-live-catalog.json
Chat-first GTM live runtime live SKIP 0.0s npm run eval:chat-intake:live -- --json-out docs/eval/chat-intake-live.json --timeout-ms 240000
Provider parser live smoke live SKIP 0.0s npm run provider-parser:smoke
Convex /free job smoke full-live SKIP 0.0s npm run free-job:smoke
V2 multi-model benchmark full-live SKIP 0.0s npm run benchmark -- --model-timeout-ms=180000 --model-reserve-ms=15000 --row-hard-timeout-ms=210000
Free-auto router ladder full-live SKIP 0.0s npm run ladder:free
Gemini UI media review ui SKIP 0.0s npx tsx scripts/gemini-ui-review.ts

Workflow Coverage

Reviewed file profile: 70 files (23 CSV, 47 XLSX).

Category Cases
gtm_company_research 11
finance_ops 5
eval_harness 2
analytics_optimization 2
legacy_agent_outputs 1

Architecture Before Eval

Research must describe the workflow, architecture gap, and existing capability fit before it proposes evals or code.

Eval trust level Meaning
candidate Generated from traces or research; useful for discussion and advisory runs, not a merge gate.
research_validated Backed by captured sources or online consensus; can create scoped handoffs, still not blocking by default.
contested Credible sources disagree; advisory only, and the eval should check that disagreement is surfaced.
human_verified Reviewed or accepted for critical use; may become blocking when deterministic and safety-safe.

Gate modes: none, advisory, blocking.

Architecture fit checks before adding code:

  • Can existing tools, prompts, context builders, and Convex mutations already handle the case?
  • If not, is the missing piece a query, action, mutation, tool schema, validator, or UI review affordance?
  • What is the smallest implementation that proves the workflow without a new subsystem?
  • What old or proposed layer can be avoided because the existing artifact/job/lock path is enough?

Root-cause labels used for HALO diagnosis:

  • stale_context: The agent acted on old state or missing refreshed context.
  • wrong_tool: The model chose a tool that could not satisfy the workflow contract.
  • missing_read_before_write: A write occurred without a current source read and version.
  • bad_mutation_contract: The server mutation allowed an unsafe or under-specified state change.
  • weak_source_evidence: The output lacked source-backed evidence or cited the wrong evidence.
  • bad_prompt_or_context: Instructions or context did not define the workflow sharply enough.
  • permission_or_visibility: Scope, privacy, or role gating was wrong or underspecified.
  • model_routing_or_budget: The chosen model, budget, or slice policy was unfit for the task.
  • ui_review_friction: The human review or approval surface obscured the right decision.
  • eval_measures_wrong_behavior: The eval target itself is suspect and needs research/calibration.

Generated Eval Ideas

Eval candidate Trust Gate Architecture fit Handoff decision
candidate-gtm-pitchbook-match candidate advisory existing_capability more_research: missing research packet evidence; candidate evals are advisory only
research-validated-finance-reconcile research_validated advisory small_gap implementation
contested-eval-harness-expansion contested advisory existing_capability eval_fixture: contested claims must stay advisory until resolved or explicitly modeled

Codex Handoff

Architecture Budget

Allowed scope: Only files or modules named by the failing trace, failing eval, or handoff evidence.

Default allowed areas:

  • src/nodeagent runtime, tools, context, and compaction
  • Convex job/tool adapters that already participate in the failing flow
  • eval fixtures and deterministic assertions for the affected workflow

Forbidden without human approval:

  • new database tables
  • new services or framework layers
  • new UI surfaces
  • graph/wiki/embedding expansion without a failing workflow eval
  • weakened CAS, lock, draft, auth, privacy, or eval gates

Recommendations

  • Unblock benchmark promotion step official-benchmark-promotion-gate: official BankerToolBench/SpreadsheetBench readiness remains blocked by external benchmark prerequisites.
  • Unblock benchmark promotion step bankertoolbench-official-contract: BTB official contract is missing external Docker/MCP/Gandalf/provenance evidence.
  • Implement scoped handoff for eval candidate research-validated-finance-reconcile.
  • Run skipped free-route-discovery once prerequisites are present: pass --live and set OPENROUTER_API_KEY to discover current free-auto candidates.
  • Run skipped halo-live-path-calibration once prerequisites are present: pass --live and set OPENROUTER_API_KEY to calibrate N=5 live path fingerprints.
  • Run skipped professional-live-catalog once prerequisites are present: pass --live and set OPENROUTER_API_KEY to prove the professional catalog with the cheap champion route.
  • Run skipped chat-intake-live-runtime once prerequisites are present: pass --live and set OPENROUTER_API_KEY to run the chat-intake room runtime against a real route.
  • Persist each new live trace into a durable eval fixture before promoting README charts.

Next Live Runs

  • npm run agent:improve -- --live
  • npm run agent:improve -- --full-live
  • npm run agent:improve -- --ui-media=docs/eval/ui-recordings/<recording-or-screenshot>
  • npm run benchmark:charts