Implementation of arXiv 2607.08010 (Kujanpää et al., Amazon Fulfillment Technologies & Robotics, Jul 2026), transplanted from Amazon's internal robotics-monitoring stack to real Kraken/Binance market data. No mocks — live REST endpoints, 312MB of historical trade archives, and human-labeled ground truth.
| System | Accuracy (training cases) | Accuracy (hold-out) | Median latency / node | Tokens / node |
|---|---|---|---|---|
| Baseline (code-writing sub-agent) | 22/25 | — | 16,717 ms | ~7,084 |
| Frozen tools (this pipeline) | 25/25 | 5/5 | 1.74 ms | 0 |
- ~9,600× median latency reduction on decision nodes, at zero tokens.
- The hold-out set was labeled after tool generation; 4 of its 5 cases use pairs (ETH, SOL) absent from training entirely — the tools encoded the labeling conventions, not the examples.
- The repair loop never fired: 6/6 tools froze at version 1 with filtered traces and convention-bearing labels as input.
LLM agents in production regenerate the same code for the same procedural steps on every request — the same instruction re-interpreted, the same schema re-discovered, similar code re-written. This repetition costs latency, tokens, and reliability (run-to-run variance). The paper replaces this loop: each repeated step is compiled once into a validated, versioned tool; at runtime the agent calls the tool, falling back to code generation only when a tool is missing or fails.
The procedure lives in an SOP (Standard Operating Procedure) — a human-written decision tree of nodes: observation nodes (did a metric cross its threshold?), root-cause nodes, constraint nodes (gates before expensive actions), and action nodes. A main agent walks this tree per alarm; in the baseline, each node is delegated to a code-writing sub-agent that queries backends via MCP.
The tool-making pipeline has three components:
- Data-collector sub-agent — runs the baseline code-writing agent on labeled cases for the node: writes code, hits errors, corrects itself, produces graded verdicts — and records everything into a trace.
- Tool-maker LLM — writes the permanent, deterministic tool for the node, conditioned on the SOP text + the labeled cases + the trace.
- Reflector LLM — tests the candidate tool against the full labeled set; on failures it diagnoses and rewrites (max 3 rounds), guarding against special-casing the failing examples.
A trace is the artifact the first component exists to produce. Every time the sub-agent's code hits a real API and gets corrected — a renamed response key, a string-typed price, a surprise header — that discovered configuration knowledge is captured. Without a trace, the tool-maker would re-hit every schema surprise at generation time; with it, the working code and the quirks it already absorbed are the raw material the tool is compiled from.
Tools return tiny, structured verdicts — true | false | no_data plus the observed value, threshold, and a short explanation. Because these outputs are a few tokens and structured, the main agent reads them directly — the sub-agent layer becomes unnecessary (direct-call architecture).
| Component | File | Notes |
|---|---|---|
| SOP decision tree | sop.yaml |
9 decision nodes + actions, 3 alarm families (spread anomaly, volume spike, feed health) |
| Exchange access layer | explore.py, backend.py |
Kraken + Binance REST, plus windowed reads over historical archives |
| MCP server | server.py |
FastMCP; live + historical data tools |
| Top-of-book recorder | recorder.py |
Daemon that recorded live snapshots for thin pairs |
| Main agent | main.py |
SOP tree walker with a pluggable evaluator — label replay, LLM, or tools |
| Baseline sub-agent | sub_agent.py |
Code-writing loop: LLM writes Python → executed in subprocess → output fed back → verdict. Max 3 attempts |
| Labeled ground truth | cases.yaml (25), holdout.yaml (5) |
Human-labeled from real historical data; conventions stated in every why |
| Tool-maker | pipeline.py |
One LLM call per node → permanent tool in tools/, registered in tools/registry.json |
| Reflector | reflector.py |
Ground-truth test + repair loop (never needed a repair round) |
| Runtime | runtime.py |
Tool-first evaluation with LLM fallback (missing tool / no_data / exception) |
| Evaluation | evaluate_all.py, charts.py |
The table and charts above |
- The baseline agent found a bug in my ground truth. My original N9 case placed its timestamp mid-gap of BTC's only >15-minute outage of the quarter — at which moment the last trade was only ~10.7 minutes old, so "has not traded in the last 15 minutes" was strictly false. The baseline correctly returned
falseagainst mytruelabel. The label was wrong; the timestamp was moved late into the gap. Labels are a spec, and specs have bugs. - Ambiguity dominates baseline error. Of the baseline's misses, none were capability failures — they were specification failures: the SOP says "significant price move" and "recent typical volume" without defining either, and the fresh LLM guessed conventions different from mine (previous-candle vs current-candle, 30-day vs 24-hour baselines). This is the paper's core argument reproduced independently: tool-making freezes the SME's conventions so they stop being re-guessed per call.
- Four labels pinned an ambiguous convention. From my N6 labels (0.18%/0.23% moves = not significant; +20%/+26.8% = significant), the tool-maker committed to a concrete |move| ≥ 5% bar — the paper's Appendix B.4 claim (N≈5 labels approach the ceiling) working live. The boundary between those extremes remains genuinely undefined: a hold-out candidate at −1.64% with 5.1× volume was dropped as unlabelable — even the SME declined to rule.
- Real data pushed back constantly. Kraken answers
XBTUSDrequests under the keyXXBTZUSD; every price arrives as a string; the trade-archive header declares 6 columns for 7 fields; Binance kline timestamps are microseconds; ETH's Kraken trade archive turned out to be Q1-only, which the verification pass caught before it produced a bogus hold-out label. None of this is in any SOP — which is exactly the argument for traces.
pip install -r requirements.txt
export ANTHROPIC_API_KEY=... # needed for the baseline / pipeline / fallback, not for frozen tools
python main.py # walker with label-replay evaluator (no LLM)
python sub_agent.py # baseline: LLM evaluator over all labeled cases (writes results/, traces/)
python pipeline.py # tool-maker: generates tools/ from traces + labels
python reflector.py # ground-truth test + repair loop
python runtime.py # tool-first runtime with LLM fallback
python evaluate_all.py # the results table
python charts.py # the chartsThe historical archives are not committed (312MB+). Sources: Kraken quarterly trade CSVs from Kraken's historical data exports, Binance monthly 1h-kline CSVs from Binance public data. Place them in data/Data-Dump-Binance-And-Kraken/ (see backend.py for expected filenames).
- The 25 training cases are also the tool-maker's inputs; the independent evidence is the 5-case hold-out (labeled post-generation, mostly unseen pairs). A larger temporal split is future work.
- Tool latency is file-scan-bound on the largest archive (N9 scans BTC's 5.75M-row CSV; 8–22 s cold). An indexed store would make it sub-millisecond; irrelevant to the architecture claim.
- N=25 labeled cases across 6 node families is small — deliberately, per the paper's own Appendix B.4 finding.
- §3.2 agent fine-tuning (LoRA distillation of a small production model): requires Amazon's production trajectories — not reproducible outside. Future work: LoRA on this pipeline's own trajectories.
- Appendix C label-free generation (voting + LLM judges instead of labels): planned as a Phase-2 extension.

