Skip to content

Repository files navigation

Tool-Making and Self-Evolving LLM Agents — on Real Market Data

Implementation of arXiv 2607.08010 (Kujanpää et al., Amazon Fulfillment Technologies & Robotics, Jul 2026), transplanted from Amazon's internal robotics-monitoring stack to real Kraken/Binance market data. No mocks — live REST endpoints, 312MB of historical trade archives, and human-labeled ground truth.

Results first

System Accuracy (training cases) Accuracy (hold-out) Median latency / node Tokens / node
Baseline (code-writing sub-agent) 22/25 — 16,717 ms ~7,084
Frozen tools (this pipeline) 25/25 5/5 1.74 ms 0

Accuracy Latency

  • ~9,600× median latency reduction on decision nodes, at zero tokens.
  • The hold-out set was labeled after tool generation; 4 of its 5 cases use pairs (ETH, SOL) absent from training entirely — the tools encoded the labeling conventions, not the examples.
  • The repair loop never fired: 6/6 tools froze at version 1 with filtered traces and convention-bearing labels as input.

Problem

LLM agents in production regenerate the same code for the same procedural steps on every request — the same instruction re-interpreted, the same schema re-discovered, similar code re-written. This repetition costs latency, tokens, and reliability (run-to-run variance). The paper replaces this loop: each repeated step is compiled once into a validated, versioned tool; at runtime the agent calls the tool, falling back to code generation only when a tool is missing or fails.

Method (as I understand it)

The procedure lives in an SOP (Standard Operating Procedure) — a human-written decision tree of nodes: observation nodes (did a metric cross its threshold?), root-cause nodes, constraint nodes (gates before expensive actions), and action nodes. A main agent walks this tree per alarm; in the baseline, each node is delegated to a code-writing sub-agent that queries backends via MCP.

The tool-making pipeline has three components:

  1. Data-collector sub-agent — runs the baseline code-writing agent on labeled cases for the node: writes code, hits errors, corrects itself, produces graded verdicts — and records everything into a trace.
  2. Tool-maker LLM — writes the permanent, deterministic tool for the node, conditioned on the SOP text + the labeled cases + the trace.
  3. Reflector LLM — tests the candidate tool against the full labeled set; on failures it diagnoses and rewrites (max 3 rounds), guarding against special-casing the failing examples.

A trace is the artifact the first component exists to produce. Every time the sub-agent's code hits a real API and gets corrected — a renamed response key, a string-typed price, a surprise header — that discovered configuration knowledge is captured. Without a trace, the tool-maker would re-hit every schema surprise at generation time; with it, the working code and the quirks it already absorbed are the raw material the tool is compiled from.

Tools return tiny, structured verdicts — true | false | no_data plus the observed value, threshold, and a short explanation. Because these outputs are a few tokens and structured, the main agent reads them directly — the sub-agent layer becomes unnecessary (direct-call architecture).

What was built

Component File Notes
SOP decision tree sop.yaml 9 decision nodes + actions, 3 alarm families (spread anomaly, volume spike, feed health)
Exchange access layer explore.py, backend.py Kraken + Binance REST, plus windowed reads over historical archives
MCP server server.py FastMCP; live + historical data tools
Top-of-book recorder recorder.py Daemon that recorded live snapshots for thin pairs
Main agent main.py SOP tree walker with a pluggable evaluator — label replay, LLM, or tools
Baseline sub-agent sub_agent.py Code-writing loop: LLM writes Python → executed in subprocess → output fed back → verdict. Max 3 attempts
Labeled ground truth cases.yaml (25), holdout.yaml (5) Human-labeled from real historical data; conventions stated in every why
Tool-maker pipeline.py One LLM call per node → permanent tool in tools/, registered in tools/registry.json
Reflector reflector.py Ground-truth test + repair loop (never needed a repair round)
Runtime runtime.py Tool-first evaluation with LLM fallback (missing tool / no_data / exception)
Evaluation evaluate_all.py, charts.py The table and charts above

Findings along the way

  1. The baseline agent found a bug in my ground truth. My original N9 case placed its timestamp mid-gap of BTC's only >15-minute outage of the quarter — at which moment the last trade was only ~10.7 minutes old, so "has not traded in the last 15 minutes" was strictly false. The baseline correctly returned false against my true label. The label was wrong; the timestamp was moved late into the gap. Labels are a spec, and specs have bugs.
  2. Ambiguity dominates baseline error. Of the baseline's misses, none were capability failures — they were specification failures: the SOP says "significant price move" and "recent typical volume" without defining either, and the fresh LLM guessed conventions different from mine (previous-candle vs current-candle, 30-day vs 24-hour baselines). This is the paper's core argument reproduced independently: tool-making freezes the SME's conventions so they stop being re-guessed per call.
  3. Four labels pinned an ambiguous convention. From my N6 labels (0.18%/0.23% moves = not significant; +20%/+26.8% = significant), the tool-maker committed to a concrete |move| ≥ 5% bar — the paper's Appendix B.4 claim (N≈5 labels approach the ceiling) working live. The boundary between those extremes remains genuinely undefined: a hold-out candidate at −1.64% with 5.1× volume was dropped as unlabelable — even the SME declined to rule.
  4. Real data pushed back constantly. Kraken answers XBTUSD requests under the key XXBTZUSD; every price arrives as a string; the trade-archive header declares 6 columns for 7 fields; Binance kline timestamps are microseconds; ETH's Kraken trade archive turned out to be Q1-only, which the verification pass caught before it produced a bogus hold-out label. None of this is in any SOP — which is exactly the argument for traces.

Running it

pip install -r requirements.txt
export ANTHROPIC_API_KEY=...   # needed for the baseline / pipeline / fallback, not for frozen tools

python main.py          # walker with label-replay evaluator (no LLM)
python sub_agent.py     # baseline: LLM evaluator over all labeled cases (writes results/, traces/)
python pipeline.py      # tool-maker: generates tools/ from traces + labels
python reflector.py     # ground-truth test + repair loop
python runtime.py       # tool-first runtime with LLM fallback
python evaluate_all.py  # the results table
python charts.py        # the charts

The historical archives are not committed (312MB+). Sources: Kraken quarterly trade CSVs from Kraken's historical data exports, Binance monthly 1h-kline CSVs from Binance public data. Place them in data/Data-Dump-Binance-And-Kraken/ (see backend.py for expected filenames).

Limitations

  • The 25 training cases are also the tool-maker's inputs; the independent evidence is the 5-case hold-out (labeled post-generation, mostly unseen pairs). A larger temporal split is future work.
  • Tool latency is file-scan-bound on the largest archive (N9 scans BTC's 5.75M-row CSV; 8–22 s cold). An indexed store would make it sub-millisecond; irrelevant to the architecture claim.
  • N=25 labeled cases across 6 node families is small — deliberately, per the paper's own Appendix B.4 finding.

Not reproduced (and why)

  • §3.2 agent fine-tuning (LoRA distillation of a small production model): requires Amazon's production trajectories — not reproducible outside. Future work: LoRA on this pipeline's own trajectories.
  • Appendix C label-free generation (voting + LLM judges instead of labels): planned as a Phase-2 extension.

About

LLM agent that compiles its own reasoning into permanent tools (arXiv 2607.08010, on real Kraken/Binance data).

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages