A multi-agent system that self-improves interpretable extraction rules (regexes) from labeled examples — with an honest, reproducible metric.
Most "self-improving agent" demos are unfalsifiable: they claim improvement without a number you can check. RSI-Agents is built around one metric that is never cheated — held-out F1. A team of agents evolves a regex from a deliberately weak starting point; the validation split is only ever used to report generalization, never to steer the search. The gain is real, and the exact curve reproduces from a seed.
It's a genuinely useful task, too: synthesizing readable extraction/validation rules from examples is everyday work in data cleaning, log parsing, and PII detection — and unlike a black-box classifier, the output is a regex a human can read, audit, and paste into code.
$ python3 -m rsi_agents.cli --all
email | baseline F1 0.714 -> final F1 1.000 (+0.286) ▁████████████████ | P=1.00 R=1.00 | ^(?:(?:[a-z]+\.[a-z]+@[a-z]+\.[a-z]+)|(?:[a-z]+@[a-z]+\.[a-z]+))$
hexcolor | baseline F1 0.286 -> final F1 0.667 (+0.381) ▁▁▅▇▇████████████ | P=0.83 R=0.56 | ^(?:(?:\#[a-z]+\d+...)|...)$
ipv4 | baseline F1 0.847 -> final F1 0.847 (+0.000) █████████████████ | P=0.73 R=1.00 | ^\d+\.\d+\.\d+\.\d+$
isodate | baseline F1 0.878 -> final F1 0.878 (+0.000) █████████████████ | P=0.78 R=1.00 | ^\d+\-\d+\-\d+$
phone | baseline F1 0.154 -> final F1 0.892 (+0.738) ▁▅███████████████ | P=1.00 R=0.81 | ^(?:(?:\d+\.\d+\.\d+)|(?:\d+\-\d+\-\d+)|(?:\(\d+\)\ \d+\.\d+)|(?:\(\d+\)\ \d+\-\d+))$
mean improvement +0.281 | mean final F1 0.857 over 5 tasks
(That is the verbatim default output — --all runs 16 generations, seed 0. No
hand-edited numbers; the hexcolor rule is abbreviated only for width.)
No API key, no dependencies — pure standard library. Set ANTHROPIC_API_KEY
and pass --llm to let Claude join the Proposer.
There are two loops, and I'm careful to claim only what each one delivers:
-
Object level — improving the rule (the real, verified result). Each generation the agents make the regex better: Propose → Verify → Critique → Mutate → Archive. Starting from a weak single-example template, the loop discovers generalizations and assembles multi-branch alternations on its own. This is where the measured held-out gains come from (mean +0.28 F1;
emailreaches a perfect rule;phoneclimbs from 0.15 → 0.89). -
Meta level — improving the search (an experiment, ablated honestly). A Meta agent watches the improvement curve and, on a plateau, escalates the search (wider beam, larger mutation budget). On these tasks that adaptive controller does not beat a well-tuned fixed budget — see the ablation below. I'm reporting that as a negative result rather than dressing it up. The controller is transparent and reproducible; it interpolates between a cheap and an expensive fixed budget rather than dominating either.
task | fixed-lo | fixed-hi | meta
------------------------------------------------------------------------
email | F1 1.000 ( 187 ev) | F1 1.000 ( 545 ev) | F1 1.000 ( 396 ev)
hexcolor | F1 0.744 ( 200 ev) | F1 0.768 ( 521 ev) | F1 0.737 ( 294 ev)
ipv4 | F1 0.872 ( 191 ev) | F1 0.872 ( 576 ev) | F1 0.872 ( 425 ev)
isodate | F1 0.880 ( 185 ev) | F1 0.880 ( 547 ev) | F1 0.880 ( 415 ev)
phone | F1 0.893 ( 187 ev) | F1 0.938 ( 449 ev) | F1 0.893 ( 298 ev)
------------------------------------------------------------------------
MEAN | F1 0.878 ( 190 ev) | F1 0.892 ( 528 ev) | F1 0.876 ( 366 ev)
ev = candidates evaluated (cost). Honest read: Meta starts identical to
fixed-lo and only escalates when stuck, so it never underperforms fixed-lo
by much, and it lands between the two budgets on both quality and cost. It does
not match fixed-hi's quality. Adaptive scheduling earning its keep would
need a smarter escape move (diversity injection / restart) — that's open work,
and the negative result is part of the point.
| Agent | Responsibility |
|---|---|
| Proposer | Seeds a weak gen-0 population by inducing templates from positive examples (optionally asks Claude). |
| Verifier | The only honest metric. Scores every candidate on train (selection) and held-out validation (reporting); rejects unsafe regexes. |
| Critic | Diagnoses the champion's false-positives/negatives into structured hints (too_loose, too_strict, needs_anchor). |
| Mutator | Breeds the next generation from elites via critic-guided mutation, crossover, and union (assembling alternations from simpler rules). |
| Archivist | Elitist top-k archive — the system's memory across generations. |
| Meta | Adaptive search controller (ablated above). |
- Validation never steers the search. Selection pressure uses train F1 only; the reported curve is the champion's held-out F1. Improvement is a real generalization gain, not val-set overfitting.
- gen-0 is deliberately weak (single-example templates). The agents must discover generalizations and multi-branch alternations themselves — the answer is never handed to them.
- Some tasks barely move, on purpose.
ipv4andisodatestart near the ceiling at gen-0 because a clean regex for "octet ∈ 0–255" or "month ∈ 1–12" doesn't exist — their residual precision gap (P≈0.73–0.78) is regex's honest expressive limit, not a bug. RSI-Agents shows where the method can't help. - The Meta controller is a reported negative result, not a headline win (see the ablation). The verified self-improvement is the object-level loop.
- Determinism. Every run reproduces from
(--seed, --data-seed); data generation uses a process-stable CRC32 hash (not Python's saltedhash()), verified byte-identical underPYTHONHASHSEED=random. - ReDoS safety.
is_safefails closed: it rejects exponential shapes ((a+)+,(a|a)+,(.*a){10}) through arbitrary nesting before any match runs. On the main thread a SIGALRM timer is armed as defense-in-depth; off the main thread (no SIGALRM) it refuses over-long inputs rather than match unguarded. A bad candidate scores 0; it never hangs the run.
git clone https://github.com/sugeerth/rsi-agents
cd rsi-agents
python3 -m rsi_agents.cli --all # all tasks (reproduces the table above)
python3 -m rsi_agents.cli --task email --verbose # watch one task evolve, see meta interventions
python3 -m rsi_agents.cli --task phone --no-meta # ablate the meta controller
python3 -m rsi_agents.cli --task ipv4 --llm # let Claude propose (needs ANTHROPIC_API_KEY)
python3 examples/ablation.py # reproduce the ablation table
python3 -m unittest discover -s tests -v # 28 tests, ~0.4semail, phone, ipv4, hexcolor, isodate — each generated from a hidden
ground-truth oracle with hard negatives (strings that look almost right), so
high F1 requires a genuinely discriminating rule.
┌── Proposer ──┐ weak templates from positives (gen 0 only)
▼ │
┌─ Verifier ◄────────┘ score on train (selection) + held-out val (report)
│ │
│ ▼
│ Critic diagnose champion's mistakes -> hints
│ │
│ ▼
│ Mutator mutate / crossover / union elites under hints
│ │
│ ▼
│ Archivist keep elite top-k (memory)
│ │
│ ▼
└─ Meta improving? hold effort. plateaued? escalate search.
rsi_agents/
datasets.py labeled tasks + hidden oracles + hard negatives
regex_ops.py induction, mutation/crossover/union, fail-closed safe matching
agents.py the six agents + scoring (precision/recall/F1)
engine.py the loop + RunReport (improvement curve, evaluations)
llm.py optional Claude backend (fails closed to local agents)
cli.py entry point + sparkline summary
tests/ datasets / regex_ops / engine / cli (28 tests)
examples/ run_all.py, ablation.py
MIT