Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

RSI-Agents

A multi-agent system that self-improves interpretable extraction rules (regexes) from labeled examples — with an honest, reproducible metric.

Most "self-improving agent" demos are unfalsifiable: they claim improvement without a number you can check. RSI-Agents is built around one metric that is never cheated — held-out F1. A team of agents evolves a regex from a deliberately weak starting point; the validation split is only ever used to report generalization, never to steer the search. The gain is real, and the exact curve reproduces from a seed.

It's a genuinely useful task, too: synthesizing readable extraction/validation rules from examples is everyday work in data cleaning, log parsing, and PII detection — and unlike a black-box classifier, the output is a regex a human can read, audit, and paste into code.

$ python3 -m rsi_agents.cli --all
email     | baseline F1 0.714 -> final F1 1.000 (+0.286) ▁████████████████ | P=1.00 R=1.00 | ^(?:(?:[a-z]+\.[a-z]+@[a-z]+\.[a-z]+)|(?:[a-z]+@[a-z]+\.[a-z]+))$
hexcolor  | baseline F1 0.286 -> final F1 0.667 (+0.381) ▁▁▅▇▇████████████ | P=0.83 R=0.56 | ^(?:(?:\#[a-z]+\d+...)|...)$
ipv4      | baseline F1 0.847 -> final F1 0.847 (+0.000) █████████████████ | P=0.73 R=1.00 | ^\d+\.\d+\.\d+\.\d+$
isodate   | baseline F1 0.878 -> final F1 0.878 (+0.000) █████████████████ | P=0.78 R=1.00 | ^\d+\-\d+\-\d+$
phone     | baseline F1 0.154 -> final F1 0.892 (+0.738) ▁▅███████████████ | P=1.00 R=0.81 | ^(?:(?:\d+\.\d+\.\d+)|(?:\d+\-\d+\-\d+)|(?:\(\d+\)\ \d+\.\d+)|(?:\(\d+\)\ \d+\-\d+))$

mean improvement +0.281 | mean final F1 0.857 over 5 tasks

(That is the verbatim default output — --all runs 16 generations, seed 0. No hand-edited numbers; the hexcolor rule is abbreviated only for width.)

No API key, no dependencies — pure standard library. Set ANTHROPIC_API_KEY and pass --llm to let Claude join the Proposer.

What "self-improvement" means here (precisely)

There are two loops, and I'm careful to claim only what each one delivers:

  1. Object level — improving the rule (the real, verified result). Each generation the agents make the regex better: Propose → Verify → Critique → Mutate → Archive. Starting from a weak single-example template, the loop discovers generalizations and assembles multi-branch alternations on its own. This is where the measured held-out gains come from (mean +0.28 F1; email reaches a perfect rule; phone climbs from 0.15 → 0.89).

  2. Meta level — improving the search (an experiment, ablated honestly). A Meta agent watches the improvement curve and, on a plateau, escalates the search (wider beam, larger mutation budget). On these tasks that adaptive controller does not beat a well-tuned fixed budget — see the ablation below. I'm reporting that as a negative result rather than dressing it up. The controller is transparent and reproducible; it interpolates between a cheap and an expensive fixed budget rather than dominating either.

Ablation (python3 examples/ablation.py, 5 seeds × 5 tasks)

task      |           fixed-lo |           fixed-hi |               meta
------------------------------------------------------------------------
email     | F1 1.000 ( 187 ev) | F1 1.000 ( 545 ev) | F1 1.000 ( 396 ev)
hexcolor  | F1 0.744 ( 200 ev) | F1 0.768 ( 521 ev) | F1 0.737 ( 294 ev)
ipv4      | F1 0.872 ( 191 ev) | F1 0.872 ( 576 ev) | F1 0.872 ( 425 ev)
isodate   | F1 0.880 ( 185 ev) | F1 0.880 ( 547 ev) | F1 0.880 ( 415 ev)
phone     | F1 0.893 ( 187 ev) | F1 0.938 ( 449 ev) | F1 0.893 ( 298 ev)
------------------------------------------------------------------------
MEAN      | F1 0.878 ( 190 ev) | F1 0.892 ( 528 ev) | F1 0.876 ( 366 ev)

ev = candidates evaluated (cost). Honest read: Meta starts identical to fixed-lo and only escalates when stuck, so it never underperforms fixed-lo by much, and it lands between the two budgets on both quality and cost. It does not match fixed-hi's quality. Adaptive scheduling earning its keep would need a smarter escape move (diversity injection / restart) — that's open work, and the negative result is part of the point.

The agent team

Agent Responsibility
Proposer Seeds a weak gen-0 population by inducing templates from positive examples (optionally asks Claude).
Verifier The only honest metric. Scores every candidate on train (selection) and held-out validation (reporting); rejects unsafe regexes.
Critic Diagnoses the champion's false-positives/negatives into structured hints (too_loose, too_strict, needs_anchor).
Mutator Breeds the next generation from elites via critic-guided mutation, crossover, and union (assembling alternations from simpler rules).
Archivist Elitist top-k archive — the system's memory across generations.
Meta Adaptive search controller (ablated above).

Honesty notes (read these)

  • Validation never steers the search. Selection pressure uses train F1 only; the reported curve is the champion's held-out F1. Improvement is a real generalization gain, not val-set overfitting.
  • gen-0 is deliberately weak (single-example templates). The agents must discover generalizations and multi-branch alternations themselves — the answer is never handed to them.
  • Some tasks barely move, on purpose. ipv4 and isodate start near the ceiling at gen-0 because a clean regex for "octet ∈ 0–255" or "month ∈ 1–12" doesn't exist — their residual precision gap (P≈0.73–0.78) is regex's honest expressive limit, not a bug. RSI-Agents shows where the method can't help.
  • The Meta controller is a reported negative result, not a headline win (see the ablation). The verified self-improvement is the object-level loop.
  • Determinism. Every run reproduces from (--seed, --data-seed); data generation uses a process-stable CRC32 hash (not Python's salted hash()), verified byte-identical under PYTHONHASHSEED=random.
  • ReDoS safety. is_safe fails closed: it rejects exponential shapes ((a+)+, (a|a)+, (.*a){10}) through arbitrary nesting before any match runs. On the main thread a SIGALRM timer is armed as defense-in-depth; off the main thread (no SIGALRM) it refuses over-long inputs rather than match unguarded. A bad candidate scores 0; it never hangs the run.

Install & run

git clone https://github.com/sugeerth/rsi-agents
cd rsi-agents

python3 -m rsi_agents.cli --all                  # all tasks (reproduces the table above)
python3 -m rsi_agents.cli --task email --verbose # watch one task evolve, see meta interventions
python3 -m rsi_agents.cli --task phone --no-meta  # ablate the meta controller
python3 -m rsi_agents.cli --task ipv4 --llm      # let Claude propose (needs ANTHROPIC_API_KEY)

python3 examples/ablation.py                     # reproduce the ablation table
python3 -m unittest discover -s tests -v         # 28 tests, ~0.4s

Tasks

email, phone, ipv4, hexcolor, isodate — each generated from a hidden ground-truth oracle with hard negatives (strings that look almost right), so high F1 requires a genuinely discriminating rule.

How a generation works

        ┌── Proposer ──┐  weak templates from positives (gen 0 only)
        ▼              │
  ┌─ Verifier ◄────────┘  score on train (selection) + held-out val (report)
  │     │
  │     ▼
  │   Critic            diagnose champion's mistakes -> hints
  │     │
  │     ▼
  │   Mutator           mutate / crossover / union elites under hints
  │     │
  │     ▼
  │   Archivist         keep elite top-k (memory)
  │     │
  │     ▼
  └─  Meta              improving? hold effort. plateaued? escalate search.

Project layout

rsi_agents/
  datasets.py   labeled tasks + hidden oracles + hard negatives
  regex_ops.py  induction, mutation/crossover/union, fail-closed safe matching
  agents.py     the six agents + scoring (precision/recall/F1)
  engine.py     the loop + RunReport (improvement curve, evaluations)
  llm.py        optional Claude backend (fails closed to local agents)
  cli.py        entry point + sparkline summary
tests/          datasets / regex_ops / engine / cli  (28 tests)
examples/       run_all.py, ablation.py

License

MIT

About

Multi-agent loop that evolves regex extraction rules from labeled examples, scored on held-out F1. Stdlib-only

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages