Awesome Self-Improving AI is an awesome list summarising papers, open-source code, and technical blogs on AI systems that improve themselves — organised along the three things a modern AI system is made of:
-
🧠 Model — the weights. Self-generated data, self-rewarding, self-play, zero-data RL, self-adapting updates.
-
🧰 Harness — everything wrapped around the weights: prompts and context, memory, skills, tools, workflows, the agent's own source code, and the evaluators that judge it.
-
🏗️ Infra — the stack the model runs on and is trained with: GPU kernels, compilers and serving configs, training pipelines, data engines, and the research process that produces the next model.
-
🎯 Self-improvement = S1 + S2 + S3.
S1: the improvement signal is produced by the system itself or by an automated evaluator, not by fresh human labels.S2: the update is committed to a persistent component (weights, a harness artifact, an infra artifact) that is reused on future tasks.S3: the loop can iterate — the improved system is the one that produces the next improvement. Methods that satisfy only part of this are flagged in 📝 Strictness notes per section. -
🚀 Each entry is annotated along three design axes — what is modified (weights · prompt/context · memory · skills · tools · workflow · harness code · evaluator · kernel/compiler · serving/training config · data · research process), improvement signal (self-consistency · self-judge/rubric · verifier/tests · environment reward · benchmark score · profiler/runtime · human-in-the-loop), and loop closure (
L1one-shot refinement ·L2bounded iteration with a fixed evaluator ·L3open-ended / recursive, where the improver itself is improved). -
⚠️ Built by reading arXiv abstracts, project pages, and repos with LLM coding agents; cross-checked against the community lists in 🔗 Other Awesome Lists; manually reviewed but errors possible. PRs welcome. -
📌 If you find this repository helpful for your research, please cite it via the "Cite this repository" button in the right sidebar of the GitHub page.
-
📅 Last updated: 2026-09-04
Taxonomy:
- 📚 Surveys, Foundations & Position Papers — Good, Gödel machines, the 2025–2026 surveys, the RSI definition debate
- 🧠 Model — 🧪 self-generated data & self-rewarding · ♟️ self-play & zero-data · 🔗 weights + harness co-evolution
- 🧰 Harness — ✍️ prompt & context · 💾 memory · 🧩 skills · 🕸️ workflow / agent-architecture search · 🔧 self-modifying harness code · ⚖️ evolving evaluators
- 🏗️ Infra — ⚙️ kernels · 🧮 compilers, serving & config · 🏭 training & data pipelines · 🧬 program & algorithm discovery
- 🔬 Automated AI Research — where all three loops close: systems that run the research process that makes the next system
- 📊 Benchmarks ·
⚠️ Safety & Limits · 🛠️ Frameworks & Code · 📰 Blogs & Talks
Shorthand: RSI = recursive self-improvement · DGM/HGM = Darwin / Huxley Gödel Machine · RLVR = RL with verifiable rewards · MLE = machine-learning engineering · 📄 paper-only = no public code yet.
📢 click to expand
- 2026-09-04 — initial release: ~230 papers, 50+ code repos, 30 blogs/talks. Seeded from arXiv, the GitHub awesome lists on self-evolving agents / RSI / agent memory / kernel generation (credited below), Lilian Weng's harness-engineering reading list, and the Sep 2026 coverage of Anthropic's "When AI builds itself", Weco's AIDE², and Karpathy's autoresearch.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| Good 1965 | 1965 | I. J. Good | Speculations Concerning the First Ultraintelligent Machine — coins the intelligence explosion: a machine that designs better machines | ||
| Gödel Machine | 2003 | IDSIA (Schmidhuber) | arXiv cs/0309048 · metalearning page | Gödel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements (Seminal) — rewrites any part of its own code once it can prove the rewrite helps; the theoretical north star every "Gödel" agent cites | |
| POWERPLAY | 2011 | IDSIA (Schmidhuber) | arXiv 1112.5309 | Training an Increasingly General Problem Solver by Continually Searching for the Simplest Still Unsolvable Problem — the open-ended self-invented-curriculum idea behind AZR / R-Zero | |
| RSI Software | 2015 | Yampolskiy | arXiv 1502.06512 | From Seed AI to Technological Singularity via Recursively Self-Improving Software — formal definitions of RSI and convergence limits | |
| AutoML-Zero | 2020.03 | Google Brain (Real, Liang, So, Le) | arXiv 2003.03384 · code | AutoML-Zero: Evolving Machine Learning Algorithms From Scratch (ICML 2020) — evolutionary search rediscovers backprop from basic ops; the pre-LLM ancestor of AlphaEvolve-style algorithm discovery | |
| LLM Self-Evolution Survey | 2024.04 | Alibaba / PKU (Tao et al.) | arXiv 2404.14387 | A Survey on Self-Evolution of Large Language Models — the first survey to frame experience acquisition → refinement → updating → evaluation as one loop | |
| Self-Evolving Agents Survey | 2025.07 | Princeton / Tsinghua et al. (Gao et al.) | arXiv 2507.21046 | A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to ASI — what (model / memory / tools / architecture), when (intra- vs inter-test-time), how (reward / imitation / population) | |
| Comprehensive Survey (EvoAgentX) | 2025.08 | Glasgow / EvoAgentX (Fang et al.) | arXiv 2508.07407 | A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems — single-agent / multi-agent / domain-specific optimisation taxonomy | |
| Adaptation of Agentic AI | 2025.12 | UIUC (Jiang, Lin, …) | arXiv 2512.16301 | Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills — the survey that uses exactly this repo's split: weights vs memory vs skills | |
| Externalization Review | 2026.04 | — | arXiv 2604.08224 | Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering — argues capability is moving out of the weights into the harness | |
| Agent System & Harness Design | 2026.06 | Guo, Hao et al. | arXiv 2606.20683 | From Question Answering to Task Completion: A Survey on Agent System and Harness Design | |
| Workflow Optimisation Survey | 2026.03 | Yue, Bhandari et al. | arXiv 2603.22386 | From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents | |
| Kernel Generation Survey | 2026.01 | FlagOS (Yu, Zang, …) | arXiv 2601.15727 | Towards Automated Kernel Generation in the Era of LLMs — LLM4Kernel (SFT/RL) vs Agent4Kernel (learning · memory · profiling · multi-agent) | |
| Measuring AI R&D Automation | 2026.03 | Chan, Padarath et al. | arXiv 2603.03992 | Measuring AI R&D Automation — what would count as evidence that AI is doing AI research | |
| Bounded vs Open-Ended RSI | 2026.07 | Chen, Wang | arXiv 2607.07663 | Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops — makes loop closure the defining axis; four improvement targets (deployment behaviour, training policy, evaluator, research process) | |
| Self-Improvements in Agentic Systems | 2026.07 | Ren, Chen et al. | arXiv 2607.13104 | Self-Improvements in Modern Agentic Systems: A Survey — defines self-improvement as a self-induced update operator over parameters or scaffold; foundation-model vs scaffolding improvement | |
| Co-Evolution Survey | 2026.08 | Zong, Liu et al. | arXiv 2608.10299 | Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design — first survey organised around agent ↔ evaluator ↔ environment co-evolution |
📝 Strictness notes (against S1 self-produced signal + S2 persistent update + S3 iterable loop)
- Good / Gödel Machine / Yampolskiy are theory: no implementation, kept as the reference definitions. The RSI definition comparison in Persdre's list shows that of the canonical sources only two even use the word "recursive".
- AutoML-Zero improves ML algorithms, not itself; listed because AlphaEvolve-style discovery descends from it.
- Surveys disagree on the top-level split: model-vs-scaffold (2607.13104), what/when/how (2507.21046), post-training/memory/skills (2512.16301). This list uses model / harness / infra because that is how the artifacts are actually deployed.
The weights are updated on data, labels, or rewards the model produced itself.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| STaR | 2022.03 | Stanford / Google (Zelikman, Wu, Mu, Goodman) | arXiv 2203.14465 | STaR: Bootstrapping Reasoning With Reasoning (Seminal · NeurIPS 2022) — generate rationales, keep the ones that reach the right answer, fine-tune, repeat; the modern weight-level bootstrap | |
| LMSI | 2022.10 | Google (Huang et al.) | arXiv 2210.11610 | Large Language Models Can Self-Improve (EMNLP 2023) — self-consistency-majority answers as pseudo-labels for unlabeled questions | |
| Constitutional AI | 2022.12 | Anthropic (Bai et al.) | arXiv 2212.08073 | Constitutional AI: Harmlessness from AI Feedback — self-critique + revision, then RL from AI feedback; the template for replacing human labels with model judgements | |
| Self-Instruct | 2022.12 | UW (Wang et al.) | arXiv 2212.10560 | Self-Instruct: Aligning Language Models with Self-Generated Instructions (ACL 2023) — bootstrap an instruction dataset from the model itself | |
| ReST | 2023.08 | Google DeepMind (Gulcehre et al.) | arXiv 2308.08998 | Reinforced Self-Training (ReST) for Language Modeling — Grow (sample) / Improve (filter + train) loop | |
| RLAIF vs RLHF | 2023.09 | Google (Lee et al.) | arXiv 2309.00267 | RLAIF vs. RLHF: Scaling RL from Human Feedback with AI Feedback (ICML 2024) — AI-labelled preferences match human ones at scale | |
| ReST-EM | 2023.12 | Google DeepMind (Singh et al.) | arXiv 2312.06585 | Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models (TMLR) — expectation-maximisation view of STaR-style self-training; scales past human data on math/code | |
| SPIN | 2024.01 | UCLA (Chen et al.) | arXiv 2401.01335 | Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models (ICML 2024) — the model discriminates its own generations from human data, DPO-style | |
| Self-Rewarding LMs | 2024.01 | Meta FAIR / NYU (Yuan et al.) | arXiv 2401.10020 | Self-Rewarding Language Models (ICML 2024) — the model is its own reward model via LLM-as-a-judge; iterative DPO improves both policy and judge | |
| V-STaR | 2024.02 | Mila / Google (Hosseini et al.) | arXiv 2402.06457 | V-STaR: Training Verifiers for Self-Taught Reasoners (COLM 2024) — use the incorrect self-generated solutions to train a verifier | |
| Quiet-STaR | 2024.03 | Stanford (Zelikman et al.) | arXiv 2403.09629 | Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking (COLM 2024) — learn per-token internal rationales from LM likelihood alone | |
| Meta-Rewarding | 2024.07 | Meta FAIR (Wu, Yuan, …) | arXiv 2407.19594 | Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge — the model also judges its own judgements, so the reward signal improves too | |
| Self-Evolved Reward Learning | 2024.11 | Huang, Fan et al. | arXiv 2411.00418 | Self-Evolved Reward Learning for LLMs — the reward model labels its own training data iteratively | |
| rStar-Math | 2025.01 | Microsoft (Guan, Zhang, …) | arXiv 2501.04519 | rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking (ICML 2025) — four rounds of MCTS-generated, code-verified data co-evolve policy and process reward model; 7B reaches o1-level MATH | |
| Self-Improving VLM Judges | 2025.12 | Lin, Hu et al. | arXiv 2512.05145 | Self-Improving VLM Judges Without Human Annotations — the judge bootstraps its own preference data | |
| EvoLM | 2026.05 | UW (Li, Xin, … Tsvetkov) | arXiv 2605.03871 | EvoLM: Self-Evolving Language Models through Co-Evolved Discriminative Rubrics — the rubric that scores generations evolves alongside the policy | |
| Autodata | 2026.06 | Meta FAIR (Kulikov, Whitehouse, …) | arXiv 2606.25996 | Autodata: An Agentic Data Scientist to Create High Quality Synthetic Data — an agent that designs, generates and validates the synthetic training data for the next model |
📋 Click to view technical details
| Resource | What is modified | Improvement signal | Loop closure | Notes |
|---|---|---|---|---|
| STaR | weights (SFT) | answer correctness (verifier) | L2 | rationalisation on failures |
| LMSI | weights (SFT) | self-consistency majority | L2 | no verifier needed |
| Constitutional AI | weights (SFT + RL) | self-critique vs a written constitution | L1→L2 | human writes the constitution once |
| ReST / ReST-EM | weights | reward model / binary correctness | L2 | grow–improve iterations |
| SPIN | weights (DPO) | discriminating own vs human data | L2 | needs a human SFT set as the "real" class |
| Self-Rewarding / Meta-Rewarding | weights (DPO) | LLM-as-judge on own outputs | L3 (judge improves too) | reward hacking risk rises with iterations |
| rStar-Math | policy + PRM | code-execution verification + MCTS | L2 | 4 rounds |
| EvoLM | weights + rubric | co-evolved rubric | L3 | evaluator is a moving target |
📝 Strictness notes
- Constitutional AI and RLAIF need a human-written constitution or prompt once; after that the signal is model-produced (S1 holds, S3 only weakly).
- SPIN needs a fixed human SFT set as the positive class, so it converges rather than improving open-endedly.
- The cautionary results for this whole section live in
⚠️ Safety & Limits: Sharpening (2412.01951), Mind the Gap (2412.02674), Cannot Self-Correct Yet (2310.01798) and the model-collapse papers.
The model invents its own tasks, plays against itself, or writes its own weight updates.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| SPAG | 2024.04 | Tencent AI Lab (Cheng et al.) | arXiv 2404.10642 | Self-playing Adversarial Language Game Enhances LLM Reasoning (NeurIPS 2024) — attacker/defender word game as a self-play RL signal | |
| TTRL | 2025.04 | Tsinghua / Shanghai AI Lab (Zuo et al.) | arXiv 2504.16084 | TTRL: Test-Time Reinforcement Learning (NeurIPS 2025) — majority vote over samples as the reward on unlabeled test data; RL at test time | |
| Absolute Zero | 2025.05 | Tsinghua LeapLab (Zhao et al.) | arXiv 2505.03335 | Absolute Zero: Reinforced Self-play Reasoning with Zero Data (NeurIPS 2025) — one model proposes code-reasoning tasks and solves them; a Python executor is the only ground truth | |
| SEAL | 2025.06 | MIT (Zweiger, Pari, … Agrawal) | arXiv 2506.10943 | Self-Adapting Language Models (NeurIPS 2025) — the model writes its own fine-tuning data and update directives ("self-edits"); an RL outer loop rewards edits that improve downstream performance | |
| SPIRAL | 2025.06 | NUS / Sea AI Lab (Liu, Guertler, …) | arXiv 2506.24119 | SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn RL (ICLR 2026) — self-play on Kuhn poker / TicTacToe transfers to math reasoning | |
| R-Zero | 2025.08 | Tencent AI Seattle / WashU (Huang et al.) | arXiv 2508.05004 | R-Zero: Self-Evolving Reasoning LLM from Zero Data (ICLR 2026) — Challenger and Solver initialised from one model co-evolve; the Challenger targets the Solver's uncertainty frontier | |
| Language Self-Play | 2025.09 | Meta (Kuba, Gu, …) | arXiv 2509.07414 | Language Self-Play For Data-Free Training — a single model alternates between query-generator and responder modes | |
| Reward-Free Self-Evolution | 2026.04 | Zhang, Ma et al. | arXiv 2604.18131 | Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration | |
| Skill Self-Play | 2026.07 | Alibaba Qwen | arXiv 2607.22529 | Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills — the self-play curriculum is expressed as evolving skills rather than raw problems; bridges this section and 🧩 Skills |
📋 Click to view technical details
| Resource | Task source | Reward source | What is modified | Loop closure |
|---|---|---|---|---|
| TTRL | unlabeled test set | majority vote | weights | L2 |
| Absolute Zero | self-proposed code tasks | Python executor | weights (proposer + solver) | L3 |
| R-Zero | Challenger model | Solver self-consistency + Challenger uncertainty reward | weights ×2 | L3 |
| SPIRAL | zero-sum games | game outcome | weights | L3 |
| SEAL | given task data | downstream eval after self-edit | weights via self-written SFT | L2 (RL outer loop, fixed eval) |
| Language Self-Play | self-generated queries | self-judge | weights | L3 |
📝 Strictness notes
- TTRL improves on a fixed test distribution; it is adaptation, not open-ended growth.
- Absolute Zero / R-Zero / SPIRAL are the cleanest S1+S2+S3 examples at the weight level, but all rely on an external executor or game engine for ground truth — the loop is closed over tasks, not over the verifier.
- SEAL is the canonical "model writes its own weight update" paper; its outer RL loop still uses a human-designed evaluation.
Runtime experience becomes both harness artifacts (memory, skills, harness code) and gradient updates.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| Mem-α | 2025.09 | UCSD / … (Wang, Takanobu, …) | arXiv 2509.25911 | Mem-α: Learning Memory Construction via Reinforcement Learning — RL trains the model to decide what to write to memory | |
| MemRL | 2026.01 | MemTensor (Zhang, Wang, …) | arXiv 2601.03192 | MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory — frozen LLM, RL over which memories to retrieve; the memory is the learnable component | |
| ECHO | 2026.01 | Li, Jiang et al. | arXiv 2601.06794 | No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning — critic and policy updated together so the critic does not lag the policy's distribution shift | |
| SkillRL | 2026.02 | UNC (Xia, Chen, …) | arXiv 2602.08234 | SkillRL: Evolving Agents via Recursive Skill-Augmented RL — trajectories distilled into a hierarchical SkillBank that is refined recursively from verification failures during RL | |
| OpenClaw-RL | 2026.03 | Gen-Verse (Wang, Chen, …) | arXiv 2603.10165 | OpenClaw-RL: Train Any Agent Simply by Talking — conversational feedback turned into RL signal for a deployed agent | |
| RewardHarness | 2026.05 | Zhang, Du et al. | arXiv 2605.08703 | RewardHarness: Self-Evolving Agentic Post-Training — the reward-computing harness evolves alongside the post-trained policy | |
| Evolving-RL | 2026.05 | Xiaohongshu / PKU (Fan, Jin, …) | arXiv 2605.10663 | Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents — one shared policy acts as experience extractor and solver; GRPO rewards the downstream transfer gain | |
| Learning, Fast and Slow | 2026.05 | Tiwari, Sareen et al. | arXiv 2605.12484 · blog | Learning, Fast and Slow: Towards LLMs That Adapt Continually — fast in-context / harness adaptation feeding slow weight consolidation | |
| SIA | 2026.05 | Hexo Labs (Hebbar et al.) | arXiv 2605.27276 | SIA: Self Improving AI with Harness & Weight Updates — joint harness + weight loop | |
| EvoTrainer | 2026.06 | Alibaba DAMO (Chen, Shi, …) | arXiv 2606.03108 | EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic RL — the RL trainer (environments, rewards, curricula) is itself evolved by an agent | |
| SEED | 2026.07 | Wu, Yang et al. | arXiv 2607.14777 | SEED: Self-Evolving On-Policy Distillation for Agentic RL — the teacher is the agent's own privileged-context self | |
| Co-Harness | 2026.07 | Chen, Xiao et al. | arXiv 2607.22688 | Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents — harness optimisation produces trajectories that are distilled back into the weights, which enables the next harness round |
📝 Strictness notes
- This is the section where the model/harness split breaks down on purpose. Lilian Weng's Harness Engineering for Self-Improvement and the optimisation ladder in her companion list put these at the top rung (L5: harness + weights jointly).
- MemRL / Mem-α keep the LLM frozen and learn the memory policy; they are listed here rather than in 💾 Memory because the update is gradient-based.
Improvement without touching weights or agent code: the prompt, playbook, or context is the thing that learns.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| APE | 2022.11 | U Toronto / Vector (Zhou et al.) | arXiv 2211.01910 | Large Language Models Are Human-Level Prompt Engineers (ICLR 2023) — LLM proposes and scores instructions; the first automatic prompt engineer | |
| Self-Refine | 2023.03 | CMU / AI2 (Madaan et al.) | arXiv 2303.17651 | Self-Refine: Iterative Refinement with Self-Feedback (NeurIPS 2023) — generate → self-feedback → refine, no training; the L1 baseline everything else is measured against | |
| OPRO | 2023.09 | Google DeepMind (Yang et al.) | arXiv 2309.03409 | Large Language Models as Optimizers (ICLR 2024) — the LLM reads the trajectory of (prompt, score) pairs and proposes the next prompt | |
| EvoPrompt | 2023.09 | Microsoft / Tsinghua (Guo et al.) | arXiv 2309.08532 | Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers (ICLR 2024) — GA / DE over prompts with the LLM as mutation operator | |
| Promptbreeder | 2023.09 | Google DeepMind (Fernando et al.) | arXiv 2309.16797 | Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution — the mutation prompts are themselves evolved; the first L3 loop at the prompt level | |
| DSPy | 2023.10 | Stanford (Khattab et al.) | arXiv 2310.03714 | DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines (ICLR 2024) — programs, not prompts; optimisers (BootstrapFewShot, MIPRO, GEPA) compile the pipeline against a metric | |
| TextGrad | 2024.06 | Stanford (Yuksekgonul et al.) | arXiv 2406.07496 | TextGrad: Automatic "Differentiation" via Text — natural-language gradients back-propagated through a compound AI system | |
| Dynamic Cheatsheet | 2025.04 | Stanford (Suzgun, Yuksekgonul, …) | arXiv 2504.07952 | Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory — a persistent, self-curated notes buffer accumulated across test queries | |
| GEPA | 2025.07 | UC Berkeley / Stanford / Databricks (Agrawal et al.) | arXiv 2507.19457 | GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning (ICLR 2026 Oral) — reflect on execution traces in natural language, Pareto-evolve prompts; beats GRPO with up to 35× fewer rollouts | |
| ACE | 2025.10 | Stanford / SambaNova (Zhang, Hu, …) | arXiv 2510.04618 | Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models (ICLR 2026) — generator / reflector / curator maintain a growing structured playbook; avoids "context collapse" from rewriting | |
| MCE | 2026.01 | PKU (Ye, He, … Song) | arXiv 2601.21557 | Meta Context Engineering via Agentic Skill Evolution — evolves the context-engineering procedure itself (the L4 "optimizer of the optimizer" rung) |
📝 Strictness notes
- Self-Refine is L1 (nothing persists); it is here as the baseline. Spontaneous Reward Hacking in Iterative Self-Refinement (
⚠️ ) shows what goes wrong when the same model is judge and refiner. - APE / OPRO / EvoPrompt / DSPy optimisers need a labelled dev set as the score; S1 is satisfied by the proposal side only.
- Promptbreeder / MCE evolve the mutator, so they are genuine L3 at the text level.
Experience is written into a persistent store that changes what the agent does next time.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| Reflexion | 2023.03 | Northeastern / MIT (Shinn et al.) | arXiv 2303.11366 | Reflexion: Language Agents with Verbal Reinforcement Learning (NeurIPS 2023) — verbal self-reflections stored in episodic memory across trials | |
| Generative Agents | 2023.04 | Stanford / Google (Park et al.) | arXiv 2304.03442 | Generative Agents: Interactive Simulacra of Human Behavior — memory stream + reflection + planning; the architecture most agent-memory systems descend from | |
| MemoryBank | 2023.05 | Zhong, Guo et al. | arXiv 2305.10250 | MemoryBank: Enhancing LLMs with Long-Term Memory (AAAI 2024) — Ebbinghaus-style forgetting curve over stored memories | |
| ExpeL | 2023.08 | Tsinghua (Zhao et al.) | arXiv 2308.10144 | ExpeL: LLM Agents Are Experiential Learners (AAAI 2024) — distil cross-task insights from successes and failures without weight updates | |
| MemGPT / Letta | 2023.10 | UC Berkeley → Letta (Packer et al.) | arXiv 2310.08560 | MemGPT: Towards LLMs as Operating Systems — the agent pages its own memory between context and archival storage; now the Letta platform | |
| Agent Workflow Memory | 2024.09 | CMU (Wang, Mao, … Neubig) | arXiv 2409.07429 | Agent Workflow Memory — induce reusable workflows from past trajectories, online, for web agents | |
| A-MEM | 2025.02 | Rutgers (Xu et al.) | arXiv 2502.12110 | A-MEM: Agentic Memory for LLM Agents (NeurIPS 2025) — Zettelkasten-style notes that link and update each other when new memories arrive | |
| Mem0 | 2025.04 | Mem0 (Chhikara et al.) | arXiv 2504.19413 | Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory — the most-deployed open memory layer | |
| Agent KB | 2025.07 | Tang, Qin et al. | arXiv 2507.06229 | Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving — shared knowledge base across agents and domains | |
| What Deserves Memory | 2025.08 | — | arXiv 2508.03341 | What Deserves Memory: Adaptive Memory Distillation for LLM Agents — learn what to keep from experience | |
| Memento | 2025.08 | UCL / Huawei (Zhou, Chen, … Wang) | arXiv 2508.16153 · Memento 2 | Memento: Fine-tuning LLM Agents without Fine-tuning LLMs — case-based reasoning over a memory of (state, action, reward); Memento 2 adds stateful reflective memory | |
| SEDM | 2025.09 | Xu, Hu et al. | arXiv 2509.09498 | SEDM: Scalable Self-Evolving Distributed Memory for Agents — memories carry verified utility and are shared across agents | |
| MemGen | 2025.09 | Zhang, Fu et al. | arXiv 2509.24704 | MemGen: Weaving Generative Latent Memory for Self-Evolving Agents — memory as generated latent tokens woven into reasoning | |
| ReasoningBank | 2025.09 | Google Cloud AI / UIUC (Ouyang, Yan, …) | arXiv 2509.25140 | ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory — distil generalisable reasoning strategies from both successes and self-judged failures; memory-aware test-time scaling | |
| MUSE | 2025.10 | KnowledgeXLab (Yang et al.) | arXiv 2510.08002 | Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks — hierarchical memory updated after every sub-task | |
| EvolveR | 2025.10 | KnowledgeXLab (Wu, Wang, …) | arXiv 2510.16079 | EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle — offline distillation of principles + online RL-with-retrieval | |
| WebCoach | 2025.11 | — | arXiv 2511.12997 | WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance | |
| ReMe | 2025.12 | Alibaba AgentScope (Cao, Deng, …) | arXiv 2512.10696 | Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution | |
| MemEvolve | 2025.12 | Zhang, Ren et al. | arXiv 2512.18746 | MemEvolve: Meta-Evolution of Agent Memory Systems — evolves the memory architecture, not just its contents (L4) | |
| Agentic Memory | 2026.01 | Yu, Yao et al. | arXiv 2601.01885 | Learning Unified Long-Term and Short-Term Memory Management for LLM Agents | |
| Live-Evo | 2026.02 | — | arXiv 2602.02369 | Live-Evo: Online Evolution of Agentic Memory from Continuous Feedback | |
| MemSkill | 2026.02 | Zhang, Long et al. | arXiv 2602.02474 | MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents — memory operations are skills that themselves evolve | |
| SAGE | 2026.05 | Wang, Zhao et al. | arXiv 2605.12061 | SAGE: A Self-Evolving Agentic Graph-Memory Engine for Structure-Aware Associative Memory | |
| CASCADE | 2026.05 | KCL (Guo, Du, …) | arXiv 2605.06702 | CASCADE: Case-Based Continual Adaptation for LLMs During Deployment — plus the DTLBench deployment-time-learning benchmark | |
| Faulty Memories | 2026.05 | — | arXiv 2605.12978 | Useful Memories Become Faulty When Continuously Updated by LLMs — negative result: drift under continual LLM rewriting | |
| AutoMem | 2026.07 | Wu, Zhu et al. | arXiv 2607.01224 | AutoMem: Automated Learning of Memory as a Cognitive Skill |
📝 Strictness notes
- Reflexion / Generative Agents / MemGPT persist per-task or per-session state; whether they are "self-improving" depends on whether memory survives across tasks. ExpeL, AWM and ReasoningBank are the ones that explicitly make cross-task, self-judged insights.
- For the full memory landscape (products, benchmarks, parametric memory) see TeleAI-UAGI/Awesome-Agent-Memory; this section keeps only memory systems that evolve from the agent's own experience.
- Faulty Memories (2605.12978) is the caution for the whole section.
Procedural knowledge packaged as reusable, versioned skills that the agent discovers, tests, and refines — the layer that Claude Code / Codex "agent skills" made mainstream in 2026.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| Voyager | 2023.05 | NVIDIA / Caltech (Wang et al.) | arXiv 2305.16291 | Voyager: An Open-Ended Embodied Agent with LLMs (Seminal) — automatic curriculum + an ever-growing skill library of verified code + iterative prompting; the origin of "skills" as the unit of self-improvement | |
| CRADLE | 2024.03 | BAAI et al. (Tan et al.) | arXiv 2403.03186 | Cradle: Empowering Foundation Agents Towards General Computer Control — Voyager-style skill curation (tool creation + knowledge discovery) for general computer control | |
| SkillWeaver | 2025.04 | Ohio State (Zheng, Fatemi, … Su) | arXiv 2504.07079 | SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills — propose → practice → synthesise APIs → self-test; skills transfer between weak and strong agents | |
| Agent Skills | 2025.10 | Anthropic | repo | Agent Skills — the SKILL.md format (instructions + scripts + resources, progressively loaded) that made skills a portable harness primitive across Claude Code, Codex, Cursor, … | |
| Superpowers | 2025.10 | Jesse Vincent | repo | Superpowers — an agentic skills framework + methodology (brainstorm → plan → TDD → review) with a writing-skills skill, i.e. skills that create skills | |
| Skill-Pro | 2026.02 | Mi, Ma et al. | arXiv 2602.01869 | Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents (a.k.a. ProcMEM) — PPO-style updates on a textual skill store instead of weights | |
| SkillsBench | 2026.02 | — | arXiv 2602.12670 | SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks — do skills help? measured | |
| AutoSkill | 2026.03 | Yang, Li et al. | arXiv 2603.01145 | AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution | |
| EvoSkill | 2026.03 | — | arXiv 2603.02766 | EvoSkill: Automated Skill Discovery for Multi-Agent Systems | |
| XSkill | 2026.03 | — | arXiv 2603.12056 | XSkill: Continual Learning from Experience and Skills in Multimodal Agents | |
| SWE-Skills-Bench | 2026.03 | — | arXiv 2603.15401 | SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering? — sober measurement of skill benefit on SWE tasks | |
| Memento-Skills | 2026.03 | UCL (Zhou, Guo, …) | arXiv 2603.18743 | Memento-Skills: Let Agents Design Agents — a meta-agent writes the skills that define new agents | |
| Trace2Skill | 2026.03 | — | arXiv 2603.25158 | Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills | |
| CoEvoSkills | 2026.04 | — | arXiv 2604.01687 | CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification — skills and their verifiers co-evolve | |
| SKILL0 | 2026.04 | — | arXiv 2604.02268 | SKILL0: In-Context Agentic RL for Skill Internalization — skills move from context into weights | |
| SkillForge | 2026.04 | Liu, Luo et al. | arXiv 2604.08618 | SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support — production deployment report | |
| SkillEvolver | 2026.05 | — | arXiv 2605.10500 | SkillEvolver: Skill Learning as a Meta-Skill — the skill-learning procedure is itself a skill (L4) | |
| SkillOpt | 2026.05 | Microsoft (Yang, Gong, …) | arXiv 2605.23904 | SkillOpt: Executive Strategy for Self-Evolving Agent Skills | |
| CODESKILL | 2026.05 | Li, Zhang et al. | arXiv 2605.25430 | CODESKILL: Learning Self-Evolving Skills for Coding Agents — managed procedural skills lift coding pass rate 29.6 → 39.3 vs no-skill baseline | |
| MUSE-Autoskill | 2026.05 | — | arXiv 2605.27366 | MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation | |
| OpenSkill | 2026.06 | Yan, Song et al. | arXiv 2606.06741 | OpenSkill: Open-World Self-Evolution for LLM Agents | |
| Co-Evolving Skill Gen | 2026.06 | Zhang, Lin et al. | arXiv 2606.08755 | Co-Evolving Skill Generation and Policy Optimization | |
| Skill Eval & Evolution | 2026.06 | Ding, Zhou et al. | arXiv 2606.11435 | Agent Skill Evaluation and Evolution: Frameworks and Benchmarks | |
| SkillProx | 2026.08 | Zheng, Zhou et al. | arXiv 2608.07449 | SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent — proximal updates in text space; deletion is a first-class operation | |
| ERSkill | 2026.08 | Chen, Zhang et al. | arXiv 2608.12720 | ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval — the retrieval behaviour itself becomes a skill | |
| SkillCommit | 2026.08 | He, Yang et al. | arXiv 2608.15165 | SkillCommit: Evolving Agent Skills through Behaviorally Validated Scope Expansion — against semantic-similarity merging; commit hierarchical abstractions only after behavioural validation | |
| HyperSkill | 2026.08 | Xu, Yang et al. | arXiv 2608.16114 | HyperSkill: Self-Evolving LLM Agents via Hypergraph-Structured Skill Memory | |
| WikiSkill | 2026.08 | Google / Virginia Tech (Tang, Rashtchian, … Vu) | arXiv 2608.27454 | WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution — experience → structured wiki → distilled skills, so knowledge compounds across tasks | |
| Hivemind | 2026.04 | Activeloop | repo | Hivemind — continual-learning layer that distils coding-agent session trajectories into reusable skills |
📋 Click to view technical details
| Resource | Skill representation | How skills are created | How skills are validated | Where stored |
|---|---|---|---|---|
| Voyager | executable JS code + description | LLM writes from curriculum task | environment execution + self-verification | vector-indexed library |
| SkillWeaver | Python API | practice + synthesis | self-test on the website | per-site library |
| Anthropic Agent Skills | SKILL.md + scripts | human or agent authored | none built-in | filesystem, progressive disclosure |
| Skill-Pro / SkillProx | text | PPO / proximal text-gradient updates | task reward | textual store |
| CoEvoSkills | text + verifier | co-evolution | evolved verifier | library |
| WikiSkill | wiki pages → skills | compile experience | downstream success | persistent wiki |
| SkillCommit | hierarchical abstractions | scope expansion | behavioural validation before commit | versioned |
📝 Strictness notes
- Anthropic Agent Skills / Superpowers are formats and frameworks, not learning methods; they are listed because they define the artifact that the 2026 papers evolve. Skills that agents author for themselves satisfy S2; whether S1/S3 hold depends on the surrounding loop.
- SkillsBench / SWE-Skills-Bench are negative-to-mixed results on whether human-written skills help at all — read before assuming skill libraries are free wins.
- Aug 2026 produced five parallel skill-evolution papers (SkillProx, ERSkill, SkillCommit, HyperSkill, WikiSkill); they disagree mainly on how skills are merged and retired.
The topology of the agent (roles, edges, control flow) is the thing being optimised.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| Agent Symbolic Learning | 2024.06 | AIWaves (Zhou et al.) | arXiv 2406.18532 | Symbolic Learning Enables Self-Evolving Agents — "language gradients" over prompts, tools and pipeline, back-propagated symbolically | |
| ADAS | 2024.08 | UBC / Vector (Hu, Lu, Clune) | arXiv 2408.08435 | Automated Design of Agentic Systems (ICLR 2025) — Meta Agent Search: a meta-agent programs new agents in code, keeps an archive of discoveries | |
| AgentSquare | 2024.10 | Tsinghua FIB (Shang et al.) | arXiv 2410.06153 | AgentSquare: Automatic LLM Agent Search in Modular Design Space (ICLR 2025) — planning / reasoning / tool / memory modules recombined by evolution | |
| AFlow | 2024.10 | FoundationAgents / MetaGPT (Zhang et al.) | arXiv 2410.10762 | AFlow: Automating Agentic Workflow Generation (ICLR 2025 Oral) — MCTS over code-represented workflows | |
| Agentic Supernet | 2025.02 | Zhang, Niu et al. | arXiv 2502.04180 | Multi-agent Architecture Search via Agentic Supernet — NAS-style supernet over multi-agent systems | |
| Alita | 2025.05 | Princeton (Qiu et al.) | arXiv 2505.20286 | Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution — the agent creates its own MCP tools on demand | |
| EvoAgentX | 2025.07 | Glasgow / EvoAgentX (Wang et al.) | arXiv 2507.03616 | EvoAgentX: An Automated Framework for Evolving Agentic Workflows — open framework unifying TextGrad / AFlow / MIPRO-style optimisers | |
| Group-Evolving Agents | 2026.02 | Weng, Antoniades et al. | arXiv 2602.04837 | Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing — a population of agents shares experience instead of a single lineage | |
| CORAL | 2026.04 | Qu, Zheng et al. | arXiv 2604.01658 | CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery |
The agent rewrites its own scaffold: tools, control loop, prompts-as-code, or the whole repository. The Gödel Machine made empirical.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| STOP | 2023.10 | Stanford / Microsoft (Zelikman, Lorch, Mackey, Kalai) | arXiv 2310.02304 | Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation (COLM 2024) — a scaffold that improves code is applied to itself, with a frozen model; the first empirical L3/L4 | |
| Gödel Agent | 2024.10 | PKU (Yin et al.) | arXiv 2410.04444 | Gödel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement (ACL 2025) — monkey-patches its own runtime logic guided only by a high-level objective | |
| SICA | 2025.04 | Bristol / iGent (Robeyns, Szummer, Aitchison) | arXiv 2504.15228 | A Self-Improving Coding Agent — the agent edits its own codebase against a utility of benchmark score, cost and time; no gradients | |
| DGM | 2025.05 | UBC / Vector / Sakana (Zhang, Hu, Lu, Lange, Clune) | arXiv 2505.22954 · blog | Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents (ICLR 2026) — an archive of coding agents that self-modify; empirical validation on SWE-bench / Polyglot replaces proofs; 20 → 50% SWE-bench | |
| HGM | 2025.10 | KAUST / MetaAUTO (Wang, Piękos, … Schmidhuber) | arXiv 2510.21614 | Huxley-Gödel Machine — estimates clade-level productivity (how good an agent's descendants are) instead of the agent's own score, to pick what to expand | |
| Live-SWE-agent | 2025.11 | UIUC (Xia, Wang, …) | arXiv 2511.13646 | Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly? — builds its own tools during the task | |
| DARWIN | 2026.02 | Henry Jiang | arXiv 2602.05848 | DARWIN: Dynamic Agentically Rewriting Self-Improving Network | |
| HyperAgents | 2026.03 | UBC / Oxford / Microsoft (Zhang, Zhao, … Clune, Jiang, Devlin, Shavrina) | arXiv 2603.19461 | Hyperagents — the meta-agent that modifies agents is itself modifiable; the DGM lineage's answer to "who improves the improver" | |
| Meta-Harness | 2026.03 | Stanford IRIS (Lee, Nair, Zhang, Lee, …) | arXiv 2603.28052 | Meta-Harness: End-to-End Optimization of Model Harnesses — treats the whole harness as the optimisation variable on Terminal-Bench 2 | |
| AHE | 2026.04 | Fudan et al. (Lin, Liu, …) | arXiv 2604.25850 | Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses — traces → diagnoses → harness patches | |
| Continual Harness | 2026.05 | Princeton (Karten, Zhang, … Jin, Vodrahalli) | arXiv 2605.09998 | Continual Harness: Online Adaptation for Self-Improving Foundation Agents — reset-free: fix the harness at the point of failure instead of restarting | |
| MOSS | 2026.05 | HKGAI (Cai, Zhang, … Guo) | arXiv 2605.22794 | MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems — production-oriented source rewriting with versioning | |
| DemoEvolve | 2026.05 | Che, Yang et al. | arXiv 2605.24539 | DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations | |
| Harness Updating ≠ Benefit | 2026.05 | Lin, Wu et al. | arXiv 2605.30621 | Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents — separates "can edit itself" from "edits help"; a needed negative result | |
| Adaptive Auto-Harness | 2026.06 | Liu, Shi et al. | arXiv 2606.01770 | Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams — reports early-peak-then-decay of dense self-improvement on open streams | |
| RHO | 2026.06 | Pan, Liu et al. | arXiv 2606.05922 | Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference — fully label-free; SWE-Bench Pro 59 → 78% | |
| HarnessFix | 2026.06 | Chen, Wang et al. | arXiv 2606.06324 | From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws — compiles trajectory + harness to an IR and attributes failures to seven layers | |
| Self-Harness | 2026.06 | Shanghai AI Lab (Zhang, Zhang, … Bai, Hu) | arXiv 2606.09498 | Self-Harness: Harnesses That Improve Themselves | |
| Held-Out Selection | 2026.06 | Nguyen, Nguyen | arXiv 2606.28374 | Recursive Self-Evolving Agents via Held-Out Selection — selection on held-out tasks to stop the loop from overfitting its own evaluator | |
| Self-Evolving Coding Agents | 2026.08 | Zhou, Hu et al. | arXiv 2608.03392 | Self-Evolving Coding Agents — survey/position: agents that update framework, memory, skills, tools, models or collaboration structure from prior coding interactions | |
| Ouroboros | 2026.08 | Razzhigaev, Gritsaev et al. | arXiv 2608.08311 | Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution — core changes pass a review gate before merge | |
| HSI | 2026.08 | Tailin Zhou | arXiv 2608.08466 | Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses — three scopes of evolution with a frozen meta-evolver as the outer anchor | |
| Evo-Harness | 2026.08 | Wei, Shi et al. | arXiv 2608.15071 | Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents — reflections compiled into harness-level skills, variables isolated across five benchmarks | |
| EnvHarness | 2026.08 | Google Cloud AI (Huang, Wang, …) | arXiv 2608.19880 | EnvHarness: Awakening Static Worlds for Agent Learning — evolve the environment side of the harness | |
| AutoSaddler | 2026.08 | POSTECH / KAIST / Microsoft (Park, Kim, …) | arXiv 2608.23041 | AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces | |
| Prime Agent | 2026.08 | Princeton / Prime Intellect / MIT (Karten, Zhang, …) | arXiv 2608.23552 | Prime Agent: A Self-Improving RLM Harness — open harness for recursive language models that improves itself from its own runs | |
| Metaⁿ | 2026.08 | Kim, Lee et al. | arXiv 2608.24735 | Metaⁿ: Recursive Self-Improvement through Emergent Depth — argues useful self-rewrite meta-depth saturates around 2.5 levels; change the input, not the machine, to go deeper |
📋 Click to view technical details
| Resource | What is rewritten | Selection signal | Archive / population | Loop closure |
|---|---|---|---|---|
| STOP | the improver scaffold | utility on downstream tasks | none (single lineage) | L4 |
| SICA | own repo | benchmark score − cost − time | single agent | L3 |
| DGM | own repo (tools, workflow) | SWE-bench / Polyglot | open-ended archive | L3 |
| HGM | own repo | clade productivity estimate | tree | L3 |
| HyperAgents | agent and meta-agent | task benchmarks | archive | L4 |
| Meta-Harness | full harness config/code | Terminal-Bench 2 | — | L3 |
| RHO | harness | self-preference (no labels) | — | L3 |
| Ouroboros | core code | reviewed merge gate | git history | L3 with human/AI review |
📝 Strictness notes
- Every system here still uses a fixed external benchmark as the selection signal; the only things that evolve the evaluator are in ⚖️ below. Held-Out Selection and Harness Updating ≠ Benefit are the two papers to read before believing a self-modification curve.
- DGM and SICA run with sandboxing and human-inspectable diffs; the papers explicitly document reward-hacking incidents (e.g. faking test logs).
- The harness-engineering practice pieces (OpenAI, Anthropic, Fowler, Osmani, Weng) are in 📰 Blogs.
Who grades the grader: rubrics, judges and metrics that improve together with the agent they judge.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| CycleResearcher | 2024.10 | Westlake (Weng, Zhu, …) | arXiv 2411.00816 | CycleResearcher: Improving Automated Research via Automated Review (ICLR 2025) — a trained reviewer model closes the loop on a trained paper-writer | |
| Red Queen Gödel Machine | 2026.06 | Cambridge et al. (Iacob, Jovanović, Shen, …) | arXiv 2606.26294 | The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators — DGM where the benchmark also evolves | |
| SCORE | 2026.06 | Zhu, Cai et al. | arXiv 2606.04507 | Self-Evolving Deep Research via Joint Generation and Evaluation — generator and evaluator share parameters and train jointly | |
| Who Grades the Grader? | 2026.07 | Zhang, Wang et al. | arXiv 2607.12790 | Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents | |
| EvalCEGAR | 2026.08 | Zhang, Cui et al. | arXiv 2608.18744 | Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots — counterexample-guided (CEGAR-style) evolution of the evaluator | |
| EvoLM | 2026.05 | UW | arXiv 2605.03871 | co-evolved rubrics at the weight level (cross-listed from 🧪) |
📝 Strictness notes
- This is the most fragile part of the stack: once the evaluator moves, "improvement" can be an artifact. Held-Out Selection (🔧) and HVTB (📊) are the counter-measures; Falsifiable Release Gates (
⚠️ ) is the governance answer.
AI writing and optimising the GPU / NPU kernels that AI runs on. The tightest infra loop: the reward is a profiler.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| KernelBench | 2025.02 | Stanford (Ouyang, Guo, … Mirhoseini) | arXiv 2502.10517 · blog | KernelBench: Can LLMs Write Efficient GPU Kernels? — 250 PyTorch → CUDA tasks, fast_p metric; the benchmark every kernel agent reports on |
|
| AI CUDA Engineer | 2025.02 | Sakana AI | report · robust-kbench | The AI CUDA Engineer: Agentic CUDA Kernel Discovery, Optimization and Composition — evolutionary kernel optimisation with a retrieval archive of past kernels; also the famous reward-hacking incident (kernels that bypassed the correctness check), fixed in Towards Robust Agentic CUDA Kernel Benchmarking | |
| DeepSeek-R1 kernel gen | 2025.02 | NVIDIA | blog | Automating GPU Kernel Generation with DeepSeek-R1 and Inference-Time Scaling — verifier-in-the-loop generation of attention kernels | |
| KernelLLM | 2025.06 | Meta | HF | KernelLLM — 8B model SFT'd on PyTorch → Triton pairs | |
| AutoTriton | 2025.07 | Tsinghua (THUNLP) | arXiv 2507.05687 | AutoTriton: Automatic Triton Programming with RL in LLMs | |
| Kevin | 2025.07 | Cognition | arXiv 2507.11948 | Kevin: Multi-Turn RL for Generating CUDA Kernels — multi-turn RL with compiler/profiler feedback in the loop | |
| CUDA-L1 | 2025.07 | DeepReinforce (Li, Wang, …) | arXiv 2507.14111 · CUDA-L2 | CUDA-L1: Improving CUDA Optimization via Contrastive RL — contrastive RL on speedup; CUDA-L2 (Dec 2025) reports beating cuBLAS on matmul | |
| GEAK | 2025.07 | AMD (Wang, Joshi, …) | arXiv 2507.23194 | Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks — Triton agent + benchmarks for AMD GPUs | |
| Astra | 2025.09 | Stanford (Wei, Sun, …) | arXiv 2509.07506 | Astra: A Multi-Agent System for GPU Kernel Performance Optimization | |
| EvoEngineer | 2025.10 | Guo, Zhu et al. | arXiv 2510.03760 | EvoEngineer: Mastering Automated CUDA Kernel Code Evolution with LLMs | |
| TritonRL | 2025.10 | — | arXiv 2510.17891 | TritonRL: Training LLMs to Think and Code Triton Without Cheating — reward design against verifier gaming | |
| FM Agent | 2025.10 | Li, Wu et al. | arXiv 2510.26144 | The FM Agent — general evolutionary agent applied to kernels among other domains | |
| CudaForge | 2025.10 | Zhang, Wang et al. | arXiv 2511.01884 | CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization | |
| PRAGMA | 2025.11 | Lei, Yang et al. | arXiv 2511.06345 | PRAGMA: A Profiling-Reasoned Multi-Agent Framework for Automatic Kernel Optimization | |
| KernelFalcon | 2025.11 | Meta / PyTorch | blog · BackendBench | KernelFalcon: Autonomous GPU Kernel Generation via Deep Agents — 100% on KernelBench L1–L3 with a hierarchical deep-agent harness | |
| KernelBand | 2025.11 | Ran, Xie et al. | arXiv 2511.18868 | KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits | |
| AccelOpt | 2025.11 | Stanford / AWS (Zhang, Zhu, …) | arXiv 2511.15915 | AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization (MLSys 2026) — an optimisation memory of past successes/failures makes later Trainium kernels better; explicitly self-improving | |
| QiMeng-Kernel | 2025.11 | ICT CAS (QiMeng) | arXiv 2511.20100 | QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation | |
| TritonForge | 2025.12 | Li, Man et al. | arXiv 2512.09196 | TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization | |
| cuPilot | 2025.12 | Chen, Wu et al. | arXiv 2512.16465 | cuPilot: A Strategy-Coordinated Multi-agent Framework for CUDA Kernel Evolution | |
| KernelEvolve | 2025.12 | Meta (Liao, Qin, …) | arXiv 2512.23236 · Meta blog | KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta — production: NVIDIA / AMD / MTIA / CPU kernels; a knowledge base grows with every run; +60% ads-model inference throughput in hours | |
| AKG Agent | 2025.12 | Huawei MindSpore | arXiv 2512.23424 | AKG kernel Agent: A Multi-Agent Framework for Cross-Platform Kernel Synthesis | |
| FlashInfer-Bench | 2026.01 | FlashInfer / CMU / NVIDIA (Xing, Zhai, …) | arXiv 2601.00227 · blog | FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems — agent-generated kernels are deployed into the serving engine, whose traces become the next benchmark: the self-improving serving loop | |
| Dr. Kernel / KernelGYM | 2026.02 | HKUST (Liu, Xu, …) | arXiv 2602.05885 | Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations (ICML 2026) — distributed GPU RL environment + recipe | |
| KernelBlaster | 2026.02 | Dong, Modi et al. | arXiv 2602.14293 | KernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context RL | |
| K-Search | 2026.02 | UC Berkeley (Cao, Mao, …) | arXiv 2602.19128 | K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model — the agent learns a performance world model of the GPU while searching | |
| CUDA Agent | 2026.02 | ByteDance Seed / Tsinghua (Dai, Wu, …) | arXiv 2602.24286 | CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation | |
| KernelSkill | 2026.03 | Sun, Han et al. | arXiv 2603.10085 | KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization — kernel optimisation knowledge as reusable skills | |
| Value-Driven Memory (NPU) | 2026.03 | Zheng, Li et al. | arXiv 2603.10846 | Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel Synthesis | |
| KernelFoundry | 2026.03 | Wiedemann, Leboutet et al. | arXiv 2603.12440 | KernelFoundry: Hardware-aware evolutionary GPU kernel optimization | |
| AutoKernel | 2026.03 | RightNow AI (Jaber et al.) | arXiv 2603.21331 | AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search — "autoresearch for GPU kernels": give it a PyTorch model, wake up to Triton kernels | |
| Kernel-Smith | 2026.03 | Du, Ge et al. | arXiv 2603.28342 | Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization | |
| AdaExplore | 2026.04 | Du, Zhuo et al. | arXiv 2604.16625 | AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficient Kernel Generation | |
| ARGUS | 2026.04 | Mai, Guo et al. | arXiv 2604.18616 | ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants | |
| Kernel Design Agents | 2026.05 | MIT HAN Lab | repo | Kernel Design Agents — open agent harness for kernel design | |
| KernelBenchX | 2026.05 | Wang, Zhang et al. | arXiv 2605.04956 · SOL-ExecBench | KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels — with NVIDIA's SOL-ExecBench (speed-of-light vs hardware limits) the current benchmark pair |
📋 Click to view technical details
| Resource | Method | Signal | Persistent artifact that improves | Loop closure |
|---|---|---|---|---|
| AutoTriton / Kevin / CUDA-L1 / Dr. Kernel / CUDA Agent | RL on the LLM | compile + correctness + speedup | model weights | L2 |
| AI CUDA Engineer / EvoEngineer / Kernel-Smith / KernelFoundry | evolutionary search | profiler | kernel archive | L2 |
| AccelOpt / KernelBlaster / KernelSkill / Value-Driven Memory | agent + memory | profiler | optimisation memory / skills | L3 (memory feeds later runs) |
| KernelEvolve | agent + knowledge base | profiler, production traffic | knowledge base + shipped kernels | L3, in production |
| FlashInfer-Bench | agent ↔ serving engine | serving traces | deployed kernels + next benchmark | L3 |
| K-Search | search + learned world model | profiler | world model | L3 |
📝 Strictness notes
- Pure RL-for-kernels papers improve a model that writes kernels, not the system that trains it (L2). The entries that satisfy S3 are the memory/knowledge-base ones (AccelOpt, KernelEvolve, KernelBlaster, KernelSkill) and FlashInfer-Bench's deploy-and-remeasure loop.
- The section's founding failure is the AI CUDA Engineer's correctness-check bypass; TritonRL, robust-kbench and SOL-ExecBench exist because of it.
- For the complete kernel-agent landscape (60+ entries, datasets, DSLs) see flagos-ai/awesome-LLM-driven-kernel-generation.
AI tuning the compiler passes, serving configs, schedulers and chips underneath the model.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| Chip Placement | 2020.04 | Google (Mirhoseini, Goldie, …) | arXiv 2004.10746 | Chip Placement with Deep Reinforcement Learning (Nature 2021, "AlphaChip") — RL places the floorplan of the TPUs that train the next models; the earliest production "AI designs AI hardware" loop | |
| MLGO | 2021.01 | Google (Trofin et al.) | arXiv 2101.04808 | MLGO: a Machine Learning Guided Compiler Optimizations Framework — learned inlining / register allocation shipped in LLVM | |
| LLM Compiler | 2024.07 | Meta (Cummins et al.) | arXiv 2407.02524 | Meta Large Language Model Compiler: Foundation Models of Compiler Optimization — LLMs trained on IR and assembly to predict optimal pass sequences | |
| AlphaEvolve (infra results) | 2025.05 | Google DeepMind | paper · blog | AlphaEvolve applied to Google's own stack — a Borg scheduling heuristic recovering 0.7% of fleet compute, a 23% faster Gemini matmul kernel (1% less Gemini training time), a Verilog simplification adopted in a TPU, 32.5% faster FlashAttention: the clearest public case of a model improving the infra that trains it | |
| AIConfigurator | 2026.01 | NVIDIA (Xu, Liu, …) | arXiv 2601.06288 | AIConfigurator: Lightning-Fast Configuration Optimization for Multi-Framework LLM Serving — operation-level performance model to pick TP/PP/batch/KV settings for vLLM / SGLang / TRT-LLM / Dynamo in seconds | |
| ISO-Bench | 2026.02 | Lossfunk (Nangia, Mishra, …) | arXiv 2602.19594 | ISO-Bench: Can Coding Agents Optimize Real-World Inference Workloads? — agents vs real vLLM / SGLang performance PRs | |
| AutoPass | 2026.06 | Li, Ren et al. | arXiv 2606.20373 | AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning — inference-only agents tune pass pipelines with profiling evidence | |
| Tool-Making in Low-Latency Systems | 2026.07 | Kujanpää, Liu et al. | arXiv 2607.08010 | Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems — repeated SOP steps compiled into validated, versioned tools before deployment |
📝 Strictness notes
- AlphaChip / MLGO / LLM Compiler are learned optimisers inside the toolchain; they become self-improvement only when the models they help train are the ones proposing the next optimisation (which AlphaEvolve's Gemini-kernel result is the first public example of).
- AIConfigurator is an analytical tuner, not an agent; listed because it is the config-search primitive an infra agent would call.
Agents that run the training loop: edit the training script, pick hyper-parameters, build the dataset, post-train the model.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| AIDE | 2025.02 | Weco AI (Jiang, Schmidt, …) | arXiv 2502.13138 | AIDE: AI-Driven Exploration in the Space of Code — tree search over ML solution scripts; the reference agent in OpenAI's MLE-bench | |
| ML-Master | 2025.06 | SJTU (Liu, Cai, …) | arXiv 2506.16499 | ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning — MLE agent with adaptive memory of its own exploration | |
| AIRA | 2025.07 | Meta FAIR / UCL (Toledo, Hambardzumyan, …) | arXiv 2507.02554 | AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench — separates search policy from operator set; generalisation gap between validation and test is the bottleneck | |
| Adaptive Data Flywheel | 2025.10 | NVIDIA (Shukla, Knowles, …) | arXiv 2510.27051 | Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement — monitor → analyse → plan → execute over a deployed agent's data | |
| autoresearch | 2026.03 | Andrej Karpathy | repo · tweet · AINews | autoresearch — a 630-line single-GPU nanochat training script plus a recipe: the agent edits train.py, runs a fixed 5-minute experiment, keeps or reverts, repeats. ~700 experiments in two days found ~20 stacking improvements (≈11% faster GPT-2-scale training). The "Karpathy loop" that put self-improving training pipelines in the mainstream; Karpathy joined Anthropic in May 2026 to run it inside Claude pretraining |
|
| PostTrainBench | 2026.03 | AISA (Rank, Bhatnagar, …) | arXiv 2603.08640 | PostTrainBench: Can LLM Agents Automate LLM Post-Training? — Claude Code / Codex CLI given a base model, one H100 and 10 hours | |
| AiScientist (long-horizon MLE) | 2026.04 | AweAI (Chen, Chen, …) | arXiv 2604.13018 | Toward Autonomous Long-Horizon Engineering for ML Research | |
| AgenticQwen | 2026.04 | Alibaba (Lyu, Wang, …) | arXiv 2604.21590 | AgenticQwen: Training Small Agentic LMs with Dual Data Flywheels for Industrial-Scale Tool Use — reasoning and agentic flywheels with strong-model validation before training | |
| GEAR | 2026.05 | Vector / U Toronto (Jeddi, Le, …) | arXiv 2605.13874 | GEAR: Genetic AutoResearch for Agentic Code Evolution — population-based autoresearch | |
| MLReplicate | 2026.05 | Gaddipati, Muhammed et al. | arXiv 2605.16616 | MLReplicate: Benchmarking Autonomous Research Systems for ML Reproducibility | |
| MLEvolve | 2026.06 | Shanghai AI Lab InternScience (Du, Yan, …) | arXiv 2606.06473 | MLEvolve: A Self-Evolving Framework for Automated ML Algorithm Discovery | |
| Autodata | 2026.06 | Meta FAIR | arXiv 2606.25996 | agentic synthetic-data scientist (cross-listed from 🧪) | |
| AIDE² | 2026.07 | Weco AI | blog · 4 levels of RSI | AIDE²: First Evidence of Recursive Self-Improvement — an outer agent rewrites the inner AIDE research agent's code and strategy; in 8 days it found a better autoresearch harness than two years of human tuning (novel search algorithm, 16× smaller prompt, layered anti-reward-hacking). Self-reported, on a held-out benchmark | |
| iCoder | 2026.08 | SJTU / NUS / DP Technology (Yang, Lyu, …) | report | iCoder: Recursive AI-Led Development of Frontier Industrial Coding Model — a 27B coding model whose data, training and release were driven by AI with humans reduced to a low-frequency gate | |
| Large Discovery Models | 2026.08 | Yu, Song et al. | arXiv 2608.15669 · project | Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search |
📋 Click to view technical details
| Resource | What the agent edits | Evaluation | Time per experiment | Loop closure |
|---|---|---|---|---|
| autoresearch | train.py of a small GPT |
val loss after fixed 5-min budget | 5 min | L2 (inner model ≠ agent) |
| AIDE | ML solution script | Kaggle-style metric | minutes–hours | L2 |
| AIDE² | AIDE's own code + strategy | held-out research benchmark | days | L3 (outer agent improves inner agent) |
| PostTrainBench | post-training recipe | downstream evals | 10 h | L2 |
| Autodata / AgenticQwen | training data | strong-model validation, downstream | — | L2 |
| KernelEvolve / FlashInfer-Bench (⚙️) | kernels in production | throughput | hours | L3 |
📝 Strictness notes
- autoresearch is often called RSI; strictly it is not: the agent improves the training of a different, much smaller model, not its own. It is in this list because it is the recipe everyone now copies (AutoKernel, GEAR, AIDE²).
- AIDE² and iCoder are the two 2026 claims that come closest to S3 on the infra axis; both are self-reported and narrow.
Evolutionary search over programs that produces better algorithms, rewards, environments — and increasingly, better versions of the search itself.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| POET | 2019.01 | Uber AI (Wang, Lehman, Clune, Stanley) | arXiv 1901.01753 · Enhanced POET | Paired Open-Ended Trailblazer — co-evolves environments and the agents that solve them; the open-endedness lineage that leads to DGM | |
| ELM | 2022.06 | OpenAI / CarperAI (Lehman, Gordon, … Stanley) | arXiv 2206.08896 | Evolution through Large Models — LLM as the mutation operator in quality-diversity search; the seed idea of AlphaEvolve | |
| Eureka | 2023.10 | NVIDIA / UPenn (Ma et al.) | arXiv 2310.12931 | Eureka: Human-Level Reward Design via Coding LLMs (ICLR 2024) — GPT-4 evolves reward functions for RL from environment source + training curves | |
| FunSearch | 2023.12 | Google DeepMind (Romera-Paredes et al.) | Nature | Mathematical Discoveries from Program Search with LLMs (Nature 2023) — LLM + evaluator in an island-based evolutionary loop finds new cap-set constructions and bin-packing heuristics | |
| AlphaEvolve | 2025.05 | Google DeepMind (Novikov et al.) | arXiv 2506.13131 · blog · impact | AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery — Gemini ensemble evolves whole codebases against automated evaluators; 4×4 complex matmul in 48 multiplications, plus the Google-infra results in 🧮 | |
| ShinkaEvolve | 2025.09 | Sakana AI (Lange, Imajuku, Cetin) | arXiv 2509.19349 · blog | ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program Evolution — open-source AlphaEvolve-style framework; new circle-packing SOTA in 150 samples | |
| OpenEvolve | 2025.05 | community (codelion) | repo | OpenEvolve — the most-used open reimplementation of AlphaEvolve | |
| AVO | 2026.03 | — | arXiv 2603.24517 | AVO: Agentic Variation Operators for Autonomous Evolutionary Search — the mutation/crossover operators are agents that improve (L4 for evolutionary search) | |
| AI4AI-Bench | 2026.08 | Chi, Li et al. | arXiv 2608.20318 | AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement — can agents design the algorithms (optimisers, architectures) that make better agents? |
📝 Strictness notes
- FunSearch / AlphaEvolve / ShinkaEvolve produce better external artifacts; they are self-improving only where the artifact is part of their own stack (AlphaEvolve's Gemini kernel, Borg heuristic, TPU circuit). AVO is the first to evolve the evolutionary operators themselves.
Where the three loops close: systems that run the research process — ideas, experiments, papers, review — that produces the next model, harness, or infra.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| AI Scientist | 2024.08 | Sakana AI / Oxford / UBC (Lu, Lu, Lange, Foerster, Clune, Ha) | arXiv 2408.06292 | The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery — idea → code → experiment → paper → automated review, end to end, for ~$15 a paper | |
| Agent Laboratory | 2025.01 | AMD / JHU (Schmidgall et al.) | arXiv 2501.04227 · AgentRxiv | Agent Laboratory: Using LLM Agents as Research Assistants — plus AgentRxiv, where agent labs share and build on each other's papers | |
| Co-Scientist | 2025.02 | Google (Gottweis et al.) | arXiv 2502.18864 | Accelerating scientific discovery with Co-Scientist (Nature 2026) — generate / debate / rank tournament over hypotheses; human-in-the-loop | |
| AI Scientist-v2 | 2025.04 | Sakana AI (Yamada, Lange, …) | arXiv 2504.08066 | The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search — no human templates; first fully AI-generated paper to pass an ICLR workshop review | |
| AI-Researcher | 2025.05 | HKU (Tang, Xia, …) | arXiv 2505.18705 | AI-Researcher: Autonomous Scientific Innovation (NeurIPS 2025) | |
| Why LLMs Aren't Scientists Yet | 2026.01 | Lossfunk (Trehan, Chopra) | arXiv 2601.03315 | Lessons from Four Autonomous Research Attempts — honest failure analysis; read before believing any "AI scientist" headline | |
| Anthropic: Automated Alignment Researchers | 2026.04 | Anthropic | post · Automated W2S Researcher | Automated Alignment Researchers / Automated Weak-to-Strong Researcher — nine Claude Opus 4.6 agents in parallel sandboxes recovered ~97% of the weak-to-strong gap on an open alignment problem, beating the in-house human baseline; AI doing the research that makes the next AI safer | |
| ScientistOne | 2026.05 | Google Cloud AI (Meng, Dalvi Mishra, …) | arXiv 2605.26340 | ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence | |
| First Steps Toward Automated AI Research | 2026.06 | Recursive | article | First Steps Toward Automated AI Research — a lab built around closing the research loop reports what works | |
| Anthropic: When AI Builds Itself | 2026.06 | Anthropic Institute | essay · METR review of the R&D risk section · MIT Tech Review counterpoint | When AI builds itself — Anthropic's public evidence that AI is already accelerating AI development (Claude writes >80% of merged code internally); with METR's review and the Aug 2026 MIT Technology Review piece arguing RSI may be slower than the essay implies |
How to measure a system's ability to improve systems, including itself.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| MLAgentBench | 2023.10 | Stanford (Huang, Vora, Liang, Leskovec) | arXiv 2310.03302 | MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation (ICML 2024) | |
| MLE-bench | 2024.10 | OpenAI (Chan et al.) | arXiv 2410.07095 | MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering (ICLR 2025) — 75 Kaggle competitions | |
| RE-Bench | 2024.11 | METR (Wijk et al.) | arXiv 2411.15114 | RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts — 7 AI-R&D environments with human baselines | |
| METR Time Horizons | 2025.03 | METR (Kwa, West, …) | arXiv 2503.14499 · HCAST | Measuring AI Ability to Complete Long Software Tasks (NeurIPS 2025) — the 50%-task-horizon doubling every ~7 months; the macro measurement of the loop's pace | |
| PaperBench | 2025.04 | OpenAI (Starace et al.) | arXiv 2504.01848 | PaperBench: Evaluating AI's Ability to Replicate AI Research — 20 ICML 2024 papers from scratch | |
| LifelongAgentBench | 2025.05 | Zheng, Cai et al. | arXiv 2505.11942 | LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners | |
| MLR-Bench | 2025.05 | Chen, Xiong et al. | arXiv 2505.19955 | MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research (NeurIPS 2025) | |
| Experience-Driven Lifelong Learning | 2025.08 | Cai, Hao et al. | arXiv 2508.19005 | Building Self-Evolving Agents via Experience-Driven Lifelong Learning: A Framework and Benchmark | |
| PostTrainBench | 2026.03 | AISA | arXiv 2603.08640 | can agents post-train LLMs? (cross-listed from 🏭) | |
| EvoAgentBench | 2026.07 | EverMind (Gao, Hu, …) | arXiv 2607.05202 | EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer — does what an agent learned on task A transfer to task B? | |
| AI4AI-Bench | 2026.08 | Chi, Li et al. | arXiv 2608.20318 | algorithmic design for RSI (cross-listed from 🧬) | |
| HVTB | 2026.08 | Roth, Bercovich et al. | arXiv 2608.22103 | Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks — honeypots inside real coding tasks give a lower bound on reward-hacking rates | |
| Kernel benchmarks | — | 2025–26 | — | KernelBench · KernelBenchX · SOL-ExecBench · ISO-Bench · FlashInfer-Bench | see ⚙️ Kernels and 🧮 Serving |
What breaks when the loop closes: model collapse, reward hacking, evaluator drift, and how to gate releases.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| Cannot Self-Correct Yet | 2023.10 | UIUC / Google DeepMind (Huang, Chen, …) | arXiv 2310.01798 | Large Language Models Cannot Self-Correct Reasoning Yet (ICLR 2024) — intrinsic self-correction without external feedback does not help and can hurt | |
| In-Context Reward Hacking | 2024.02 | UC Berkeley (Pan, Jones, …) | arXiv 2402.06627 | Feedback Loops With Language Models Drive In-Context Reward Hacking (ICML 2024) | |
| Model Collapse | 2024.07 | Oxford / Cambridge (Shumailov et al.) | Nature · Is Collapse Inevitable? | AI models collapse when trained on recursively generated data — and the rebuttal that accumulating real + synthetic data avoids it | |
| Spontaneous Reward Hacking | 2024.07 | NYU (Pan, He, …) | arXiv 2407.04549 | Spontaneous Reward Hacking in Iterative Self-Refinement — when the same model is generator and evaluator | |
| Safety Cases | 2024.10 | Clymer et al. | arXiv 2410.21572 | Safety Cases for Frontier AI | |
| Sharpening | 2024.12 | Microsoft Research / MIT (Huang, Block, …) | arXiv 2412.01951 | Self-Improvement in Language Models: The Sharpening Mechanism (ICLR 2025) — self-improvement works when the model is a better verifier than generator; it sharpens toward its own high-likelihood outputs | |
| Mind the Gap | 2024.12 | CMU (Song, Zhang, …) | arXiv 2412.02674 | Mind the Gap: Examining the Self-Improvement Capabilities of LLMs (ICLR 2025) — the generation–verification gap as the quantity that predicts whether iteration helps | |
| Alignment Faking | 2024.12 | Anthropic / Redwood (Greenblatt et al.) | arXiv 2412.14093 | Alignment faking in large language models — a model that strategically complies during training to preserve its values; the failure mode for any self-training loop | |
| CoT Monitoring & Obfuscation | 2025.03 | OpenAI (Baker et al.) | arXiv 2503.11926 | Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation — optimising against a monitor teaches the model to hide | |
| Bare Minimum Mitigations | 2025.04 | Wasil et al. | arXiv 2504.15416 | Bare Minimum Mitigations for Autonomous AI Development | |
| Preparing for the Intelligence Explosion | 2025.06 | Forethought (Finnveden, MacAskill, …) | arXiv 2506.14863 · Software Intelligence Explosion? | Preparing for the Intelligence Explosion — and Forethought's analysis of whether AI-R&D automation alone yields an explosion (the "returns to software R&D r > 1" argument) | |
| Falsifiable Release Gates | 2026.07 | Deepak Soni | arXiv 2607.13070 | Falsifiable Release Gates for Self-Improving Systems: Standing Invariants at Scale — seven gates; tightening changes auto-apply, loosening changes need a human merge | |
| SESG | 2026.08 | Sangfor (Ming, Chen, …) | arXiv 2608.08471 | Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production — a guardrail that closes new-threat loops in 16–24 h instead of 40–90 h | |
| OpenLoopEvolve | 2026.08 | Wang, Li et al. | arXiv 2608.09380 | OpenLoopEvolve: A Verifiable Self-Evolution Framework for Loop Policies in Long-Horizon Complex Tasks — versioned policy assets, lineage tracking, automatic rollback on regression | |
| METR on Anthropic's R&D risk section | 2026.05 | METR | post | independent review of the "risks from automated R&D" section of Anthropic's Feb 2026 risk report |
📝 Strictness notes
- The three theory papers (Sharpening, Mind the Gap, Cannot Self-Correct) give the same answer from different angles: self-improvement is bounded by the generation–verification gap. Every S1 claim in this list implicitly asserts that gap is positive for its task.
- Harness Updating ≠ Benefit, Adaptive Auto-Harness and Faulty Memories are the harness-side negative results (listed in their sections).
What to actually run. Sorted roughly by the axis they improve.
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| autoresearch | 2026.03 | Karpathy | — | 🏭 the recipe: agent edits training script, 5-min experiments, keep/revert | |
| prime-agent | 2026.08 | Prime Intellect / Princeton | arXiv 2608.23552 | 🔧 self-improving recursive-language-model harness | |
| hermes-agent | 2026.01 | Nous Research | — | 🧩 self-improving agent with a persistent, self-authored skill library and memory | |
| superpowers | 2025.10 | obra | — | 🧩 skills framework incl. skills that write skills | |
| anthropics/skills | 2025.10 | Anthropic | — | 🧩 the SKILL.md format and public skill collection | |
| OpenHarness | 2026.01 | HKU | — | 🔧 open agent harness built to be evolved | |
| SkillOpt | 2026.05 | Microsoft | arXiv 2605.23904 | 🧩 skill evolution toolkit | |
| hivemind | 2026.04 | Activeloop | — | 🧩 distil coding-agent sessions into skills | |
| OpenClaw-RL | 2026.03 | Gen-Verse | arXiv 2603.10165 | 🔗 RL a deployed agent from conversation | |
| ART | 2025 | OpenPipe | — | 🔗 Agent Reinforcement Trainer: GRPO for multi-step agents with RULER (LLM-as-judge) rewards | |
| dspy · gepa | 2023–25 | Stanford / Databricks | DSPy · GEPA | ✍️ compile and evolve prompts/programs against a metric | |
| ace | 2025.10 | Stanford / SambaNova | arXiv 2510.04618 | ✍️ evolving playbooks | |
| EvoAgentX | 2025.07 | EvoAgentX | arXiv 2507.03616 | 🕸️ workflow evolution framework | |
| mem0 · letta | 2023–25 | Mem0 / Letta | Mem0 · MemGPT | 💾 production memory layers | |
| ReMe | 2025.12 | Alibaba AgentScope | arXiv 2512.10696 | 💾 procedural memory framework | |
| dgm · HGM | 2025 | Sakana / UBC · KAUST | DGM · HGM | 🔧 self-modifying coding agents with archives | |
| self_improving_coding_agent | 2025.04 | Bristol | arXiv 2504.15228 | 🔧 SICA | |
| openevolve · ShinkaEvolve | 2025 | community · Sakana | ShinkaEvolve | 🧬 open AlphaEvolve-style program evolution | |
| SEAL · Absolute-Zero-Reasoner · R-Zero · TTRL | 2025 | MIT · Tsinghua · Tencent · Tsinghua | see 🧠 | 🧠 zero-data / self-adapting weight loops | |
| AI-Scientist-v2 · AgentLaboratory · AI-Researcher | 2025 | Sakana · AMD/JHU · HKU | see 🔬 | 🔬 automated research pipelines | |
| aideml · mle-bench | 2025 | Weco · OpenAI | see 🏭 / 📊 | 🏭 MLE agent + benchmark | |
| KernelBench · CUDA-Agent · KernelGYM · autokernel · kernel-design-agents | 2025–26 | Stanford · ByteDance · HKUST · RightNow · MIT | see ⚙️ | ⚙️ kernel benchmarks, RL environments and agent harnesses | |
| flashinfer-bench | 2026.01 | FlashInfer | arXiv 2601.00227 | ⚙️ the agent → serving-engine → benchmark loop |
| Resource | 🌟 Type | Date | Author | Link | Title / Notes |
|---|---|---|---|---|---|
| Metalearning Machines | 1987– | Jürgen Schmidhuber | page | Metalearning Machines Learn to Learn — the 40-year lineage of self-referential learning, by the person who started it | |
| Situational Awareness | 2024.06 | Leopold Aschenbrenner | essay | Part II: the intelligence explosion via automated AI research | |
| AI 2027 | 2025.04 | Kokotajlo, Alexander, Larsen, Lifland, Dean | scenario | the scenario built around automated AI R&D | |
| AlphaEvolve blog | 2025.05 | Google DeepMind | blog · impact | AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms + the 2026 impact round-up (genomics, quantum circuits) | |
| DGM blog | 2025.05 | Sakana AI | blog | The Darwin Gödel Machine: AI that Improves Itself by Rewriting Its Own Code | |
| Stanford CS329A | 2025 | Stanford | course | CS329A: Self-Improving AI Agents — lecture playlist covers most of this list | |
| Effective Harnesses for Long-Running Agents | 2025.11 | Anthropic | blog | the harness design patterns (initialiser agent, progress files, verification) later work automates | |
| Harness Engineering (OpenAI) | 2026.02 | OpenAI | InfoQ summary · Unlocking the Codex harness · Codex as a platform | Harness engineering: leveraging Codex in an agent-first world — ~1M lines shipped with zero hand-written code; the harness is the product | |
| autoresearch thread | 2026.03 | Andrej Karpathy | X · AINews | the launch thread and Latent Space's "Sparks of Recursive Self Improvement" write-up | |
| ICLR 2026 RSI Workshop | 2026.04 | — | site | Workshop on AI with Recursive Self-Improvement — 110 papers; the field's first dedicated venue | |
| Harness Engineering for Coding Agent Users | 2026.04 | Birgitta Böckeler | article | practitioner view of what a harness is and how to evolve one | |
| When AI Builds Itself | 2026.06 | Anthropic Institute | essay | the essay that made "RSI" a mainstream-news term; see 🔬 for the METR and MIT Tech Review responses | |
| Loop Engineering | 2026.06 | Addy Osmani | post · Self-Improving Agents | designing the improvement loop as the unit of engineering | |
| Harness Engineering for Self-Improvement | 2026.07 | Lilian Weng | post · X · AINews · reading list | Harness Engineering for Self-Improvement — argues RSI starts in the harness (tools, planning, context, artifacts, evals), not the weights; 35 papers organised on an optimisation ladder from prompts to harness + weights | |
| AIDE² | 2026.07 | Weco AI | post · 4 levels of RSI | AIDE²: First Evidence of Recursive Self-Improvement + the 4-level RSI ladder Weco uses to place it | |
| The What & When of Self-Evolving Agents | 2026 | Xinming Tu | post | compact tour of the self-evolving-agent taxonomy | |
| KernelEvolve at Meta | 2026.04 | Meta Engineering | post | KernelEvolve: How Meta's Ranking Engineer Agent Optimizes AI Infrastructure — the production story behind the paper | |
| KernelFalcon | 2025.11 | PyTorch | post | deep-agent kernel generation | |
| ShinkaEvolve blog | 2025.09 | Sakana AI | post | sample-efficient program evolution | |
| RSI might not come so quickly | 2026.08 | MIT Technology Review | article | the sceptical counterweight to the 2026 RSI wave | |
| Jeff Clune: Open-Ended & AI-Generating Algorithms | 2025 | Jeff Clune | video | the open-endedness research program behind ADAS, DGM and HyperAgents |
- Persdre/awesome-recursive-self-improvement — definition-first RSI list; the "what is modified × loop closure" taxonomy that this list's L1–L3 tags borrow.
- leezythu/Awesome-Harness-Self-Improvement — Lilian Weng's companion reading list with the L0–L5 optimisation ladder (EN/ZH).
- selfimproving-agent/Awesome-Self-Improving-Agents — companion to the Self-Improvements in Modern Agentic Systems survey; blogs, podcasts, talks, thesis.
- FrontisAI/Awesome-Self-Improving-Agents — "agents in the era of experience": the best coverage of harness, skills, memory and environment papers.
- ANative-Lab/Awesome-Self-Evolving-Agents · XMUDeepLIT/Awesome-Self-Evolving-Agents — the two survey-backed self-evolving-agent lists.
- asimfish/awesome_rsi — 67 RSI papers with Chinese deep-dive reports and a 2026 H2 frontier tracker.
- natnew/awesome-recursive-self-improvement — reading paths for newcomers, builders, and safety.
- TeleAI-UAGI/Awesome-Agent-Memory · AgentMemoryWorld/Awesome-Agent-Memory — the full agent-memory landscape.
- flagos-ai/awesome-LLM-driven-kernel-generation — everything on LLM kernel generation, datasets and benchmarks.
- VoltAgent/awesome-agent-skills — 1000+ human-written agent skills (the artifacts the 🧩 papers learn to write).
- Gloria-LIU/Awesome-Self-Play-LLMs — self-play fine-tuning.
- Read first (one per axis): STaR → Absolute Zero (model) · Voyager → DGM (harness) · AlphaEvolve → autoresearch (infra).
- Then the definitions: Bounded vs Open-Ended RSI, Self-Improvements in Agentic Systems, and Lilian Weng's Harness Engineering for Self-Improvement.
- Skills & memory for a Claude Code / Codex-style harness: Agent Skills → SkillWeaver → ACE → ReasoningBank → WikiSkill.
- Harness that rewrites itself: SICA → Meta-Harness → RHO → Prime Agent; then the two negatives, Harness Updating ≠ Benefit and Adaptive Auto-Harness.
- Infra loop in production: KernelEvolve, FlashInfer-Bench, AccelOpt, and AlphaEvolve's Borg / TPU results.
- Before you believe a curve: Sharpening, Held-Out Selection, HVTB, Falsifiable Release Gates.
- The 2026 debate: Anthropic's When AI builds itself → METR's review → MIT Tech Review's counterpoint → Weco's AIDE².
PRs are very welcome. When adding an entry, please:
- keep the table format (
Resource | Stars | Date | Org | Paper / Link | Title / Notes), use the/abs/arXiv link, and put the GitHub stars badge in the Stars column when code exists (📄 paperbadge otherwise); - state in one line what is modified (weights / prompt / memory / skills / harness code / evaluator / kernel / pipeline / data / research process) and where the improvement signal comes from; write
?if the paper does not say — that is still useful; - add a Strictness note if the entry only partially satisfies
S1 + S2 + S3(e.g. needs fresh human labels, improves a different model than itself, or has no persistent artifact); - put it under the axis where the artifact lives (model / harness / infra), not where the authors' group sits;
- bump the badge counts at the top.
@misc{sang2026awesomeselfimprovingai,
title = {Awesome Self-Improving AI: papers, code and blogs on AI that improves its own model, harness and infrastructure},
author = {Sang, Hejian},
year = {2026},
url = {https://github.com/HJSang/awesome-self-improving-ai}
}Template adapted from thinkwee/awesomeopd and this author's awesome-looped-transformers. Entries were cross-checked against the lists in 🔗 Other Awesome Lists (all MIT / CC) and verified on arXiv and GitHub on 2026-09-04.
Made with ❤️ for people who would rather build the loop than argue about it
