A self-improvement research testbed for LLMs, built around Dream-RSI — recursive self-improvement through evolving worlds — together with a classic weight-upgrade track, both running on one shared substrate of tasks, verifiers, and sandbox.
The mechanism at the heart of this project belongs to the Dream-RSI paper (Google; Tong Zheng et al., arXiv:2609.14858): treat accumulated discovery history as a replay simulator over the realized search space, "dream" in it to refine the exploration policy at zero model cost, then redeploy the policy so it expands the simulator in turn. All credit for that idea belongs to the paper's authors; this repository is an unofficial, independent engineering implementation of it, documented honestly where our protocol choices diverge.
dream-loop runs two self-improvement loops that share the same tasks, the same
verifiers, and the same sandbox:
- Line A — weight upgrades. Classic STaR-style rejection-sampling self-training: solve tasks by day and bank every attempt, then build an SFT corpus from successes and DPO pairs from same-task success/fail contrasts, train a LoRA, and promote it only after it passes a frozen probe-set gate. The improved artifact is the model's weights.
- Line B — Dream-RSI (implements the loop of arXiv 2609.14858 Dream-RSI: Recursive Self-Improvement through Evolving Worlds). The model is frozen for the entire loop; the only artifact that changes is a small file of Python exploration-policy code, rewritten every cycle by a fixed developer agent and selected by replay scoring.
LINE A — weights change ----------------------------------------------------
wake ──────────► hippocampus ──────► night-build ─────► night-train ────► gate-eval
solve tasks, append-only SFT corpus from LoRA (SFT+DPO, frozen probe set:
k samples each, episode store successes, DPO GPU) promote / reject
verify, log from contrasts
LINE B — policy code changes, model frozen ---------------------------------
explore ───────► trees.jsonl ─────► dream: replay ───► developer agent ─► argmax redeploy
Stage 1: the one tree = one Stage 2: reveal Stage 3: a fixed best revision
policy expands replay world; stored results LLM rewrites becomes the next
a discovery tree equation 1 deterministically, the policy file incumbent policy
node by node scores each zero model calls
revision
pip install -r requirements.txt # numpy/sklearn cover the paper domains
# Self-check the whole pipeline without spending a single API call:
python3 tests/test_replay.py # replay simulator + equation 1 semantics
python3 tests/test_graded.py # continuous (per-assert) scorer
python3 tests/test_paper_tasks.py # all 5 paper domains, incl. GPU when present
DREAMLOOP_MOCK=1 python3 -m dreamloop explore --source packing --rounds 3 --workers 2
DREAMLOOP_MOCK=1 python3 -m dreamloop dream --revisions 2
# Real run — same loop, mock flag removed (set DREAMLOOP_API_KEY first):
python -m dreamloop init-probes # freeze the gate's probe set (one-time)
python -m dreamloop explore --source lasso --rounds 8 --workers 4
python -m dreamloop dream --revisions 3 # the next explore automatically uses the winning policyDREAMLOOP_MOCK=1 replaces every LLM call with deterministic stubs (weak but
feasible artifacts), so the full loop runs end-to-end offline.
Dream-RSI's contract, as implemented here:
| Paper mechanism | Where |
|---|---|
| Discovery tree (parent chains / workspace snapshots / artifacts / diagnostics / continuous score s_v) | dreamloop/discovery.py, persisted to data/dreams/trees.jsonl |
| Stage 1 online exploration: the policy picks ≤W nodes per round; each expansion is one real discovery-agent generate+evaluate | dreamloop/explore.py |
Policy observation API (prefix visibility; legal_actions = root ∪ revealed leaves) |
dreamloop/question.py |
| Stage 2 replay simulator: reveals stored results deterministically, never executes any model | replay mode in question.py |
| Equation 1: max s − β1·N + β2·N/max(1,k), averaged over all worlds | Question.equation1 |
| Stage 3 dreaming: a fixed LLM developer agent rewrites the policy code, M revisions | dreamloop/dream.py + prompts.DREAM_DEVELOPER |
| argmax selection (incumbent included as m=0; replay score never regresses) + redeployment | dream.py, chosen/redeploy logic |
| Only the policy changes | data/policies/policy.py is the single rewritten file |
Replay reveal rule (the paper's own semantics): picking the root reveals the root's earliest-created unrevealed child (opens a new branch); picking a non-root node reveals its only recorded child, if any; otherwise that slot of the round stays empty.
Prefix visibility is enforced by construction: the policy receives a facade
object (Question.view) that exposes only the observation/decision surface —
the tree, the revealed set, and the trail are unreachable. Bypassing via
reflection is treated as cheating: a revision that reveals nodes in zero rounds
is disqualified entirely (V=None, excluded from argmax).
Two original domains plus five domains from the paper's experiment section,
each a task package in dreamloop/tasks/ verified by dreamloop/verifiers/.
All candidate code runs in a restricted subprocess (python3 -I + rlimits +
private process group, under the dlbox sandbox when available; BLAS threads
pinned to 1 for fair timing).
--source |
Task | Score s_v (higher is better) | Correctness gate |
|---|---|---|---|
humaneval |
HumanEval code synthesis | per-assert pass ratio | all asserts must pass |
lean |
Lean 4 theorem proving (8 core-library tasks built in) | binary | lean kernel accepts, no sorry |
lasso |
Lasso regularization-path solver (17 synthetic timing + 17 fresh gate instances) | geometric-mean speedup R = (∏ t_sklearn_i/t_cand_i)^(1/17); R ≥ 1 beats sklearn | per-λ gate F_k(w̃_k) ≤ F_k(w_k^sklearn) + 1e-6 on fresh instances; any failure → 0 |
packing |
n-circle packing in the unit square (n ∈ {26, 32}) | Σ r_i when feasible; paper optimum 2.635983 (n=26) | containment + non-overlap at tolerance 1e-7; violation → 0 |
sumdiff |
Sum–Difference problem: choose A ⊂ ℤ maximizing Γ(A) = log(|A+A|/|A|)/log(|A−A|/|A|) | Γ(A); paper optimum 1.145427 | set validity (distinct ints, ≥2, size/span guardrails); Γ recomputed by the verifier |
autocorr |
Autocorrelation inequality (first variant): f ≥ 0, support [−1/4,1/4], ∫f = 1, minimize Φ₁ = max(f∗f) | −Φ₁; paper 1.456375, SOTA 1.453675 | non-negativity and normalization at tolerance 1e-6; violation → 0 |
kernels |
KernelBench: VGG16, LayerNorm, ConvDiv, ConvMax | 1/t_cand_ms (inverse runtime); correctness vs reference via allclose(atol=rtol=1e-2) | wrong output → 0; requires torch + CUDA (missing → readable error, score 0; the tree still grows) |
Dependencies: the lasso verifier needs scikit-learn (it is both the reference
implementation and the timing baseline; reference solutions and reference
runtimes are cached per instance to data/cache/paper/lasso/*.npz on first use,
so later verifications carry no sklearn overhead). All domains need numpy;
kernels additionally needs CUDA torch (see requirements.txt). Lasso
instance recipes and timeouts live in the paper: section of config.yaml.
Every piece of model-generated code is untrusted and is treated as such:
-
dlbox(sandbox/dlbox/, Rust) is a minimal rootless sandbox built for this project: user + mount + network + PID/IPC/UTS namespaces, whole root remounted read-only (CUDA/numpy stay usable, GPU included), writable surface narrowed to the workspace and a fresh/dev/shm, rlimits on CPU / memory / process count,PR_SET_NO_NEW_PRIVS, and an explicit environment whitelist. Build it with:cd sandbox/dlbox && cargo build --release # then either: cp target/release/dlbox ../../sandbox/dlbox-bin # or: put it on PATH, or export DREAMLOOP_DLBOX=<path>
Discovery order at runtime:
$DREAMLOOP_DLBOX→sandbox/dlbox-bin→dlboxonPATH.DREAMLOOP_SANDBOX=processforce-disables it (bare subprocess + rlimits fallback, which is strictly weaker). -
Secrets never reach candidate code. API keys are read from the environment only, and candidate subprocesses get a minimal allowlisted environment.
-
Lasso cache hardening: reference
.npzfiles enter the sandbox as copies (cache originals are invisible to candidates, closing the rewrite-t_ref-and-poison-scoring channel), and each file carries an instance signature — any recipe change or corruption triggers a rebuild. -
Policy code is executed in-process during
explore(SIGALRM soft timeout only) and in a subprocess duringdreamreplay. This is the one intentional weak spot: there is no strong isolation anywhere in this repo. For serious runs, put the whole thing in a container.
Definitions were implemented from primary sources and are labeled accordingly in each module's docstring:
- Lasso objective/gate/scoring, packing constraints, the sumdiff Γ, and the autocorr first-variant Φ₁ + constraints follow the paper text (arXiv 2609.14858, checked against the HTML full text).
- KernelBench reference implementations are embedded verbatim and verified against the upstream main branch (VGG16 = level3/11, LayerNorm = level1/40, ConvDiv = level2/71, ConvMax = level2/82).
- Unverified (self-devised protocol, the paper gives no values): the exact
lasso instance recipes (following SimpleTES), the autocorr grid N=1024, and
the sumdiff size/span guardrails.
tests/test_autocorr_sensitivity.pydemonstrates the autocorr ranking is stable across N ∈ {512, 1024, 2048}. - Third-party code bundled in this repository:
- KernelBench reference implementations (
dreamloop/tasks/paper/kernels.py) — MIT License, © Scaling Intelligence Lab, Stanford University (upstream). - HumanEval is downloaded at runtime from openai/human-eval (MIT), not bundled.
- KernelBench reference implementations (
- The paper's six downstream held-out datasets (Gisette, RCV1, DNA, Leukemia, Colon, Duke Breast) require external download and are not included.
The paper fixes only the symbols K1, K2, M, β1, β2. Everything else is this
repo's choice, made explicit: the developer agent reuses the config.api
endpoint; baseline_score is the task's best historical score; code tasks use
per-assert continuous scoring to feed equation 1 (lean falls back to binary);
ties keep the incumbent; revisions that crash, time out, reveal nothing, or
show tampering are disqualified wholesale (V=None).
All domain adaptation, not mechanism substitution:
- The original domains are HumanEval + Lean; the paper's Lasso / math
optimization / KernelBench domains are implemented in
tasks/paper/+verifiers/paper/. The paper's "workspace snapshot" is realized as a text context (inherited history + task statement). - A "discovery" manifests as iterative improvement of failed attempts (parent chains), rather than the paper's refinement of engineering artifacts.
All knobs live in config.yaml: W (workers) / K1 (rounds) / K2
(replay_rounds) / M (revisions) / β1 / β2 / temperatures / soft timeouts /
policy_config (passed to OptimalPolicy.config). Relative paths are anchored
to the config file's directory, so the CLI behaves identically from any cwd.
Environment variables:
| Variable | Meaning |
|---|---|
DREAMLOOP_API_KEY |
API key for the OpenAI-compatible endpoint (required outside mock mode) |
DREAMLOOP_MOCK |
1 = deterministic stubs for every LLM call (offline self-checks) |
DREAMLOOP_DLBOX |
explicit path to the dlbox binary |
DREAMLOOP_SANDBOX |
process = disable dlbox, use the bare-subprocess fallback |
dream-loop/
├── dreamloop/ # the package
│ ├── cli.py # entry point: python -m dreamloop <command>
│ ├── agent.py, llm.py # OpenAI-compatible clients (line A / line B)
│ ├── wake.py # daytime loop: solve → verify → hippocampus
│ ├── hippocampus.py # append-only experience store (flock-guarded)
│ ├── night/ # SFT/DPO corpus building, LoRA training, probe gate
│ ├── question.py # the policy-facing world: one facade, two modes
│ ├── discovery.py # discovery trees — the replay "worlds"
│ ├── explore.py # Stage 1: online exploration
│ ├── dream.py # Stage 2+3: replay scoring, rewrite, argmax
│ ├── policy/ # built-in hand-written exploration policy
│ ├── tasks/ # humaneval, lean, and the five paper domains
│ ├── verifiers/ # code (binary + graded), lean, paper-domain verifiers
│ └── prompts.py # every LLM prompt in one place
├── sandbox/dlbox/ # Rust rootless sandbox (source of truth)
├── tests/ # dependency-light self-tests, no pytest required
├── probes/ # frozen probe set for the gate (regenerate with
│ # `python -m dreamloop init-probes`; not shipped)
├── config.yaml # every knob
└── data/ # runtime artifacts (gitignored, created on first run)
python3 tests/test_replay.py # replay semantics, equation 1, determinism
python3 tests/test_graded.py # scorer: partial credit, crashes, quotas
python3 tests/test_paper_tasks.py # all domains; uses the GPU when present
python3 tests/test_autocorr_sensitivity.pyAll self-tests are offline and free; test_paper_tasks exercises real GPU and
sklearn paths when those are available and readable-error paths when not.
- No strong isolation anywhere (see Sandbox and security). Containerize before running untrusted workloads seriously.
- The lasso
t_refdenominator is frozen once at cache-build time; under changing machine load, scores near R = 1.0 will wobble — an inherent trade-off of the reference cache. - Line A's training assumes a single GPU box with
torch transformers peft trl datasetsinstalled; the rest of the pipeline runs CPU-only.
If you use this implementation, please cite the paper it implements:
@misc{dreamrsi2026,
title = {Dream-RSI: Recursive Self-Improvement through Evolving Worlds},
year = {2026},
howpublished = {\url{https://arxiv.org/abs/2609.14858}},
note = {arXiv:2609.14858}
}This repository is released under the Apache License 2.0.
Third-party code bundled here keeps its own license: the KernelBench reference
implementations in dreamloop/tasks/paper/kernels.py are MIT (© Scaling
Intelligence Lab, Stanford University,
upstream).
This project would not exist without Dream-RSI: Recursive Self-Improvement through Evolving Worlds (Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo — arXiv:2609.14858). We are grateful to the authors for publishing the mechanism with enough precision to build on. Thanks also to the OpenAI HumanEval team and to Scaling Intelligence Lab's KernelBench for the benchmarks this testbed stands on.