Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

dream-loop

A self-improvement research testbed for LLMs, built around Dream-RSI — recursive self-improvement through evolving worlds — together with a classic weight-upgrade track, both running on one shared substrate of tasks, verifiers, and sandbox.

The mechanism at the heart of this project belongs to the Dream-RSI paper (Google; Tong Zheng et al., arXiv:2609.14858): treat accumulated discovery history as a replay simulator over the realized search space, "dream" in it to refine the exploration policy at zero model cost, then redeploy the policy so it expands the simulator in turn. All credit for that idea belongs to the paper's authors; this repository is an unofficial, independent engineering implementation of it, documented honestly where our protocol choices diverge.

arXiv Python Platform GPU License

dream-loop runs two self-improvement loops that share the same tasks, the same verifiers, and the same sandbox:

  • Line A — weight upgrades. Classic STaR-style rejection-sampling self-training: solve tasks by day and bank every attempt, then build an SFT corpus from successes and DPO pairs from same-task success/fail contrasts, train a LoRA, and promote it only after it passes a frozen probe-set gate. The improved artifact is the model's weights.
  • Line B — Dream-RSI (implements the loop of arXiv 2609.14858 Dream-RSI: Recursive Self-Improvement through Evolving Worlds). The model is frozen for the entire loop; the only artifact that changes is a small file of Python exploration-policy code, rewritten every cycle by a fixed developer agent and selected by replay scoring.
LINE A — weights change ----------------------------------------------------

  wake ──────────► hippocampus ──────► night-build ─────► night-train ────► gate-eval
  solve tasks,     append-only         SFT corpus from    LoRA (SFT+DPO,    frozen probe set:
  k samples each,  episode store       successes, DPO     GPU)              promote / reject
  verify, log                          from contrasts

LINE B — policy code changes, model frozen ---------------------------------

  explore ───────► trees.jsonl ─────► dream: replay ───► developer agent ─► argmax redeploy
  Stage 1: the     one tree = one    Stage 2: reveal    Stage 3: a fixed   best revision
  policy expands   replay world;     stored results     LLM rewrites       becomes the next
  a discovery tree equation 1        deterministically, the policy file    incumbent policy
  node by node     scores each       zero model calls
                   revision

Quick start

pip install -r requirements.txt          # numpy/sklearn cover the paper domains

# Self-check the whole pipeline without spending a single API call:
python3 tests/test_replay.py             # replay simulator + equation 1 semantics
python3 tests/test_graded.py             # continuous (per-assert) scorer
python3 tests/test_paper_tasks.py        # all 5 paper domains, incl. GPU when present
DREAMLOOP_MOCK=1 python3 -m dreamloop explore --source packing --rounds 3 --workers 2
DREAMLOOP_MOCK=1 python3 -m dreamloop dream --revisions 2

# Real run — same loop, mock flag removed (set DREAMLOOP_API_KEY first):
python -m dreamloop init-probes            # freeze the gate's probe set (one-time)
python -m dreamloop explore --source lasso --rounds 8 --workers 4
python -m dreamloop dream --revisions 3  # the next explore automatically uses the winning policy

DREAMLOOP_MOCK=1 replaces every LLM call with deterministic stubs (weak but feasible artifacts), so the full loop runs end-to-end offline.

How Line B works

Dream-RSI's contract, as implemented here:

Paper mechanism Where
Discovery tree (parent chains / workspace snapshots / artifacts / diagnostics / continuous score s_v) dreamloop/discovery.py, persisted to data/dreams/trees.jsonl
Stage 1 online exploration: the policy picks ≤W nodes per round; each expansion is one real discovery-agent generate+evaluate dreamloop/explore.py
Policy observation API (prefix visibility; legal_actions = root ∪ revealed leaves) dreamloop/question.py
Stage 2 replay simulator: reveals stored results deterministically, never executes any model replay mode in question.py
Equation 1: max s − β1·N + β2·N/max(1,k), averaged over all worlds Question.equation1
Stage 3 dreaming: a fixed LLM developer agent rewrites the policy code, M revisions dreamloop/dream.py + prompts.DREAM_DEVELOPER
argmax selection (incumbent included as m=0; replay score never regresses) + redeployment dream.py, chosen/redeploy logic
Only the policy changes data/policies/policy.py is the single rewritten file

Replay reveal rule (the paper's own semantics): picking the root reveals the root's earliest-created unrevealed child (opens a new branch); picking a non-root node reveals its only recorded child, if any; otherwise that slot of the round stays empty.

Prefix visibility is enforced by construction: the policy receives a facade object (Question.view) that exposes only the observation/decision surface — the tree, the revealed set, and the trail are unreachable. Bypassing via reflection is treated as cheating: a revision that reveals nodes in zero rounds is disqualified entirely (V=None, excluded from argmax).

Task domains

Two original domains plus five domains from the paper's experiment section, each a task package in dreamloop/tasks/ verified by dreamloop/verifiers/. All candidate code runs in a restricted subprocess (python3 -I + rlimits + private process group, under the dlbox sandbox when available; BLAS threads pinned to 1 for fair timing).

--source Task Score s_v (higher is better) Correctness gate
humaneval HumanEval code synthesis per-assert pass ratio all asserts must pass
lean Lean 4 theorem proving (8 core-library tasks built in) binary lean kernel accepts, no sorry
lasso Lasso regularization-path solver (17 synthetic timing + 17 fresh gate instances) geometric-mean speedup R = (∏ t_sklearn_i/t_cand_i)^(1/17); R ≥ 1 beats sklearn per-λ gate F_k(w̃_k) ≤ F_k(w_k^sklearn) + 1e-6 on fresh instances; any failure → 0
packing n-circle packing in the unit square (n ∈ {26, 32}) Σ r_i when feasible; paper optimum 2.635983 (n=26) containment + non-overlap at tolerance 1e-7; violation → 0
sumdiff Sum–Difference problem: choose A ⊂ ℤ maximizing Γ(A) = log(|A+A|/|A|)/log(|A−A|/|A|) Γ(A); paper optimum 1.145427 set validity (distinct ints, ≥2, size/span guardrails); Γ recomputed by the verifier
autocorr Autocorrelation inequality (first variant): f ≥ 0, support [−1/4,1/4], ∫f = 1, minimize Φ₁ = max(f∗f) −Φ₁; paper 1.456375, SOTA 1.453675 non-negativity and normalization at tolerance 1e-6; violation → 0
kernels KernelBench: VGG16, LayerNorm, ConvDiv, ConvMax 1/t_cand_ms (inverse runtime); correctness vs reference via allclose(atol=rtol=1e-2) wrong output → 0; requires torch + CUDA (missing → readable error, score 0; the tree still grows)

Dependencies: the lasso verifier needs scikit-learn (it is both the reference implementation and the timing baseline; reference solutions and reference runtimes are cached per instance to data/cache/paper/lasso/*.npz on first use, so later verifications carry no sklearn overhead). All domains need numpy; kernels additionally needs CUDA torch (see requirements.txt). Lasso instance recipes and timeouts live in the paper: section of config.yaml.

Sandbox and security

Every piece of model-generated code is untrusted and is treated as such:

  • dlbox (sandbox/dlbox/, Rust) is a minimal rootless sandbox built for this project: user + mount + network + PID/IPC/UTS namespaces, whole root remounted read-only (CUDA/numpy stay usable, GPU included), writable surface narrowed to the workspace and a fresh /dev/shm, rlimits on CPU / memory / process count, PR_SET_NO_NEW_PRIVS, and an explicit environment whitelist. Build it with:

    cd sandbox/dlbox && cargo build --release
    # then either: cp target/release/dlbox ../../sandbox/dlbox-bin
    #              or: put it on PATH, or export DREAMLOOP_DLBOX=<path>

    Discovery order at runtime: $DREAMLOOP_DLBOX → sandbox/dlbox-bin → dlbox on PATH. DREAMLOOP_SANDBOX=process force-disables it (bare subprocess + rlimits fallback, which is strictly weaker).

  • Secrets never reach candidate code. API keys are read from the environment only, and candidate subprocesses get a minimal allowlisted environment.

  • Lasso cache hardening: reference .npz files enter the sandbox as copies (cache originals are invisible to candidates, closing the rewrite-t_ref-and-poison-scoring channel), and each file carries an instance signature — any recipe change or corruption triggers a rebuild.

  • Policy code is executed in-process during explore (SIGALRM soft timeout only) and in a subprocess during dream replay. This is the one intentional weak spot: there is no strong isolation anywhere in this repo. For serious runs, put the whole thing in a container.

Provenance and verification status

Definitions were implemented from primary sources and are labeled accordingly in each module's docstring:

  • Lasso objective/gate/scoring, packing constraints, the sumdiff Γ, and the autocorr first-variant Φ₁ + constraints follow the paper text (arXiv 2609.14858, checked against the HTML full text).
  • KernelBench reference implementations are embedded verbatim and verified against the upstream main branch (VGG16 = level3/11, LayerNorm = level1/40, ConvDiv = level2/71, ConvMax = level2/82).
  • Unverified (self-devised protocol, the paper gives no values): the exact lasso instance recipes (following SimpleTES), the autocorr grid N=1024, and the sumdiff size/span guardrails. tests/test_autocorr_sensitivity.py demonstrates the autocorr ranking is stable across N ∈ {512, 1024, 2048}.
  • Third-party code bundled in this repository:
    • KernelBench reference implementations (dreamloop/tasks/paper/kernels.py) — MIT License, © Scaling Intelligence Lab, Stanford University (upstream).
    • HumanEval is downloaded at runtime from openai/human-eval (MIT), not bundled.
  • The paper's six downstream held-out datasets (Gisette, RCV1, DNA, Leukemia, Colon, Duke Breast) require external download and are not included.

Choices the paper leaves open

The paper fixes only the symbols K1, K2, M, β1, β2. Everything else is this repo's choice, made explicit: the developer agent reuses the config.api endpoint; baseline_score is the task's best historical score; code tasks use per-assert continuous scoring to feed equation 1 (lean falls back to binary); ties keep the incumbent; revisions that crash, time out, reveal nothing, or show tampering are disqualified wholesale (V=None).

Known deviations from the paper

All domain adaptation, not mechanism substitution:

  • The original domains are HumanEval + Lean; the paper's Lasso / math optimization / KernelBench domains are implemented in tasks/paper/ + verifiers/paper/. The paper's "workspace snapshot" is realized as a text context (inherited history + task statement).
  • A "discovery" manifests as iterative improvement of failed attempts (parent chains), rather than the paper's refinement of engineering artifacts.

Configuration

All knobs live in config.yaml: W (workers) / K1 (rounds) / K2 (replay_rounds) / M (revisions) / β1 / β2 / temperatures / soft timeouts / policy_config (passed to OptimalPolicy.config). Relative paths are anchored to the config file's directory, so the CLI behaves identically from any cwd.

Environment variables:

Variable Meaning
DREAMLOOP_API_KEY API key for the OpenAI-compatible endpoint (required outside mock mode)
DREAMLOOP_MOCK 1 = deterministic stubs for every LLM call (offline self-checks)
DREAMLOOP_DLBOX explicit path to the dlbox binary
DREAMLOOP_SANDBOX process = disable dlbox, use the bare-subprocess fallback

Repository layout

dream-loop/
├── dreamloop/               # the package
│   ├── cli.py               # entry point: python -m dreamloop <command>
│   ├── agent.py, llm.py     # OpenAI-compatible clients (line A / line B)
│   ├── wake.py              # daytime loop: solve → verify → hippocampus
│   ├── hippocampus.py       # append-only experience store (flock-guarded)
│   ├── night/               # SFT/DPO corpus building, LoRA training, probe gate
│   ├── question.py          # the policy-facing world: one facade, two modes
│   ├── discovery.py         # discovery trees — the replay "worlds"
│   ├── explore.py           # Stage 1: online exploration
│   ├── dream.py             # Stage 2+3: replay scoring, rewrite, argmax
│   ├── policy/              # built-in hand-written exploration policy
│   ├── tasks/               # humaneval, lean, and the five paper domains
│   ├── verifiers/           # code (binary + graded), lean, paper-domain verifiers
│   └── prompts.py           # every LLM prompt in one place
├── sandbox/dlbox/           # Rust rootless sandbox (source of truth)
├── tests/                   # dependency-light self-tests, no pytest required
├── probes/                  # frozen probe set for the gate (regenerate with
│                            #   `python -m dreamloop init-probes`; not shipped)
├── config.yaml              # every knob
└── data/                    # runtime artifacts (gitignored, created on first run)

Testing

python3 tests/test_replay.py             # replay semantics, equation 1, determinism
python3 tests/test_graded.py             # scorer: partial credit, crashes, quotas
python3 tests/test_paper_tasks.py        # all domains; uses the GPU when present
python3 tests/test_autocorr_sensitivity.py

All self-tests are offline and free; test_paper_tasks exercises real GPU and sklearn paths when those are available and readable-error paths when not.

Limitations

  • No strong isolation anywhere (see Sandbox and security). Containerize before running untrusted workloads seriously.
  • The lasso t_ref denominator is frozen once at cache-build time; under changing machine load, scores near R = 1.0 will wobble — an inherent trade-off of the reference cache.
  • Line A's training assumes a single GPU box with torch transformers peft trl datasets installed; the rest of the pipeline runs CPU-only.

Citation

If you use this implementation, please cite the paper it implements:

@misc{dreamrsi2026,
  title         = {Dream-RSI: Recursive Self-Improvement through Evolving Worlds},
  year          = {2026},
  howpublished  = {\url{https://arxiv.org/abs/2609.14858}},
  note          = {arXiv:2609.14858}
}

License

This repository is released under the Apache License 2.0.

Third-party code bundled here keeps its own license: the KernelBench reference implementations in dreamloop/tasks/paper/kernels.py are MIT (© Scaling Intelligence Lab, Stanford University, upstream).

Acknowledgments

This project would not exist without Dream-RSI: Recursive Self-Improvement through Evolving Worlds (Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo — arXiv:2609.14858). We are grateful to the authors for publishing the mechanism with enough precision to build on. Thanks also to the OpenAI HumanEval team and to Scaling Intelligence Lab's KernelBench for the benchmarks this testbed stands on.

About

Self-improvement research testbed for LLMs: Dream-RSI-style policy evolution (arXiv:2609.14858) + STaR-style weight upgrades, with a Rust rootless sandbox. Apache-2.0.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages