Reference implementation of “Smarter by the Moment: Environment-Driven Dynamic Policies for Continual LLM Improvement” (COLM 2026).
Paper (arXiv:2609.16800) · Project page · 繁體中文說明
DRPG lets an LLM agent improve while it works through a stream of tasks, without touching its weights. At each step the agent answers with the help of two things drawn from its own history: retrieved examples it previously got right, and a short policy: a handful of actionable rules that a separate policy generator distills from past correct and incorrect cases. The environment returns binary correctness feedback, that feedback goes into memory, and the next policy is written from a slightly wiser history.
The difference from prior memory-based methods is what gets reused. Self-StreamICL reuses instances; DRPG also reuses strategies, so a mistake that keeps recurring can be named and avoided rather than merely re-retrieved.
t: receive x_t
R_t = retrieve(x_t, correct cases) # for the agent
R'_t = retrieve(x_t, correct) ∪ retrieve(x_t, wrong) # contrastive, k/2 each
P_t = policy_generator(R'_t) # ≤ 5 bullet points
ŷ_t = agent(x_t, R_t, P_t)
fb_t = environment(x_t, ŷ_t) # 0 / 1
memory ← (x_t, ŷ_t, fb_t)
Across six benchmarks and seven LLMs, DRPG beats Zero-shot, Self-Refine and Self-StreamICL on most model × dataset pairs, at the cost of roughly one extra LLM call per query.
| Path | What it is |
|---|---|
stream_bench/agents/policy.py |
DRPG (agent_name: rag_policy): retrieval, policy generation, prediction, memory update |
stream_bench/agents/policy_rag.py |
DRPG variant whose memory also stores the policy in force at write time |
stream_bench/agents/dynamic_cheatsheet.py |
Dynamic Cheatsheet, used as the no-feedback ablation |
stream_bench/agents/fewshot_rag.py |
Self-StreamICL and MemPrompt |
stream_bench/agents/iter_prompt.py |
Self-Refine |
stream_bench/agents/zeroshot.py |
Zero-shot |
stream_bench/agents/utils.py |
The BGE + FAISS memory (RAG) and the backend factory (get_llm) |
stream_bench/benchmarks/ |
The six streaming environments and their metrics |
stream_bench/llms/ |
One thin client per LLM provider |
stream_bench/pipelines/run_bench.py |
The evaluation loop, and the entry point for every run |
configs/agent/, configs/bench/ |
Method and benchmark configs (generated by scripts/gen_configs.py) |
docs/configuration.md |
Every configuration field, and what changing it does |
docs/reproduce.md |
The exact grid behind the paper's tables |
This repository is a trimmed fork of StreamBench (Apache-2.0) containing the six benchmarks and the methods used in the paper. See NOTICE for what came from where.
Requires Python 3.10+. All commands use uv; dependencies live in pyproject.toml and are pinned in uv.lock.
git clone https://github.com/tingwei161803/drpg.git
cd drpg
uv syncDS-1000 grades a solution by executing it in the same process, so the libraries its tasks are written against have to be importable too. Add them only if you plan to run DS-1000:
uv sync --extra ds1000Because that evaluation really does run the generated code, DS-1000 scores are sensitive to the versions of NumPy, pandas and friends: a handful of reference outputs change across major releases.
DDXPlus, HotpotQA and DS-1000 are pulled from the appier-ai-research/StreamBench dataset on first use, so there is nothing to do for them.
The three text-to-SQL benchmarks additionally need their SQLite databases, which land in ./data:
uv run download_text2sql_data.pyEvery backend reads its key from the environment; no key is ever written to a config file.
cp .env.example .env
$EDITOR .env
source .envThe paper's runs used Google AI Studio (series: gemini_dev, needs GOOGLE_API_KEY) and NVIDIA NIM (series: nvidia, needs NVIDIA_API_KEY). Both had free tiers that were sufficient. Other providers are listed in .env.example.
Before spending a single token, run the oracle agent: it answers every question with the ground-truth label, so a score near 100% means the dataset, the feedback loop and the metric are wired up correctly. It needs no API key.
uv run python -m stream_bench.pipelines.run_bench \
--agent_cfg configs/agent/gt.yml \
--bench_cfg configs/bench/ddxplus.ymlYou can also check every agent config without any network access at all:
uv run scripts/validate_configs.pyuv run python -m stream_bench.pipelines.run_bench \
--agent_cfg configs/agent/drpg/gemini-2.0-flash.yml \
--bench_cfg configs/bench/spider.ymlThat streams the whole Spider test split through DRPG and prints the final metric. Add --use_wandb --entity "$WANDB_ENTITY" to log each step to Weights & Biases, and --slow_task (5 s between steps) or --slow_slow_task (10 s) if your provider rate-limits you.
Per-step records (prompts, raw outputs, token counts, the generated policy) are appended to log/<benchmark>/test/<run name>.jsonl, and the memory index is persisted next to it as .jsonl.db.
scripts/run_drpg.sh runs one method across all six benchmarks.
| Paper | Config | agent_name |
|---|---|---|
| DRPG (Ours) | configs/agent/drpg/<model>.yml |
rag_policy |
| DRPG, cross-generator (Sec. 5.3) | configs/agent/drpg-cross-generator/<agent>__policy-<generator>.yml |
rag_policy |
| Dynamic Cheatsheet, no feedback (App. F) | configs/agent/dynamic_cheatsheet/<model>.yml |
dynamic_cheatsheet |
| Self-StreamICL (main baseline) | configs/agent/self_stream_icl/<model>.yml |
self_stream_icl |
| Self-Refine | configs/agent/self_refine/<model>.yml |
self_refine |
| Zero-shot | configs/agent/zeroshot/<model>.yml |
zeroshot |
| Config | Task | Metric | Stream length |
|---|---|---|---|
configs/bench/spider.yml |
Text-to-SQL | Execution accuracy | 2,147 |
configs/bench/cosql.yml |
Text-to-SQL (conversational) | Execution accuracy | 1,007 |
configs/bench/bird.yml |
Text-to-SQL (hard) | Execution accuracy | 1,534 |
configs/bench/hotpotqa.yml |
Multi-hop QA, distractor setting | Exact match | 1,500 |
configs/bench/ddxplus.yml |
Medical diagnosis, 49 classes | Accuracy | 1,764 |
configs/bench/ds_1000.yml |
Python data science | pass@1 | 1,000 |
Order is fixed by seed: 42, so every method sees the same stream.
A different model. Add it to MODELS in scripts/gen_configs.py and regenerate:
uv run scripts/gen_configs.py
uv run scripts/validate_configs.py # instantiates all configs, no networkseries picks a client from stream_bench/llms/; model_name is passed through to that provider.
A different policy generator. Point policy_llm at another model; it does not have to match llm. Section 5.3 of the paper shows smaller and cross-family generators working well.
A new agent. Subclass Agent in stream_bench/agents/base.py, implement __init__, __call__ and update, then register it in stream_bench/agents/__init__.py.
A new LLM backend. Subclass LLM in stream_bench/llms/base.py, return (text, info) with the token counts filled in, and add a branch to get_llm in stream_bench/agents/utils.py.
Every knob (policy_generator_mode, use_previous_policy, use_correct_fewshots, top_k, and the rest) is documented in docs/configuration.md.
docs/reproduce.md lists the grid: which config produces which table, and what to expect. One caveat worth stating up front: all seven models are hosted APIs, and API-served checkpoints drift or get retired. Re-running today reproduces the pattern of the results, not the exact decimals.
@inproceedings{chang2026drpg,
title = {Smarter by the Moment: Environment-Driven Dynamic Policies for
Continual {LLM} Improvement},
author = {Chang, Ting-Wei and Chen, Po-Chun and Huang, Hen-Hsen and Chen, Hsin-Hsi},
booktitle = {Conference on Language Modeling (COLM)},
year = {2026}
}If you use the benchmark environments, please also cite StreamBench:
@article{wu2024streambench,
title = {StreamBench: Towards Benchmarking Continuous Improvement of Language Agents},
author = {Wu, Cheng-Kuang and Tam, Zhi Rui and Lin, Chieh-Yen and Chen, Yun-Nung and Lee, Hung-yi},
journal = {Advances in Neural Information Processing Systems},
volume = {37},
pages = {107039--107063},
year = {2024}
}Apache License 2.0. See LICENSE and NOTICE. The benchmark datasets remain under their own licenses.
