Skip to content

Repository files navigation

DRPG: Dynamic Retrieval-based Policy Generation

Reference implementation of “Smarter by the Moment: Environment-Driven Dynamic Policies for Continual LLM Improvement” (COLM 2026).

Paper (arXiv:2609.16800) · Project page · 繁體中文說明

The DRPG framework

DRPG lets an LLM agent improve while it works through a stream of tasks, without touching its weights. At each step the agent answers with the help of two things drawn from its own history: retrieved examples it previously got right, and a short policy: a handful of actionable rules that a separate policy generator distills from past correct and incorrect cases. The environment returns binary correctness feedback, that feedback goes into memory, and the next policy is written from a slightly wiser history.

The difference from prior memory-based methods is what gets reused. Self-StreamICL reuses instances; DRPG also reuses strategies, so a mistake that keeps recurring can be named and avoided rather than merely re-retrieved.

t:  receive x_t
    R_t   = retrieve(x_t, correct cases)                # for the agent
    R'_t  = retrieve(x_t, correct) ∪ retrieve(x_t, wrong)   # contrastive, k/2 each
    P_t   = policy_generator(R'_t)                      # ≤ 5 bullet points
    ŷ_t   = agent(x_t, R_t, P_t)
    fb_t  = environment(x_t, ŷ_t)                       # 0 / 1
    memory ← (x_t, ŷ_t, fb_t)

Across six benchmarks and seven LLMs, DRPG beats Zero-shot, Self-Refine and Self-StreamICL on most model × dataset pairs, at the cost of roughly one extra LLM call per query.


Repository layout

Path What it is
stream_bench/agents/policy.py DRPG (agent_name: rag_policy): retrieval, policy generation, prediction, memory update
stream_bench/agents/policy_rag.py DRPG variant whose memory also stores the policy in force at write time
stream_bench/agents/dynamic_cheatsheet.py Dynamic Cheatsheet, used as the no-feedback ablation
stream_bench/agents/fewshot_rag.py Self-StreamICL and MemPrompt
stream_bench/agents/iter_prompt.py Self-Refine
stream_bench/agents/zeroshot.py Zero-shot
stream_bench/agents/utils.py The BGE + FAISS memory (RAG) and the backend factory (get_llm)
stream_bench/benchmarks/ The six streaming environments and their metrics
stream_bench/llms/ One thin client per LLM provider
stream_bench/pipelines/run_bench.py The evaluation loop, and the entry point for every run
configs/agent/, configs/bench/ Method and benchmark configs (generated by scripts/gen_configs.py)
docs/configuration.md Every configuration field, and what changing it does
docs/reproduce.md The exact grid behind the paper's tables

This repository is a trimmed fork of StreamBench (Apache-2.0) containing the six benchmarks and the methods used in the paper. See NOTICE for what came from where.


Setup

Requires Python 3.10+. All commands use uv; dependencies live in pyproject.toml and are pinned in uv.lock.

git clone https://github.com/tingwei161803/drpg.git
cd drpg
uv sync

DS-1000 grades a solution by executing it in the same process, so the libraries its tasks are written against have to be importable too. Add them only if you plan to run DS-1000:

uv sync --extra ds1000

Because that evaluation really does run the generated code, DS-1000 scores are sensitive to the versions of NumPy, pandas and friends: a handful of reference outputs change across major releases.

Data

DDXPlus, HotpotQA and DS-1000 are pulled from the appier-ai-research/StreamBench dataset on first use, so there is nothing to do for them.

The three text-to-SQL benchmarks additionally need their SQLite databases, which land in ./data:

uv run download_text2sql_data.py

API keys

Every backend reads its key from the environment; no key is ever written to a config file.

cp .env.example .env
$EDITOR .env
source .env

The paper's runs used Google AI Studio (series: gemini_dev, needs GOOGLE_API_KEY) and NVIDIA NIM (series: nvidia, needs NVIDIA_API_KEY). Both had free tiers that were sufficient. Other providers are listed in .env.example.

Check the setup

Before spending a single token, run the oracle agent: it answers every question with the ground-truth label, so a score near 100% means the dataset, the feedback loop and the metric are wired up correctly. It needs no API key.

uv run python -m stream_bench.pipelines.run_bench \
    --agent_cfg configs/agent/gt.yml \
    --bench_cfg configs/bench/ddxplus.yml

You can also check every agent config without any network access at all:

uv run scripts/validate_configs.py

Running DRPG

uv run python -m stream_bench.pipelines.run_bench \
    --agent_cfg configs/agent/drpg/gemini-2.0-flash.yml \
    --bench_cfg configs/bench/spider.yml

That streams the whole Spider test split through DRPG and prints the final metric. Add --use_wandb --entity "$WANDB_ENTITY" to log each step to Weights & Biases, and --slow_task (5 s between steps) or --slow_slow_task (10 s) if your provider rate-limits you.

Per-step records (prompts, raw outputs, token counts, the generated policy) are appended to log/<benchmark>/test/<run name>.jsonl, and the memory index is persisted next to it as .jsonl.db.

scripts/run_drpg.sh runs one method across all six benchmarks.

The methods

Paper Config agent_name
DRPG (Ours) configs/agent/drpg/<model>.yml rag_policy
DRPG, cross-generator (Sec. 5.3) configs/agent/drpg-cross-generator/<agent>__policy-<generator>.yml rag_policy
Dynamic Cheatsheet, no feedback (App. F) configs/agent/dynamic_cheatsheet/<model>.yml dynamic_cheatsheet
Self-StreamICL (main baseline) configs/agent/self_stream_icl/<model>.yml self_stream_icl
Self-Refine configs/agent/self_refine/<model>.yml self_refine
Zero-shot configs/agent/zeroshot/<model>.yml zeroshot

The benchmarks

Config Task Metric Stream length
configs/bench/spider.yml Text-to-SQL Execution accuracy 2,147
configs/bench/cosql.yml Text-to-SQL (conversational) Execution accuracy 1,007
configs/bench/bird.yml Text-to-SQL (hard) Execution accuracy 1,534
configs/bench/hotpotqa.yml Multi-hop QA, distractor setting Exact match 1,500
configs/bench/ddxplus.yml Medical diagnosis, 49 classes Accuracy 1,764
configs/bench/ds_1000.yml Python data science pass@1 1,000

Order is fixed by seed: 42, so every method sees the same stream.


Changing things

A different model. Add it to MODELS in scripts/gen_configs.py and regenerate:

uv run scripts/gen_configs.py
uv run scripts/validate_configs.py   # instantiates all configs, no network

series picks a client from stream_bench/llms/; model_name is passed through to that provider.

A different policy generator. Point policy_llm at another model; it does not have to match llm. Section 5.3 of the paper shows smaller and cross-family generators working well.

A new agent. Subclass Agent in stream_bench/agents/base.py, implement __init__, __call__ and update, then register it in stream_bench/agents/__init__.py.

A new LLM backend. Subclass LLM in stream_bench/llms/base.py, return (text, info) with the token counts filled in, and add a branch to get_llm in stream_bench/agents/utils.py.

Every knob (policy_generator_mode, use_previous_policy, use_correct_fewshots, top_k, and the rest) is documented in docs/configuration.md.


Reproducing the paper

docs/reproduce.md lists the grid: which config produces which table, and what to expect. One caveat worth stating up front: all seven models are hosted APIs, and API-served checkpoints drift or get retired. Re-running today reproduces the pattern of the results, not the exact decimals.


Citation

@inproceedings{chang2026drpg,
  title     = {Smarter by the Moment: Environment-Driven Dynamic Policies for
               Continual {LLM} Improvement},
  author    = {Chang, Ting-Wei and Chen, Po-Chun and Huang, Hen-Hsen and Chen, Hsin-Hsi},
  booktitle = {Conference on Language Modeling (COLM)},
  year      = {2026}
}

If you use the benchmark environments, please also cite StreamBench:

@article{wu2024streambench,
  title   = {StreamBench: Towards Benchmarking Continuous Improvement of Language Agents},
  author  = {Wu, Cheng-Kuang and Tam, Zhi Rui and Lin, Chieh-Yen and Chen, Yun-Nung and Lee, Hung-yi},
  journal = {Advances in Neural Information Processing Systems},
  volume  = {37},
  pages   = {107039--107063},
  year    = {2024}
}

License

Apache License 2.0. See LICENSE and NOTICE. The benchmark datasets remain under their own licenses.

About

Code for the COLM 2026 paper "Smarter by the Moment: Environment-Driven Dynamic Policies for Continual LLM Improvement" (DRPG)

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages