A production-quality Rust implementation of the Recursive Language Model architecture from the MIT paper "RLM: A Recursive Language Model for Long Contexts".
RLM is a novel prompting technique that enables language models to process inputs far exceeding their native context window by:
- Context Offloading β Moving context into an interactive REPL environment instead of the prompt
- Recursive Decomposition β Breaking complex tasks into subtasks via recursive
llm_query()calls - Minimal Overhead β Achieving linear or sub-linear complexity compared to quadratic costs of full-context processing
| Task | Context Size | Baseline Acc | RLM Acc | Complexity |
|---|---|---|---|---|
| S-NIAH | 8K-128K | 95.0% | 97.5% | O(1) |
| OOLONG | 32K-131K | 72.3% | 89.4% | O(n) |
| OOLONG-Pairs | 8K-32K | 45.2% | 78.6% | O(nΒ²) |
| BrowseComp | 100 docs | 61.8% | 83.2% | O(n log n) |
| Code Repos | ~50K tokens | 68.5% | 84.1% | O(n) |
This implementation has three primary objectives:
Integrate RLM as an execution strategy within the Prometheus Universal Agent Runtime (UAR):
- Expose RLM as a UAR adapter that implements UAR's execution ports
- Stream RLM events (REPL operations, recursive calls, context chunks) as UAR-normalized events
- Support all UAR protocol surfaces: OpenAI REST, MCP, A2A, AG-UI
- Enable workflows to trigger RLM execution for long-context tasks
Provide JavaScript/WASM bindings for direct integration into Cherry Studio:
- Compile
rlm-coreandrlm-repl-rhaito WASM - Export simple FFI:
rlm_execute(request, event_callback) - Stream events to Cherry Studio UI for real-time visualization
- Support configuration via Cherry Studio settings
Build a reference implementation following Rust best practices:
- Safety-first design (no panics, proper error handling)
- Async-first with Tokio
- Minimal dependencies in core
- Comprehensive testing (unit, integration, golden fixtures)
- Full compliance with Rust API Guidelines
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β UAR / Cherry Studio β
β (OpenAI, MCP, A2A, AG-UI) β
ββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββ
β rlm-uar-adapter β β UAR integration
β (maps UAR β RLM events) β
βββββββββββββββ¬ββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββ
β rlm-server β β HTTP + SSE streaming
β (Axum, SSE events) β
βββββββββββββββ¬ββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββ
β rlm-core β β Core executor
β (RlmExecutor + ports) β
βββββββββββββββ¬ββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββ
β rlm-repl-rhai β β Rhai REPL backend
β (sandboxed execution) β
βββββββββββββββββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββ
β rlm-ffi β β WASM bindings
β (wasm-bindgen exports) β
βββββββββββββββββββββββββββββββ
| Crate | Purpose | Dependencies |
|---|---|---|
rlm-core |
Core types, executor, ports | serde, thiserror, tracing (minimal) |
rlm-repl-rhai |
Rhai-based REPL backend | rhai, tokio |
rlm-server |
HTTP server with SSE streaming | axum, tokio, tower |
rlm-ffi |
WASM FFI for Cherry Studio | wasm-bindgen, serde-wasm-bindgen |
rlm-uar-adapter |
UAR integration adapter | rlm-core, UAR types |
# Clone the repository
git clone https://github.com/prometheus/rlm.git
cd rlm
# Build the workspace
cargo build --release
# Run the server
cargo run --bin rlm-server -- --port 8080# Build WASM module
cd rlm-ffi
wasm-pack build --target web --out-dir pkg
# Use in Cherry Studio
import init, { rlm_execute } from './rlm-ffi/pkg/rlm_ffi.js';
await init();
const events = [];
await rlm_execute(request, (event) => {
events.push(event);
console.log('RLM Event:', event);
});use rlm_core::{RlmExecutor, RlmRequest, RlmConfig};
use rlm_repl_rhai::RhaiReplBackend;
#[tokio::main]
async fn main() {
let config = RlmConfig::default();
let repl = RhaiReplBackend::new();
let llm = MyLlmProvider::new();
let executor = RlmExecutor::new(config, repl, llm);
let request = RlmRequest {
query: "Summarize key points from these 100 documents".to_string(),
context: load_documents(),
max_iterations: 50,
recursion_depth: 1,
};
let response = executor.execute(request).await?;
println!("Answer: {}", response.answer);
}# Start server
cargo run --bin rlm-server
# Make request
curl -X POST http://localhost:8080/v1/rlm/execute \
-H "Content-Type: application/json" \
-H "Accept: text/event-stream" \
-d '{
"query": "Find the needle in the haystack",
"context": "...",
"max_iterations": 50
}'use rlm_uar_adapter::RlmUarAdapter;
use uar::{UarRuntime, ExecutionRequest};
let runtime = UarRuntime::new();
runtime.register_adapter("rlm", RlmUarAdapter::new());
let request = ExecutionRequest {
strategy: "rlm".to_string(),
messages: vec![...],
tools: vec![...],
};
let mut stream = runtime.execute(request).await?;
while let Some(event) = stream.next().await {
match event {
UarEvent::RlmReplOp(op) => { /* handle REPL operation */ },
UarEvent::RlmRecursiveCall(call) => { /* handle recursive call */ },
UarEvent::Chunk(chunk) => { /* handle LLM chunk */ },
_ => {}
}
}RLM implements the three-stage process from the paper:
# Instead of:
prompt = f"{long_context}\n\nQuestion: {query}"
# RLM does:
repl.set_variable("context", long_context)
prompt = f"Context is in REPL variable 'context'. Question: {query}"The LLM can call llm_query(subquery) to spawn recursive sub-calls:
# REPL available function
def llm_query(q: str) -> str:
"""Recursively query the LLM with reduced context"""
return rlm_executor.execute_recursive(q, current_depth + 1)Sub-call results are stored in REPL and aggregated:
results = []
for item in context:
results.append(llm_query(f"Process {item}"))
final_answer = aggregate(results)[rlm]
max_iterations = 50 # Max REPL iterations per execution
recursion_depth = 1 # Max recursive call depth
chunk_size = 4096 # Context chunk size (tokens)
timeout_seconds = 300 # Execution timeout
enable_code_exec = true # Allow REPL code execution
streaming = true # Enable SSE streaming
[rlm.llm]
provider = "openai"
model = "gpt-4-turbo"
temperature = 0.0
max_tokens = 4096
[rlm.repl]
backend = "rhai" # "rhai" or "python" (future)
sandbox = true # Enable sandboxing
max_memory_mb = 512 # Memory limit# Run all tests
cargo test --workspace
# Run golden tests (paper benchmarks)
cargo test --package rlm-core --test golden_tests
# Run with coverage
cargo tarpaulin --workspace --out Html
# Benchmark against paper results
cargo bench --package rlm-coreThe rlm-test-utils crate provides comprehensive integration testing capabilities with real LLM providers and secure key vault management. This enables 100% integration test coverage with actual OpenAI GPT-4/GPT-5 models.
| Provider | Use Case | Features |
|---|---|---|
| Local Files | Development & CI | File-based secrets, fast setup |
| Supabase Vault | Production | Database-backed, RLS policies, self-hosted/cloud |
| AWS Secrets Manager | Enterprise | AWS native, fine-grained IAM |
| Azure Key Vault | Enterprise | Azure native, AD integration |
- Create local secrets file:
mkdir -p ~/.rlm/secrets
cat > ~/.rlm/secrets/test-secrets.json << EOF
{
"openai_api_key": "sk-proj-your-openai-key-here",
"anthropic_api_key": "sk-ant-your-anthropic-key-here"
}
EOF- Set environment variable:
export RLM_LOCAL_SECRETS_FILE="~/.rlm/secrets/test-secrets.json"- Run integration tests:
# Run with local vault (default)
cargo test --package rlm-test-utils
# Run with all vault providers
cargo test --package rlm-test-utils --features all-providersSupabase provides the most comprehensive vault solution with database-backed storage, Row Level Security (RLS), and support for both self-hosted and cloud installations.
- Environment variables:
export SUPABASE_URL="https://your-project.supabase.co"
export SUPABASE_SERVICE_ROLE_KEY="your-service-role-key"
export SUPABASE_SECRETS_TABLE="vault_secrets" # optional- Initialize and store secrets:
use rlm_test_utils::SupabaseVaultProvider;
#[tokio::test]
async fn setup_supabase_vault() {
let vault = SupabaseVaultProvider::from_env().await.unwrap();
// Initialize the secrets table with RLS policies
vault.init_table().await.unwrap();
// Store your OpenAI API key securely
vault.store_secret("openai_api_key", "sk-proj-your-key").await.unwrap();
vault.store_secret("anthropic_api_key", "sk-ant-your-key").await.unwrap();
}- Run tests with Supabase:
cargo test --package rlm-test-utils --features supabase-vault- Configure AWS credentials:
export AWS_REGION="us-west-2"
# Use AWS CLI: aws configure
# Or use environment variables: AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY- Create secrets in AWS:
aws secretsmanager create-secret \
--name "openai_api_key" \
--secret-string "sk-proj-your-openai-key"
aws secretsmanager create-secret \
--name "anthropic_api_key" \
--secret-string "sk-ant-your-anthropic-key"- Run tests:
cargo test --package rlm-test-utils --features aws-secrets- Environment variables:
export AZURE_KEYVAULT_URL="https://your-vault.vault.azure.net/"
# Use Azure CLI: az login
# Or set: AZURE_CLIENT_ID, AZURE_CLIENT_SECRET, AZURE_TENANT_ID- Create secrets in Azure:
az keyvault secret set --vault-name "your-vault" \
--name "openai-api-key" --value "sk-proj-your-openai-key"
az keyvault secret set --vault-name "your-vault" \
--name "anthropic-api-key" --value "sk-ant-your-anthropic-key"- Run tests:
cargo test --package rlm-test-utils --features azure-keyvaultuse rlm_test_utils::{
TestFramework, LlmTestFramework, VaultFactory,
OpenAiTestProvider, AccuracyEvaluator, PerformanceEvaluator,
MarkdownReportGenerator, LlmTestProviderConfig,
};
#[tokio::test]
async fn comprehensive_rlm_integration_test() {
// 1. Initialize secure vault (auto-detects provider from environment)
let vault = VaultFactory::from_env().await.unwrap();
// 2. Create OpenAI test provider with real API
let config = LlmTestProviderConfig {
model: "gpt-4".to_string(),
api_key_vault_key: "openai_api_key".to_string(),
max_tokens: Some(500),
temperature: Some(0.1),
rate_limit_rps: 1.0, // Respect rate limits
max_retries: 3,
};
let provider = OpenAiTestProvider::new(config, vault.clone()).await.unwrap();
let framework = LlmTestFramework::new(Arc::new(provider), vault);
// 3. Execute RLM request with real LLM
let request = RlmRequest {
query: "What is the main theme of the provided context?".to_string(),
context: "Long document content here...".to_string(),
max_iterations: 10,
recursion_depth: 1,
metadata: HashMap::new(),
};
let response = framework.execute_test(&request).await.unwrap();
// 4. Evaluate response quality
let accuracy = AccuracyEvaluator::fuzzy(0.8, vec![
"theme".to_string(),
"main idea".to_string(),
]);
let result = accuracy.evaluate_response(&request, &response).await.unwrap();
assert!(result.quality_score > 0.8);
// 5. Generate comprehensive report
let report = framework.generate_report("RLM Integration Test").await.unwrap();
let generator = MarkdownReportGenerator::new();
let markdown = generator.generate_report(&report).await.unwrap();
tokio::fs::write("integration-test-report.md", markdown).await.unwrap();
}The framework includes automated testing workflows:
# .github/workflows/integration-tests.yml
name: RLM Integration Tests
on: [push, pull_request, schedule]
jobs:
test:
runs-on: ubuntu-latest
strategy:
matrix:
vault: [local, supabase, aws]
steps:
- uses: actions/checkout@v4
- name: Setup Rust
uses: actions-rs/toolchain@v1
with:
toolchain: stable
- name: Run Integration Tests
env:
# Local vault
RLM_LOCAL_SECRETS_FILE: ${{ secrets.LOCAL_SECRETS_FILE }}
# Supabase vault
SUPABASE_URL: ${{ secrets.SUPABASE_URL }}
SUPABASE_SERVICE_ROLE_KEY: ${{ secrets.SUPABASE_SERVICE_ROLE_KEY }}
# AWS vault
AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
AWS_REGION: us-west-2
# API Keys (stored in your vault)
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
cargo test --package rlm-test-utils --features ${{ matrix.vault }}-vaultThe testing framework includes built-in cost controls:
use rlm_test_utils::CostEvaluator;
let cost_evaluator = CostEvaluator::openai_gpt4()
.with_daily_budget(25.0) // $25 daily limit
.with_max_cost_per_test(0.10); // $0.10 per test limit
// Automatic cost tracking and budget enforcement
let cost_result = cost_evaluator.evaluate_response(&request, &response).await?;
println!("Test cost: ${:.4}", cost_result.metadata.get("total_cost").unwrap());- π API keys never exposed in logs, error messages, or console output
- π Encrypted at rest (vault provider dependent)
- π Encrypted in transit via TLS
- π Audit logging for key access
- π‘οΈ Row-level security for database vaults (Supabase)
- π Key rotation support
Located in tests/fixtures/:
s_niah/β Simple needle-in-haystack (O(1) complexity)oolong/β OOLONG dataset (O(n) aggregation)oolong_pairs/β Pairwise reasoning (O(nΒ²) complexity)browsecomp/β Multi-hop QA over documentscode_repo/β Code understanding tasksstreaming/β Event streaming validation
Based on paper results:
| Complexity Class | Traditional LLM | RLM | Improvement |
|---|---|---|---|
| O(1) β S-NIAH | O(nΒ²) context | O(1) | ~128x |
| O(n) β OOLONG | O(nΒ³) | O(n) | ~nΒ² |
| O(nΒ²) β Pairs | O(nβ΄) | O(nΒ²) | ~nΒ² |
| O(n log n) β Browse | O(nΒ³) | O(n log n) | ~nΒ²/log n |
Memory usage: Linear in context size (vs quadratic for attention)
RLM integrates with UAR via the adapter pattern:
// UAR adapter maps RLM events β UAR events
impl UarExecutionAdapter for RlmUarAdapter {
async fn execute(&self, request: UarRequest) -> UarEventStream {
let rlm_request = self.convert_request(request);
let rlm_stream = self.executor.execute_stream(rlm_request).await?;
rlm_stream.map(|event| match event {
RlmEvent::ReplOp(op) => UarEvent::ToolCall {
tool: "repl.execute",
args: op.code,
},
RlmEvent::RecursiveCall(call) => UarEvent::SubCall {
query: call.query,
depth: call.depth,
},
RlmEvent::Chunk(chunk) => UarEvent::Chunk(chunk),
// ... more mappings
})
}
}UAR can then expose RLM via:
- OpenAI
/v1/chat/completionswithstrategy: "rlm" - MCP tool:
rlm.execute - A2A event:
rlm.stream - AG-UI widget: RLM visualization
Cherry Studio uses the WASM FFI:
import init, { RlmExecutor } from '@rlm/ffi';
await init();
const executor = new RlmExecutor({
maxIterations: 50,
recursionDepth: 1,
});
// Stream events to UI
executor.execute(request, (event) => {
if (event.type === 'repl_op') {
ui.showReplExecution(event.code, event.result);
} else if (event.type === 'recursive_call') {
ui.showRecursiveCall(event.query, event.depth);
} else if (event.type === 'chunk') {
ui.appendChunk(event.content);
}
});- β Core types and executor
- β Rhai REPL backend
- β HTTP server with SSE
- β Golden test fixtures
- β¬ UAR adapter
- β¬ WASM FFI
- β¬ Cherry Studio plugin
- β¬ Performance benchmarks
- β¬ Python REPL backend
- β¬ Workflow integration
- β¬ Custom decomposition strategies
- β¬ Distributed execution
MIT
Contributions welcome! Please read docs/IMPLEMENTATION_PLAN.md for development guidelines.
Status: π§ Active Development | Target: Production-ready v0.1 by Q2 2024