Research Agent • Training • Evaluation • OpenPipe ART
Gradient is an open-source reference system for training tool-using research agents with reinforcement learning. It is designed around two core abstractions:
- The Research Environment places an agent inside a reproducible company workspace with a defined task, a limited set of tools, and outcomes that can be checked against known facts and sources.
- The Learning Loop records complete tool-use trajectories, scores the result, trains the model with GRPO, and evaluates the trained adapter against the base model on held-out tasks.
Gradient makes the full agent reinforcement-learning workflow concrete and inspectable:
- Work happens inside an environment: the included workspace contains emails, contracts, policies, meeting notes, customer records, and deliberately distributed evidence.
- Agents use tools: agents can search the full workspace, search emails or documents separately, open individual records, and submit answers with cited sources.
- The full trajectory is recorded: searches, opened records, tool outputs, final answers, citations, and reward components are stored for inspection.
- Rewards measure research quality: each episode is scored for answer correctness, citation quality, and tool efficiency.
- Models learn with GRPO: Gradient uses OpenPipe ART to generate trajectory groups and train lightweight model adapters.
- Improvement is evaluated on unseen work: held-out tasks compare reward, correctness, citation F1, efficiency, malformed responses, and average tool calls.
Gradient requires Python 3.12 or later.
git clone https://github.com/RobertGolds1/Gradient.git
cd Gradient
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"Run the included research-agent demo:
gradient-demoThe demo compares a weak policy with a stronger research policy on the same task, then writes the complete result to artifacts/demo.json.
Run the held-out evaluation suite:
gradient-evalRun the training pipeline without starting a model-training job:
gradient-train --dry-run --policy heuristic --max-steps 1Gradient uses OpenPipe ART and its serverless backend for GRPO training. Copy the environment template and add your Weights & Biases API key:
cp .env.example .envWANDB_API_KEY=your_api_keyStart a training run:
gradient-train --policy art --max-steps 2The default base model is OpenPipe/Qwen3-14B-Instruct. Training parameters are defined in gradient_rl/training/config.py.
Serverless training may use paid compute. Review the configuration before starting a run.
The included environment is a synthetic company knowledge workspace built to test retrieval, evidence selection, multi-step reasoning, and disciplined tool use.
It contains:
- 118 workspace records across emails and documents;
- 50 reinforcement-learning tasks;
- 20 held-out evaluation tasks;
- known facts and source IDs for deterministic scoring; and
- distractor records that make simple keyword matching unreliable.
The agent has six tools:
search(query)
search_email(query)
search_documents(query)
open_email(id)
open_document(id)
submit_answer(answer, sources)
Keeping the tool surface small isolates the learning problem. The objective is to improve how the model searches, gathers evidence, connects information across sources, and decides when it has enough support to answer.
Gradient separates research quality into three visible signals:
- Correctness: how many required facts appear in the answer;
- Citation quality: citation precision and recall against the required sources; and
- Efficiency: whether the agent completed the task without unnecessary tool calls.
The default reward is:
0.50 × correctness + 0.30 × citation quality + 0.20 × efficiency
Malformed trajectories receive zero reward. Every component remains available in the generated artifacts so a high aggregate score can be inspected rather than accepted at face value.
Gradient writes its outputs to artifacts/:
- complete training trajectories in JSONL;
- per-step training metrics;
- evaluation trajectories for each policy;
- aggregate evaluation results; and
- a readable comparison of baseline and improved research behavior.
This makes it possible to inspect not only whether a score changed, but how the agent's behavior changed.
Gradient currently supports one focused workflow end to end: run a research agent, capture its behavior, assign rewards, train it with reinforcement learning, and evaluate what changed.
The current environment is intentionally narrow. It is a baseline for studying research behavior, reward design, and agent training, not a finished general-purpose agent platform. New tasks, tools, reward components, models, and environments can be added while keeping the same trajectory and evaluation interfaces.
- Research-agent example
- Agent rollout loop
- Research environment
- Tool definitions
- Reward functions
- GRPO training pipeline
- Held-out evaluation
- Training and evaluation data
Run the test suite:
pytest -qGradient uses OpenPipe ART for reinforcement-learning training, adapters, and inference. The initial implementation draws on OpenPipe's open-source email-deep-research project as a reference for training tool-using research agents.
Gradient is fully open source and released under the Apache License 2.0.



