Skip to content
acounts1112027-cloudPublic
forked from denni9999/Gradient

About

Gradient

Resources

Stars

1 star

Watchers

0 watching

Forks

 
 

Repository files navigation

gradient

Gradient: Research Agents That Learn From Experience

Research Agent • Training • Evaluation • OpenPipe ART

Gradient is an open-source reference system for training tool-using research agents with reinforcement learning. It is designed around two core abstractions:

  • The Research Environment places an agent inside a reproducible company workspace with a defined task, a limited set of tools, and outcomes that can be checked against known facts and sources.
  • The Learning Loop records complete tool-use trajectories, scores the result, trains the model with GRPO, and evaluates the trained adapter against the base model on held-out tasks.

Gradient makes the full agent reinforcement-learning workflow concrete and inspectable:

  • Work happens inside an environment: the included workspace contains emails, contracts, policies, meeting notes, customer records, and deliberately distributed evidence.
  • Agents use tools: agents can search the full workspace, search emails or documents separately, open individual records, and submit answers with cited sources.
  • The full trajectory is recorded: searches, opened records, tool outputs, final answers, citations, and reward components are stored for inspection.
  • Rewards measure research quality: each episode is scored for answer correctness, citation quality, and tool efficiency.
  • Models learn with GRPO: Gradient uses OpenPipe ART to generate trajectory groups and train lightweight model adapters.
  • Improvement is evaluated on unseen work: held-out tasks compare reward, correctness, citation F1, efficiency, malformed responses, and average tool calls.

Getting Started

Gradient requires Python 3.12 or later.

git clone https://github.com/RobertGolds1/Gradient.git
cd Gradient

python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"

Run the included research-agent demo:

gradient-demo

The demo compares a weak policy with a stronger research policy on the same task, then writes the complete result to artifacts/demo.json.

Run the held-out evaluation suite:

gradient-eval

Run the training pipeline without starting a model-training job:

gradient-train --dry-run --policy heuristic --max-steps 1

Training With Reinforcement Learning

The Learning Loop

Gradient uses OpenPipe ART and its serverless backend for GRPO training. Copy the environment template and add your Weights & Biases API key:

cp .env.example .env
WANDB_API_KEY=your_api_key

Start a training run:

gradient-train --policy art --max-steps 2

The default base model is OpenPipe/Qwen3-14B-Instruct. Training parameters are defined in gradient_rl/training/config.py.

Serverless training may use paid compute. Review the configuration before starting a run.

The Research Environment

The Research Environment

The included environment is a synthetic company knowledge workspace built to test retrieval, evidence selection, multi-step reasoning, and disciplined tool use.

It contains:

  • 118 workspace records across emails and documents;
  • 50 reinforcement-learning tasks;
  • 20 held-out evaluation tasks;
  • known facts and source IDs for deterministic scoring; and
  • distractor records that make simple keyword matching unreliable.

The agent has six tools:

search(query)
search_email(query)
search_documents(query)
open_email(id)
open_document(id)
submit_answer(answer, sources)

Keeping the tool surface small isolates the learning problem. The objective is to improve how the model searches, gathers evidence, connects information across sources, and decides when it has enough support to answer.

Rewards

Gradient separates research quality into three visible signals:

  • Correctness: how many required facts appear in the answer;
  • Citation quality: citation precision and recall against the required sources; and
  • Efficiency: whether the agent completed the task without unnecessary tool calls.

The default reward is:

0.50 × correctness + 0.30 × citation quality + 0.20 × efficiency

Malformed trajectories receive zero reward. Every component remains available in the generated artifacts so a high aggregate score can be inspected rather than accepted at face value.

Training and Evaluation Artifacts

Held-Out Evaluation

Gradient writes its outputs to artifacts/:

  • complete training trajectories in JSONL;
  • per-step training metrics;
  • evaluation trajectories for each policy;
  • aggregate evaluation results; and
  • a readable comparison of baseline and improved research behavior.

This makes it possible to inspect not only whether a score changed, but how the agent's behavior changed.

Built as an Open Research Baseline

Gradient currently supports one focused workflow end to end: run a research agent, capture its behavior, assign rewards, train it with reinforcement learning, and evaluate what changed.

The current environment is intentionally narrow. It is a baseline for studying research behavior, reward design, and agent training, not a finished general-purpose agent platform. New tasks, tools, reward components, models, and environments can be added while keeping the same trajectory and evaluation interfaces.

Documentation

Development

Run the test suite:

pytest -q

Acknowledgements

Gradient uses OpenPipe ART for reinforcement-learning training, adapters, and inference. The initial implementation draws on OpenPipe's open-source email-deep-research project as a reference for training tool-using research agents.

License

Gradient is fully open source and released under the Apache License 2.0.

About

Gradient

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages