Skip to content

About

Reinforcement Learning Coding Agent with GRPO training

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

478 Commits

Folders and files

Repository files navigation

RL Coding Agent

An autonomous coding agent trained using Group Relative Policy Optimization (GRPO) — a lightweight RL algorithm — on real-world software engineering tasks from SWE-bench.

What it does

  • Loads a coding issue from SWE-bench
  • Runs a language model (Qwen 2.5 Coder 0.5B) as the "agent" through a mock shell environment
  • Samples multiple rollouts of the agent attempting to fix the bug
  • Scores each rollout (reward function)
  • Performs one GRPO training step using LoRA — updating the model to prefer higher-reward actions

Architecture

rl_coding_agent/
├── agent.py            # Core logic: load model, run agent, GRPO step
├── utils.py            # MockEnv, generate(), shell parser, system prompt
├── api/
│   └── index.py        # Vercel serverless entry point
├── requirements.txt    # Python dependencies
├── vercel.json         # Vercel deployment config
└── README.md

How to run locally

# Install dependencies
pip install -r requirements.txt

# Run the agent
python agent.py

How to deploy on Vercel

npm i -g vercel
vercel --prod

Tech Stack

  • Qwen/Qwen2.5-Coder-0.5B-Instruct — the language model / policy
  • transformers + peft — model loading and LoRA fine-tuning
  • datasets — SWE-bench task loading
  • torch — training loop
  • Vercel Python Serverless — deployment

How GRPO works (simply)

  1. Run the agent N times on the same task → get N rollouts
  2. Score each rollout → get N rewards
  3. Compute advantage = how much better was this rollout vs the average?
  4. Update the model to increase the probability of high-advantage actions

That's it. No value network. No critic. Just group-relative comparison.

About

Reinforcement Learning Coding Agent with GRPO training

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages