An implementation of GRPO for Unsloth's VLMs training
-
Updated
Aug 7, 2025 - Python
An implementation of GRPO for Unsloth's VLMs training
PARL (Parallel-Agent Reinforcement Learning) is a training paradigm that teaches models to decompose complex tasks into parallel subtasks and coordinate multiple agents simultaneously.
Your efficient and accurate answer verification system for RL training.
simpleR1: A Simple Framework for Training R1-like Models
Recreating the minimal training methods of DeepSeek-R1 for small langauge models.
Public showcase of NovaLiveSystem: a biomimetic cognitive architecture with interoception and distributed intelligence.
This project implements the reasoning training pipeline introduced in the DeepSeek-R1 paper, applying Group Relative Policy Optimization (GRPO) to teach a 7B language model to reason step-by-step through mathematical problems.
Low-cost GRPO fine-tuning pipeline for valence-arousal state estimation and tone-aligned responses in Llama 3.2 1B.
To associate your repository with the grpotrainer topic, visit your repo's landing page and select "manage topics."