GRPO reinforcement learning for LLMs on math problems with verifiable reward — trained on Ascend NPU with verl
-
Updated
Aug 30, 2026
GRPO reinforcement learning for LLMs on math problems with verifiable reward — trained on Ascend NPU with verl
A verifiable benchmark + RL environment for agents that design neural-network architectures. Models emit graph edits; a deterministic verifier grades the result. No GPU, no human judge.
To associate your repository with the verifiable-reward topic, visit your repo's landing page and select "manage topics."