- π Undergraduate student at Beijing University of Posts and Telecommunications (BUPT), School of Computer Science
- π¬ Research interests: RLVR Β· RLHF Β· Optimization Algorithms
- π± Currently exploring the intersection of reinforcement learning and large language model alignment
- π Beijing, China
| Area | Description |
|---|---|
| RLVR | Reinforcement Learning from Verifiable Rewards β scalable reward signals beyond human feedback |
| RLHF | Reinforcement Learning from Human Feedback β aligning LLMs with human preferences |
| Optimizer | Adaptive optimization methods (AdamW, Muon, Shampoo, etc.) for deep learning |
APO_OFFICAL β [ICML 2026] The official repository for Anchored Policy Optimization: Mitigating Exploration Collapse via Support-Constrained Rectification
β 16 π΄ 2
SPPO β [ACL 2026 Oral] SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks official repos.
β 3 π΄ 3
No recent public activity.
Blog RSS not configured or no posts found. Set BLOG_RSS_URL to enable.


