Skip to content
View 1BIMU's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report 1BIMU

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
1BIMU/README.md

Typing SVG

Profile Views

πŸ§‘β€πŸ’» About Me

  • πŸŽ“ Undergraduate student at Beijing University of Posts and Telecommunications (BUPT), School of Computer Science
  • πŸ”¬ Research interests: RLVR Β· RLHF Β· Optimization Algorithms
  • 🌱 Currently exploring the intersection of reinforcement learning and large language model alignment
  • πŸ“ Beijing, China

πŸ”­ Research Interests

Area Description
RLVR Reinforcement Learning from Verifiable Rewards β€” scalable reward signals beyond human feedback
RLHF Reinforcement Learning from Human Feedback β€” aligning LLMs with human preferences
Optimizer Adaptive optimization methods (AdamW, Muon, Shampoo, etc.) for deep learning

πŸ“Œ Pinned Repositories

APO_OFFICAL β€” [ICML 2026] The official repository for Anchored Policy Optimization: Mitigating Exploration Collapse via Support-Constrained Rectification Python ⭐ 16 🍴 2

SPPO β€” [ACL 2026 Oral] SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks official repos. Python ⭐ 3 🍴 3


⚑ Recent Activity

No recent public activity.


πŸ“ Latest Blog Posts

Blog RSS not configured or no posts found. Set BLOG_RSS_URL to enable.


πŸ› οΈ Tech Stack

Python PyTorch C++ Linux Git LaTeX


πŸ“Š GitHub Stats

GitHub Streak


πŸ“« Contact

GitHub Email Zhihu


"The pursuit of intelligence β€” from theory to practice." Β· Last updated: auto-refreshed every 3 hours

Pinned Loading

  1. APO_OFFICAL APO_OFFICAL Public

    [ICML 2026] The official repository for Anchored Policy Optimization: Mitigating Exploration Collapse via Support-Constrained Rectification

    Python 16 2

  2. SPPO SPPO Public

    [ACL 2026 Oral] SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks official repos.

    Python 3 3