-
Notifications
You must be signed in to change notification settings - Fork 86
Pull requests: RL-Align/RL-Kernel
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
feat(ascend): Qwen-Image qk_rmsnorm & multi_axis_rope kernels (issue #386)
#411
opened Sep 13, 2026 by
erfgss
Contributor
Loading…
4 of 7 tasks
Qwen image latent pack unpack
multimodal
Features, bugs, or optimizations specific to multimodal support.
#410
opened Sep 12, 2026 by
nodeeeeee
Loading…
style: format Python sources and benchmark results
#408
opened Sep 12, 2026 by
maxiaosong1124
Collaborator
Loading…
Add fixed-order MoE shared/residual merge reference
deepseek-P6
DSv4
platform: cuda
Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)
#404
opened Sep 11, 2026 by
AsterWang
Loading…
[DSv4][P5-4] MXFP8×MXFP4 Routed Expert grouped GEMM: strict CUDA kernel
deepseek-P5
DSv4
platform: cuda
Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)
#399
opened Sep 10, 2026 by
Bignonia7
Loading…
feat(musa): add native deterministic gemm kernel
MUSA
#395
opened Sep 9, 2026 by
Arlo-mt
Collaborator
Loading…
[CI] Fix black formatting on main and pin line-length in pyproject.toml
#384
opened Sep 3, 2026 by
Dnoob
Contributor
Loading…
[DSv4][P1-S0] Start kit for the P1 work package (mHC + RMSNorm)
deepseek-P1
DSv4
platform: cuda
Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)
#383
opened Sep 3, 2026 by
zhangj1an
Collaborator
Loading…
10 tasks
[DSv4][P5-0] Start kit for the P5 work package (MXFP4 Routed Expert + LoRA + Shared Expert)
deepseek-P5
DSv4
platform: cuda
Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)
#368
opened Sep 1, 2026 by
KJLdefeated
Collaborator
Loading…
[WS1] Add batch-invariant h_aggregate kernel
deepseek-P1
DSv4
platform: cuda
Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)
#366
opened Aug 30, 2026 by
nodeeeeee
Loading…
docs: add DCO 1.1 text and contributor sign-off guide
type: ci-cd
Modify GitHub Actions, automated tests, and packaging/deployment tasks.
#359
opened Aug 29, 2026 by
Zhifu-Liu
Contributor
Loading…
[WS2][Logp] Deterministic config option for operator
platform: cuda
Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)
#314
opened Aug 16, 2026 by
KJLdefeated
Collaborator
Loading…
[CI][refator]: migrate GPU workflow
needs-gpu-ci
#309
opened Aug 14, 2026 by
Flink-ddd
Collaborator
Loading…
[WS2][PR1][Logp] Add TP-aware logprob contract and dispatch metadata
#259
opened Aug 2, 2026 by
ryankert01
Contributor
Loading…
[FEAT][kernels] Add fused ratio-clip-aggregate loss primitive
needs-gpu-ci
#255
opened Jul 31, 2026 by
Chen-BUPT
Loading…
feat(observability): NVTX kernel tracing + Prometheus metrics endpoint
#250
opened Jul 27, 2026 by
Billy1900
Loading…
3 of 4 tasks
feat(observability): add NVTX tracing and Prometheus metrics
#242
opened Jul 22, 2026 by
BruceLoveDecimal
•
Draft
feat(kernels): SM90 TMA + mma.sync FlashAttention, forward + backward (causal, varlen, LSE)
needs-gpu-ci
stale
#237
opened Jul 20, 2026 by
Billy1900
Loading…
5 tasks done
[FEAT][kernels]: add backward support for fused logp CUDA kernels
needs-gpu-ci
#234
opened Jul 19, 2026 by
Dnoob
Contributor
Loading…
feat(kernels): Triton FlashAttention LSE export + varlen packing
needs-gpu-ci
stale
#233
opened Jul 19, 2026 by
Billy1900
Loading…
3 tasks done
docs(bench): add H100 SXM5 benchmark results alongside A100 baseline
stale
#231
opened Jul 19, 2026 by
Billy1900
Loading…
4 of 5 tasks
feat(executors): add colocated dual-engine with vLLM sleep/wake
stale
#218
opened Jul 11, 2026 by
icenfly
Loading…
Previous Next
ProTip!
What’s not been updated in a month: updated:<2026-08-13.