forked from xLLM-AI/xllm
-
Notifications
You must be signed in to change notification settings - Fork 0
Pull requests: Wang-1F/xllm
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
feat: Optimize kernels (activation, cache), add piecewise CUDA graph, and revert fused rope.
#63
opened Jan 21, 2026 by
yingxudeng
Loading…
perf: optimize data type conversions and activation kernel for improved performance.
#53
opened Jan 17, 2026 by
yingxudeng
Loading…
feat: add rec worker and layer for pure device pipeline[3/3].
#31
opened Jan 13, 2026 by
Wang-1F
Owner
Loading…
feat: add Cutlass support for Qwen3 W8A8 quantization.
#24
opened Jan 13, 2026 by
yingxudeng
•
Draft
ProTip!
Exclude everything labeled
bug with -label:bug.