Skip to content

Pull requests: FlashML-org/FreeToken

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

fix(gemma4-gguf): accept scalar attention.head_count_kv
#190 opened Aug 25, 2026 by qxZap Loading…
fix(deps): use PyTorch 2.13 to fix SM89 hangs
#184 opened Aug 25, 2026 by endenis Loading…
fix(engine): load prefill triton kernels at startup, not mid-request
#169 opened Aug 25, 2026 by jason-fxz Collaborator Loading…
GGUF: read multi-shard checkpoints
#154 opened Aug 24, 2026 by vcruz305 Loading…
feat(rocm): serve on AMD GPUs through the HIP toolchain
#137 opened Aug 24, 2026 by paralin Loading…
feat(rocm): add RDNA3 and RDNA4 runtime foundation
#132 opened Aug 24, 2026 by zihaomu Loading…
GGUF: support all quant types, add qwen35moe
#131 opened Aug 24, 2026 by vcruz305 Loading…
Add fp8 KV-cache quantization (--kv-dtype)
#126 opened Aug 24, 2026 by mkornreich Loading…
test(pinned): use mapped memory for UVA pointer check
#125 opened Aug 24, 2026 by huo-ju Loading…
ProTip! Follow long discussions with comments:>50.