-
Notifications
You must be signed in to change notification settings - Fork 238
Pull requests: lightseekorg/tokenspeed
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
perf(kimi3): fuse batched AttnRes graph on gfx950
#1089
opened Aug 13, 2026 by
qedawkins
Contributor
Loading…
perf(kernel): split gfx950 DSA decode across KV rows
#1088
opened Aug 13, 2026 by
Max191
Contributor
Loading…
feat(zmq): materialize precomputed multimodal inputs on msgpack ingest
#1081
opened Aug 13, 2026 by
slin1237
Contributor
Loading…
feat(zmq): piggyback a scheduler-load snapshot on slim output batches
#1079
opened Aug 13, 2026 by
slin1237
Contributor
Loading…
fix(layers): pass a 0-token batch through SiluAndMul without a launch
#1077
opened Aug 13, 2026 by
slin1237
Contributor
Loading…
perf(comm): route generic MNNVL AR patterns to upstream flashinfer
#1075
opened Aug 13, 2026 by
dongjiyingdjy
Contributor
•
Draft
[Stacked on #957] perf(kimi3): MNNVL CuTe-DSL finalize tail + K3 comm layer extraction
#1062
opened Aug 12, 2026 by
dongjiyingdjy
Contributor
•
Draft
[WIP] feat(kda): replay-based lazy commit for speculative target verify
#1058
opened Aug 11, 2026 by
nperrin-fr
Collaborator
•
Draft
refactor(pd): unify paged-cache contracts across models and transfer modes
#1057
opened Aug 11, 2026 by
chenht2022
Contributor
Loading…
feat(kimi-k3): serve DSpark drafts (fc_norm + AttnRes tap)
#1031
opened Aug 10, 2026 by
torchspec-bot
Collaborator
Loading…
refactor(amd): split gfx950 mxfp4 fused.py into fused/ subpackage
#1024
opened Aug 10, 2026 by
antiagainst
Member
•
Draft
fix(kimi-k3): a DSpark draft's MLA cache cannot diverge from the target's
#1015
opened Aug 9, 2026 by
torchspec-bot
Collaborator
Loading…
fix(pd): support DeepSeek V4 grouped layerwise cache handoff
#997
opened Aug 8, 2026 by
lucifer1004
Contributor
Loading…
feat(deepseek-v4): support Flash serving on SM120
#992
opened Aug 8, 2026 by
lucifer1004
Contributor
Loading…
[WIP] feat(gpt-oss-megakernel): fused decode megakernel for GPT-OSS-120B
#988
opened Aug 7, 2026 by
vmalepati1
•
Draft
feat(scheduler): LoRA adapter identity, max_loras batch cap, and per-adapter KV namespace
#976
opened Aug 7, 2026 by
towillwu
Contributor
•
1/4
Loading…
feat(runtime): add Nemotron architecture dispatch alias for PR1
#962
opened Aug 6, 2026 by
nathon-lee
Loading…
[WIP] perf(k3): latent moe multicast tail
#957
opened Aug 6, 2026 by
nperrin-fr
Collaborator
•
Draft
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.