forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 0
Pull requests: ddvnguyen/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
arm: adaptive-KV-streaming correctness/perf on the 2x-RTX RPC-split topology (design phase)
#117
opened Sep 10, 2026 by
hydra-z
Bot
Loading…
Arm spec: context-shift hybrid-state correctness for Qwen3.8 (design phase)
#116
opened Sep 10, 2026 by
ddvnguyen
Owner
Loading…
feat(cuda): PR103.0 GDN cache-cpy fusion — A/B toggle + test coverage (implements #112)
#115
opened Sep 10, 2026 by
ddvnguyen
Owner
Loading…
3 of 5 tasks
arm PR106.x: EXL3 DFlash2 kit reference series (PR106.0 external / PR106.1 native 3.5bpw single-GPU)
#114
opened Sep 9, 2026 by
ddvnguyen
Owner
Loading…
3 tasks
arm PR104.x: mixed-quant row sharding arm series (PR104.0 roofline / PR104.1 spike / PR104.2 production)
#113
opened Sep 9, 2026 by
ddvnguyen
Owner
Loading…
4 tasks
arm PR103.0: CUDA GDN cache-cpy fusion (arm spec + implementation PR spec)
#112
opened Sep 9, 2026 by
ddvnguyen
Owner
Loading…
5 tasks
baseline arm PR105.0: single-machine dual-GPU (no RPC), PR #110 config reference
#111
opened Sep 9, 2026 by
ddvnguyen
Owner
Loading…
7 tasks
fix(713): EAGAIN retry helper + slot quarantine on STATE_PUT restore failure
#106
opened Aug 29, 2026 by
ddvnguyen
Owner
Loading…
epic(610): extract Hydra server logic out of server-context.cpp (mechanical, A/B toggle)
#105
opened Aug 21, 2026 by
ddvnguyen
Owner
Loading…
3 tasks
ci: hydra-engine PR validation (build+ctest, server pytest, C# parity)
#81
opened Aug 6, 2026 by
ddvnguyen
Owner
Loading…
3 tasks done
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.