Popular repositories Loading
-
-
llama-moe-cache
llama-moe-cache PublicForked from ongunm/llama-moe-cache
Expert cache + predictive prefetch for MoE inference in llama.cpp. A 12GB GPU can run a 120GB model at native speed. At 3% sparsity, a single workstation could theoretically run a 1T parameter mode…
C++
-
-
-
mcp-plan
mcp-plan PublicForked from haras-unicorn/mcp-plan
MCP server that provides planning tooling.
Nix
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.

