GhostCacher is a distributed Key-Value (KV) prompt caching orchestrator that dramatically reduces LLM inference latency and cost by storing and reusing the computed attention states of frequently used prompt prefixes across a distributed GPU cluster.
python rust redis dockerfile typescript key-value prometheus grafana-dashboard orchestrator prompt-caching llm-inference inference-latency prompt-prefixes distributed-key-value-caching
-
Updated
Apr 30, 2026 - Rust