Wire the 1CatAI Split-D D256 FlashAttention kernel (fishlikeX/sm70-attn, MIT) into NInfer on Tesla V100 sm_70: +36-41% prefill, TTFT -3min, decode unchanged. Measured data + integration guide. Published by the user with AI assistance.
cuda volta prefill v100 inference-optimization kv-cache long-context llama-cpp llm-inference flash-attention d256 tesla-v100 qwen3 gpu-kernel sm70 attention-kernel 27b ninfer split-d splitkv
-
Updated
Sep 29, 2026 - Cuda