Skip to content
#

ninfer

Here are 16 public repositories matching this topic...

Wire the 1CatAI Split-D D256 FlashAttention kernel (fishlikeX/sm70-attn, MIT) into NInfer on Tesla V100 sm_70: +36-41% prefill, TTFT -3min, decode unchanged. Measured data + integration guide. Published by the user with AI assistance.

  • Updated Sep 29, 2026
  • Cuda

Tesla V100 (sm_70) LLM inference on one card: sm70 decode kernel port + KV context-cache tuning for long-context agents, measured on a real 53-request Qwen3.8-27B agent session. 单卡 Tesla V100 32GB 跑 NInfer + Qwen3.8-27B:sm70 解码内核移植 + 上下文缓存调参,附 53 个真实 agent 请求的满载实测与全部原始日志(中文为主,含英文版)。Posted by an AI on behalf of the machine's owner.

  • Updated Sep 29, 2026
  • Shell

Add this topic to your repo

To associate your repository with the ninfer topic, visit your repo's landing page and select "manage topics."

Learn more