High-performance single-GPU inference for selected model checkpoints and GPUs.
-
Updated
Sep 28, 2026 - C++
High-performance single-GPU inference for selected model checkpoints and GPUs.
NInfer fork with end-user improvements such as model router and jev-alike decisions endpoint. Experimental repo, changes are subject to be wiped without notice based on my own usage observations. Packaged with nix.
Port NInfer, a single-GPU CUDA inference engine, to the NVIDIA L20 (Ada sm_89, 92 SMs, 48 GB): patch set, build tooling, and measured results
笔记本 RTX 5070 Ti Laptop 12G(12,227 MiB / 140 W TGP)· Bonsai-2-27B 三元量化 · MTP vs DFlash2 同上下文 A/B。Laptop GPU only — NOT the desktop 5070 Ti (16 GB / ~300 W); numbers are not comparable across the two.
Qwen3.8-27B on RTX 4090 D (48GB): production deployment of NInfer with MTP7 + E8 KV + NVMe disk cache, 195 tok/s decode, crash forensics for WDDM desktop GPUs
Rust CLI to run local LLMs through rootless Podman Compose
Measuring proxy + throughput dashboard for local LLM engines (llama.cpp, NInfer): live tok/s, cache-hit rate, TTFT, and history charts. Single-file panel, stdlib-only proxy, MIT.
Unofficial Windows distribution of iamwavecut/ninfer-all with target-specific builds and an integration contract for NInferEZ.
Wire the 1CatAI Split-D D256 FlashAttention kernel (fishlikeX/sm70-attn, MIT) into NInfer on Tesla V100 sm_70: +36-41% prefill, TTFT -3min, decode unchanged. Measured data + integration guide. Published by the user with AI assistance.
Tesla V100 (sm_70) LLM inference on one card: sm70 decode kernel port + KV context-cache tuning for long-context agents, measured on a real 53-request Qwen3.8-27B agent session. 单卡 Tesla V100 32GB 跑 NInfer + Qwen3.8-27B:sm70 解码内核移植 + 上下文缓存调参,附 53 个真实 agent 请求的满载实测与全部原始日志(中文为主,含英文版)。Posted by an AI on behalf of the machine's owner.
NInfer for Windows and around 16GB VRAM: RTX 5070 Ti / 5080 / 5090, Qwen3.8-27B GSQ-RCO Q3, CUDA 13 Native engine, tray manager, model conversion and measured setup guides.
To associate your repository with the ninfer topic, visit your repo's landing page and select "manage topics."