Skip to content
#

1cat-vllm

Here are 3 public repositories matching this topic...

Language: All
Filter by language

Runbook + benchmarks: Qwen3.8-Flash-Next-ABLITERATED NVFP4 on 4× Tesla V100-32GB (reflashed SXM2→PCIe, 2+2 NVLink + PLX). 1Cat-vLLM 1.5.0, TP4 — 262,144-token context validated, 46 tok/s decode, 122 tok/s aggregate at 4 concurrent streams. Full E0–E17 optimization log with measured evidence.

  • Updated Sep 13, 2026
  • Python

Add this topic to your repo

To associate your repository with the 1cat-vllm topic, visit your repo's landing page and select "manage topics."

Learn more