Every attempt — successful, failed, or statically impossible — to run GLM-5.3-Flash on NVIDIA CMP 170HX (GA100, SM80, 64 GiB HBM2e) cards.
One repository per model family/workload: all quantizations, runtimes, and attempt outcomes live here. See AGENTS.md for the publication boundary and evidence rules. DGX Spark deployment is documented separately in PixelML/GLM-5.3-Flash-NVFP4-Dual-DGX-Spark.
Four-card CMP 170HX test node: 4 x 64 GiB = 256 GiB aggregate VRAM, SM80 (Ampere), PCIe Gen2 x4 per card in the current test guest, forced airflow. Per-card power limits were not queried by the measured harness this phase; see the release manifest telemetry note. The node exposed three cards when the earliest attempts below were evaluated (2026-08-30); those records keep their three-card arithmetic as history. Generic labels only; see the club repository for full node documentation.
| # | Checkpoint (revision) | Exact bytes | Quantization | Runtime | SM80 support | Static fit (per card, current 4-card topology) | Execution status | Blocker | Evidence |
|---|---|---|---|---|---|---|---|---|---|
| 1 | LibertAIDAI/GLM-5.3-Flash-NVFP4 @ 11d73216cd636238e82e1d77fe1042ffab36e7fa |
~194.4 GiB download (community-reported) | NVFP4 (ModelOpt, marlin MoE) | vLLM fork (SM121 image) | No — sparse-MLA path targets SM12x FlashInfer backends (measured: registry + PR review) | ~48.6 GiB/card at TP=4 — fits on paper (inferred, community-reported size); three-card era fit was ~64.8 GiB/card at TP=3, over budget | Failed (compatibility-only, 2026-08-30) | glm5_next absent from upstream vLLM registry; support PR vllm#53906 open, SM90+ only |
attempt record |
| 2 | Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw @ 25a44fdbf16862a46b7cc9921142c6c81350af2f |
175,642,157,752 B = 175.64 GB = 163.58 GiB total (measured, HF blob sum at pinned rev) | EXL3/TR3 4 bpw (uniform K4, routed experts) | ExLlamaV3 via SM121 fork image | Untested on SM80 — quant gate allows Ampere+ (inferred from get_min_capability()), but kernels ship as sm_121a cubins |
~40.9 GiB/card at TP=4 with ~23 GiB/card headroom (inferred; expert-layout skew unverified); three-card era fit was ~54.5 GiB/card at TP=3 | Failed (compatibility-only, 2026-08-30) | SM121-only binary distribution; no SM80 build exists or is documented; NoPE-MLA attention needs a non-sparse SM80 backend that no runtime provides | attempt record |
| 3 | cyankiwi/GLM-5.3-Flash-AWQ-INT4 @ 3999f9bf2c3e3064790af5a2d12d19090fc97f4d |
212,721,952,636 B = 198.1 GiB (measured, HF API blob sizes) | AWQ INT4 (pack-quantized, group 32, symmetric=false) | upstream vLLM (if glm5_next were supported) |
No — same runtime blocker as row 1 | ~49.5 GiB/card at TP=4 — fits on paper (measured blob sizes); three-card era fit was 66.0 GiB/card at TP=3, 2.04 GiB over budget | Failed on the three-card node (static fit, 2026-08-30); fit arithmetic cleared by the 4-card topology | Runtime blocker unchanged: glm5_next absent from upstream vLLM |
attempt record |
| 4 | zai-org/GLM-5.3-Flash (official FP8) and BF16 | >= ~328 GB FP8; BF16 larger (community-reported) | FP8 / BF16 | any | n/a — memory | No fit — ~76+ GiB/card at TP=4 (FP8; BF16 larger) | Failed (static fit) | Size alone, at either topology | attempt record |
| 5 | Future: true 3-bit / W4A16-with-exclusions <= ~55 GiB per card | — | AWQ/GPTQ W4A16-class | upstream vLLM, pending glm5_next support |
Untested — depends on upstream runtime work | Would fit at TP=4 (inferred); the AWQ row above already fits on paper at 4 cards | Not attempted | Runtime blocker, not size, on the 4-card node | attempt record |
| 6 | Intel/GLM-5.3-Flash-W4A16-AutoRound @ 5eee1846f0321058ed73745f9aa16f2aaf0fc0a0 |
181,505,393,058 B = 169.04 GiB (measured, HF API blob sizes, 34 shards + extra-tensor files) | W4A16, AutoRound, 4-bit | upstream vLLM (if glm5_next were supported) |
No — same runtime blocker as rows 1 and 3; registry re-checked 2026-09-02, still zero glm5_next entries; PR #53906 still open, SM90+ only |
~42.3 GiB/card at TP=4 — fits on paper (measured blob sizes), ~21.7 GiB/card headroom | Not attempted, registry-blocked before download (2026-09-02) | glm5_next absent from upstream vLLM registry, on any GPU architecture; no SM80 path exists or is planned |
attempt record |
| 7 | turboderp/GLM-5.3-Flash-exl3 branch 4.05bpw @ 2a30229e67012798ba9f0cd832bb78abf4c363d5 |
165,151,555,504 B = 153.81 GiB (measured, download manifest, 31 manifest files) | EXL3, 4.05 bpw | exllamav3 1.4.6+cu128.torch2.10.0 via TabbyAPI | Yes — measured. First row in this table to reach sustained, multi-request serving on CMP 170HX itself (a separate GGUF/llama.cpp pairing served earlier off-table; see "Current conclusion" below) | Fits: 153.81 GiB weights across 4x64 GiB = 256 GiB, at manual gpu_split: [48, 48, 48, 48] GiB/card (measured; tensor_parallel raises NotImplementedError on this architecture, so manual gpu_split is required) |
Success — server booted, all gates passed, full benchmark ladder run at the verified 180 W power cap (canonical): 25.2-44.6 tok/s aggregate, C1-C8 (250 W accidental-default run, 2026-09-02: 26.9-44.8 tok/s, retained for comparison — no consistent difference outside noise); long-context follow-up (2026-09-03): validated 262,144-token context (250,000 prompt tokens tested) with cache_mode: FP16 |
None. The original ~2,048-token cap under cache_mode: Q8 was root-caused (2026-09-03) to DSA sparse attention asserting against a quantized MLA cache, and resolved by switching to cache_mode: FP16 — now the recommended default for long-context use; throughput ladder at 262k context not yet re-run |
attempt record, results |
Quantization legend: NVFP4 = NVIDIA 4-bit floating point; EXL3/TR3 = TurboDerp 3-bit-class quant with group-wise codebooks; AWQ = activation-aware weight quantization. "Static fit" = weights-only budget per card at the node's current 4-card topology (TP=4), before CUDA context (~0.5-1 GiB), activations, and KV cache; rows retain their three-card-era TP=3 arithmetic as history where it differs.
The two blockers that killed every early attempt are resolved on paper as of 2026-08-30:
- An SM80 runtime exists: the unslothai/llama.cpp
DSA fork builds for
sm_80(measured CMake configure; see the GGUF attempt). vLLM paths remain blocked:glm5_nextis absent from upstream vLLM's model registry (measured 2026-08-30); the only known support PR (vllm#53906) is open, unmerged, and targets SM90+. - Fitting checkpoints exist: the UD-IQ4_XS GGUF (146.05 GiB measured) fits 4 x 64 GiB with ~27.5 GiB/card margin at an even split, and the EXL3/TR3 4 bpw (163.58 GiB measured) and AWQ INT4 (198.1 GiB measured) also fit on paper at TP=4.
Measured validation (Phase C, 2026-08-30, four-card node): the GGUF pairing served successfully. The UD-IQ4_XS GGUF on the sm_80 llama.cpp fork reached 17.73 tok/s median single-stream (5 reps), 17.71 tok/s aggregate at c=4 (two-run median), with a 41-repetition soak completing cleanly. The corrected 26-task evaluation scored 21/26 overall — math 8/8, instruction 4/5, long-context 3/3 on the tuning set, plus 4/4 held-out math and 1/1 held-out code — after two harness defects were found and fixed during the phase (a coding sandbox that dropped the candidate module, and completion budgets that starved reasoning output before any answer bytes). The two remaining coding misses and two held-out instruction misses are recorded as genuine failures at the evaluated budget; the committed receipts lack the finish/reasoning fields needed to support per-miss diagnosis, so none is claimed. Full receipt, per-request data, charts, and rebench entry: results/summary.csv, results/phase63/README.md, results/phase63/charts/, bench/rebench.sh. The EXL3/TR3, AWQ, NVFP4, and FP8/BF16 rows above remain blocked as recorded; all failed attempts stay in this repository as history.
Row 7 above (EXL3 4.05bpw on exllamav3 1.4.6 via TabbyAPI) is the first
comparison-table row to reach sustained, multi-request serving through
the full benchmark ladder (25.2-44.6 tok/s aggregate across concurrency
1-8 at the club's verified 180 W power cap; 20/20 golden-corpus pass). It
reuses the same node-level findings that unblocked the GGUF pairing above
(a fitting quant, an SM80 runtime for glm5_next), on a different runtime
and quant format. Full record, boot-attempt failure ladder, and the
load-bearing Q8-cache context-length limitation:
attempts/exl3-4.05bpw-exllamav3/README.md.
Power cap correction (2026-09-03). The first ladder run (2026-09-02, 26.9-44.8 tok/s) used the vBIOS default 250 W by accident — no per-card cap had been set that session. A same-server, no-restart re-measure at the verified 180 W club-standard cap found no consistent throughput difference outside run-to-run noise at any concurrency level. 180 W is now the canonical cap; the 250 W numbers are retained, labeled, for comparison. Full comparison table: attempts/exl3-4.05bpw-exllamav3/README.md.
vLLM and SGLang serving lanes remain terminal / out of scope on this node, and were not re-run for this update:
- vLLM: terminal.
glm5_next(this model's architecture) is absent from upstream vLLM's model registry on any GPU architecture (rows 1, 3, and 6 above; re-checked 2026-09-02); the only known support PR (vllm#53906) is open, unmerged, and targets SM90+ only. See attempts/w4a16-autoround-vllm/README.md for the fullest writeup of this blocker. - SGLang: no attempt record exists for SGLang in this repository, and no prior documented reason to rule it out was found during this update. This is not a claim that SGLang is terminal — it is simply undocumented here. TODO: a future attempt record should evaluate SGLang against this architecture on SM80 before this repository can state a status for it.
Same-day follow-up (2026-09-03): the ~2,048-token context cap on row 7
is resolved. A root-cause investigation found that the cap came from
cache_mode: Q8 colliding with the model's DSA sparse-attention indexer,
not from max_seq_len or TabbyAPI settings. Switching to cache_mode: FP16 (with max_seq_len/cache_size: 262144, chunk_size: 4096,
gpu_split unchanged) validated context up to 262,144 tokens (250,000
prompt tokens actually tested), with passing needle-in-haystack retrieval
at 32k and 250k tokens and no OOM, crash, or Xid/ECC events anywhere in
the ladder. cache_mode: FP16 is now the recommended default for this
recipe whenever long context matters; the Q8/32k config remains
documented as a lower-VRAM short-context alternative. See
attempts/exl3-4.05bpw-exllamav3/README.md.
Apache-2.0 for this repository's own content. Checkpoints and drafts referenced here carry their own licenses (base model, EXL3 source-available, DFlash2 CC BY-NC-ND); record and respect them per attempt.