Skip to content

About

GLM-5.3-Flash on CMP 170HX (SM80): every attempt - NVFP4, AWQ, GPTQ, EXL3, FP8/BF16 - with static-fit math, runtime compatibility, and reproducible evidence

Resources

Stars

3 stars

Watchers

0 watching

Forks

Latest commit

 

History

26 Commits

Folders and files

Repository files navigation

GLM-5.3-Flash on CMP 170HX

Every attempt — successful, failed, or statically impossible — to run GLM-5.3-Flash on NVIDIA CMP 170HX (GA100, SM80, 64 GiB HBM2e) cards.

One repository per model family/workload: all quantizations, runtimes, and attempt outcomes live here. See AGENTS.md for the publication boundary and evidence rules. DGX Spark deployment is documented separately in PixelML/GLM-5.3-Flash-NVFP4-Dual-DGX-Spark.

Hardware target

Four-card CMP 170HX test node: 4 x 64 GiB = 256 GiB aggregate VRAM, SM80 (Ampere), PCIe Gen2 x4 per card in the current test guest, forced airflow. Per-card power limits were not queried by the measured harness this phase; see the release manifest telemetry note. The node exposed three cards when the earliest attempts below were evaluated (2026-08-30); those records keep their three-card arithmetic as history. Generic labels only; see the club repository for full node documentation.

Comparison table

# Checkpoint (revision) Exact bytes Quantization Runtime SM80 support Static fit (per card, current 4-card topology) Execution status Blocker Evidence
1 LibertAIDAI/GLM-5.3-Flash-NVFP4 @ 11d73216cd636238e82e1d77fe1042ffab36e7fa ~194.4 GiB download (community-reported) NVFP4 (ModelOpt, marlin MoE) vLLM fork (SM121 image) No — sparse-MLA path targets SM12x FlashInfer backends (measured: registry + PR review) ~48.6 GiB/card at TP=4 — fits on paper (inferred, community-reported size); three-card era fit was ~64.8 GiB/card at TP=3, over budget Failed (compatibility-only, 2026-08-30) glm5_next absent from upstream vLLM registry; support PR vllm#53906 open, SM90+ only attempt record
2 Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw @ 25a44fdbf16862a46b7cc9921142c6c81350af2f 175,642,157,752 B = 175.64 GB = 163.58 GiB total (measured, HF blob sum at pinned rev) EXL3/TR3 4 bpw (uniform K4, routed experts) ExLlamaV3 via SM121 fork image Untested on SM80 — quant gate allows Ampere+ (inferred from get_min_capability()), but kernels ship as sm_121a cubins ~40.9 GiB/card at TP=4 with ~23 GiB/card headroom (inferred; expert-layout skew unverified); three-card era fit was ~54.5 GiB/card at TP=3 Failed (compatibility-only, 2026-08-30) SM121-only binary distribution; no SM80 build exists or is documented; NoPE-MLA attention needs a non-sparse SM80 backend that no runtime provides attempt record
3 cyankiwi/GLM-5.3-Flash-AWQ-INT4 @ 3999f9bf2c3e3064790af5a2d12d19090fc97f4d 212,721,952,636 B = 198.1 GiB (measured, HF API blob sizes) AWQ INT4 (pack-quantized, group 32, symmetric=false) upstream vLLM (if glm5_next were supported) No — same runtime blocker as row 1 ~49.5 GiB/card at TP=4 — fits on paper (measured blob sizes); three-card era fit was 66.0 GiB/card at TP=3, 2.04 GiB over budget Failed on the three-card node (static fit, 2026-08-30); fit arithmetic cleared by the 4-card topology Runtime blocker unchanged: glm5_next absent from upstream vLLM attempt record
4 zai-org/GLM-5.3-Flash (official FP8) and BF16 >= ~328 GB FP8; BF16 larger (community-reported) FP8 / BF16 any n/a — memory No fit — ~76+ GiB/card at TP=4 (FP8; BF16 larger) Failed (static fit) Size alone, at either topology attempt record
5 Future: true 3-bit / W4A16-with-exclusions <= ~55 GiB per card — AWQ/GPTQ W4A16-class upstream vLLM, pending glm5_next support Untested — depends on upstream runtime work Would fit at TP=4 (inferred); the AWQ row above already fits on paper at 4 cards Not attempted Runtime blocker, not size, on the 4-card node attempt record
6 Intel/GLM-5.3-Flash-W4A16-AutoRound @ 5eee1846f0321058ed73745f9aa16f2aaf0fc0a0 181,505,393,058 B = 169.04 GiB (measured, HF API blob sizes, 34 shards + extra-tensor files) W4A16, AutoRound, 4-bit upstream vLLM (if glm5_next were supported) No — same runtime blocker as rows 1 and 3; registry re-checked 2026-09-02, still zero glm5_next entries; PR #53906 still open, SM90+ only ~42.3 GiB/card at TP=4 — fits on paper (measured blob sizes), ~21.7 GiB/card headroom Not attempted, registry-blocked before download (2026-09-02) glm5_next absent from upstream vLLM registry, on any GPU architecture; no SM80 path exists or is planned attempt record
7 turboderp/GLM-5.3-Flash-exl3 branch 4.05bpw @ 2a30229e67012798ba9f0cd832bb78abf4c363d5 165,151,555,504 B = 153.81 GiB (measured, download manifest, 31 manifest files) EXL3, 4.05 bpw exllamav3 1.4.6+cu128.torch2.10.0 via TabbyAPI Yes — measured. First row in this table to reach sustained, multi-request serving on CMP 170HX itself (a separate GGUF/llama.cpp pairing served earlier off-table; see "Current conclusion" below) Fits: 153.81 GiB weights across 4x64 GiB = 256 GiB, at manual gpu_split: [48, 48, 48, 48] GiB/card (measured; tensor_parallel raises NotImplementedError on this architecture, so manual gpu_split is required) Success — server booted, all gates passed, full benchmark ladder run at the verified 180 W power cap (canonical): 25.2-44.6 tok/s aggregate, C1-C8 (250 W accidental-default run, 2026-09-02: 26.9-44.8 tok/s, retained for comparison — no consistent difference outside noise); long-context follow-up (2026-09-03): validated 262,144-token context (250,000 prompt tokens tested) with cache_mode: FP16 None. The original ~2,048-token cap under cache_mode: Q8 was root-caused (2026-09-03) to DSA sparse attention asserting against a quantized MLA cache, and resolved by switching to cache_mode: FP16 — now the recommended default for long-context use; throughput ladder at 262k context not yet re-run attempt record, results

Quantization legend: NVFP4 = NVIDIA 4-bit floating point; EXL3/TR3 = TurboDerp 3-bit-class quant with group-wise codebooks; AWQ = activation-aware weight quantization. "Static fit" = weights-only budget per card at the node's current 4-card topology (TP=4), before CUDA context (~0.5-1 GiB), activations, and KV cache; rows retain their three-card-era TP=3 arithmetic as history where it differs.

Current conclusion (2026-08-30, four-card node)

The two blockers that killed every early attempt are resolved on paper as of 2026-08-30:

  1. An SM80 runtime exists: the unslothai/llama.cpp DSA fork builds for sm_80 (measured CMake configure; see the GGUF attempt). vLLM paths remain blocked: glm5_next is absent from upstream vLLM's model registry (measured 2026-08-30); the only known support PR (vllm#53906) is open, unmerged, and targets SM90+.
  2. Fitting checkpoints exist: the UD-IQ4_XS GGUF (146.05 GiB measured) fits 4 x 64 GiB with ~27.5 GiB/card margin at an even split, and the EXL3/TR3 4 bpw (163.58 GiB measured) and AWQ INT4 (198.1 GiB measured) also fit on paper at TP=4.

Measured validation (Phase C, 2026-08-30, four-card node): the GGUF pairing served successfully. The UD-IQ4_XS GGUF on the sm_80 llama.cpp fork reached 17.73 tok/s median single-stream (5 reps), 17.71 tok/s aggregate at c=4 (two-run median), with a 41-repetition soak completing cleanly. The corrected 26-task evaluation scored 21/26 overall — math 8/8, instruction 4/5, long-context 3/3 on the tuning set, plus 4/4 held-out math and 1/1 held-out code — after two harness defects were found and fixed during the phase (a coding sandbox that dropped the candidate module, and completion budgets that starved reasoning output before any answer bytes). The two remaining coding misses and two held-out instruction misses are recorded as genuine failures at the evaluated budget; the committed receipts lack the finish/reasoning fields needed to support per-miss diagnosis, so none is claimed. Full receipt, per-request data, charts, and rebench entry: results/summary.csv, results/phase63/README.md, results/phase63/charts/, bench/rebench.sh. The EXL3/TR3, AWQ, NVFP4, and FP8/BF16 rows above remain blocked as recorded; all failed attempts stay in this repository as history.

Update (2026-09-03): first successful row in the comparison table

Row 7 above (EXL3 4.05bpw on exllamav3 1.4.6 via TabbyAPI) is the first comparison-table row to reach sustained, multi-request serving through the full benchmark ladder (25.2-44.6 tok/s aggregate across concurrency 1-8 at the club's verified 180 W power cap; 20/20 golden-corpus pass). It reuses the same node-level findings that unblocked the GGUF pairing above (a fitting quant, an SM80 runtime for glm5_next), on a different runtime and quant format. Full record, boot-attempt failure ladder, and the load-bearing Q8-cache context-length limitation: attempts/exl3-4.05bpw-exllamav3/README.md.

Power cap correction (2026-09-03). The first ladder run (2026-09-02, 26.9-44.8 tok/s) used the vBIOS default 250 W by accident — no per-card cap had been set that session. A same-server, no-restart re-measure at the verified 180 W club-standard cap found no consistent throughput difference outside run-to-run noise at any concurrency level. 180 W is now the canonical cap; the 250 W numbers are retained, labeled, for comparison. Full comparison table: attempts/exl3-4.05bpw-exllamav3/README.md.

vLLM and SGLang serving lanes remain terminal / out of scope on this node, and were not re-run for this update:

  • vLLM: terminal. glm5_next (this model's architecture) is absent from upstream vLLM's model registry on any GPU architecture (rows 1, 3, and 6 above; re-checked 2026-09-02); the only known support PR (vllm#53906) is open, unmerged, and targets SM90+ only. See attempts/w4a16-autoround-vllm/README.md for the fullest writeup of this blocker.
  • SGLang: no attempt record exists for SGLang in this repository, and no prior documented reason to rule it out was found during this update. This is not a claim that SGLang is terminal — it is simply undocumented here. TODO: a future attempt record should evaluate SGLang against this architecture on SM80 before this repository can state a status for it.

Same-day follow-up (2026-09-03): the ~2,048-token context cap on row 7 is resolved. A root-cause investigation found that the cap came from cache_mode: Q8 colliding with the model's DSA sparse-attention indexer, not from max_seq_len or TabbyAPI settings. Switching to cache_mode: FP16 (with max_seq_len/cache_size: 262144, chunk_size: 4096, gpu_split unchanged) validated context up to 262,144 tokens (250,000 prompt tokens actually tested), with passing needle-in-haystack retrieval at 32k and 250k tokens and no OOM, crash, or Xid/ECC events anywhere in the ladder. cache_mode: FP16 is now the recommended default for this recipe whenever long context matters; the Q8/32k config remains documented as a lower-VRAM short-context alternative. See attempts/exl3-4.05bpw-exllamav3/README.md.

License

Apache-2.0 for this repository's own content. Checkpoints and drafts referenced here carry their own licenses (base model, EXL3 source-available, DFlash2 CC BY-NC-ND); record and respect them per attempt.

About

GLM-5.3-Flash on CMP 170HX (SM80): every attempt - NVFP4, AWQ, GPTQ, EXL3, FP8/BF16 - with static-fit math, runtime compatibility, and reproducible evidence

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages