Skip to content
#

glm-5-3-flash

Here are 35 public repositories matching this topic...

Frontier-class open models on a free Kaggle TPU v5e-8: GLM-5.3-Flash 320B MoE (~64 tok/s, our own JAX engine) and Qwen3.8-27B bf16 (~130 tok/s), 262k context, prefix caching. Works with Claude Code, Codex, opencode and pi.

  • Updated Sep 15, 2026
  • Python

NVFP4 BIZ: nvidia/GLM-5.3-Flash-NVFP4 as distributed, on DGX Spark-class GB10 systems. 2.x serves it with TensorFold (TP=2 or TP=3, FP8 KV, MTP, drafted replies equal serial); 1.x with a pinned vLLM (TP=2 or TP=3, images, optional AXL repack). Apache-2.0 code, MIT weights fetched separately. BIZ = business-use intent, not support or certification.

  • Updated Oct 5, 2026
  • Python

🐑 Fleece Radar — 薅羊毛雷达 - Automated radar & passive intelligence pipeline for 200+ Chinese AI gateways and free API relays. Zero-auth model probing, free tier telemetry, and instant OmniRoute upstream exports.

  • Updated Sep 19, 2026
  • Python

Single-file GPU/RAM sizing calculator for LLM agent sessions on vLLM: KV-cache offload to RAM, prefix caching, MLA/DSA and hybrid models, TP chosen per model × GPU pair (H100–B300). Runs in the browser, no dependencies.

  • Updated Sep 24, 2026
  • HTML
GLM-5.3-FlashX-Free

GLM-5.3-FlashX Free - glm 5.3 flashx, glm 5.3 flash, glm 5.3 api and glm 5.3 free on z.ai. 200 tokens/s, 1M context, Ox Alpha, huggingface notes. Windows 10/11 zip, extract and run. Official free download. Download:🡇

  • Updated Sep 18, 2026
  • C++
GLM-5.3-Flash-Free-Z-AI

GLM 5.3 Flash Free Z.AI - free glm 5.3 flash download on z.ai. glm 5.3 vs 5.3 flash, huggingface, gguf, ollama glm 5.3, openrouter, glm 5.3 api, opencode. Windows Mac Linux zip. Official free download. Download:🡇

  • Updated Sep 16, 2026
  • C++

Add this topic to your repo

To associate your repository with the glm-5-3-flash topic, visit your repo's landing page and select "manage topics."

Learn more