Skip to content
View tonyd2wild's full-sized avatar

Sponsoring

@rationalsa

Block or report tonyd2wild

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. DeepSeek-v4-Flash-0731-DSpark-1M-NVFP4-KV-2x-DGX-Spark DeepSeek-v4-Flash-0731-DSpark-1M-NVFP4-KV-2x-DGX-Spark Public

    DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark

    Python 450 53

  2. GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark Public

    GLM-5.3-Flash (NVFP4) on 2x NVIDIA DGX Spark - vLLM TP2, 262K context, MTP. World-first deploy recipe: 7 day-0 bugs found and fixed, patched sm121 image, probes and full report.

    Python 97 9

  3. Deepseek-v4-Flash-TP2-DGX-Spark-500k-CTX Deepseek-v4-Flash-TP2-DGX-Spark-500k-CTX Public

    Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — Docker image, launch scripts, RDMA/NCCL setup, and the gotchas.

    Shell 81 12

  4. GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s Public

    Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster

    Python 76 9

  5. Qwen3.8-Flash-Next-NVFP4-DGX-Spark Qwen3.8-Flash-Next-NVFP4-DGX-Spark Public

    Day-0 deployment of Qwen3.8-Flash-Next (NVFP4) on 2x DGX Spark TP2 via SGLang. Includes the SM121 QSA kernel-guard fix, launcher, benchmarks, and full report.

    Shell 43 3

  6. MiniMax-M3-2x-DGX-Spark-36-tok-s MiniMax-M3-2x-DGX-Spark-36-tok-s Public

    MiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three serving lanes: speed / balanced / long-context.

    42 4