Research preview for reproducible MoE routing and memory oversubscription on consumer GPUs. Not a production inference engine.
-
Updated
Aug 26, 2026 - Python
Research preview for reproducible MoE routing and memory oversubscription on consumer GPUs. Not a production inference engine.
DualDeadline adds separate gate/up and down-projection transfer deadlines to exact MoE offloading, with H200 validation and Triton-optimized predictors.
Fits BF16 models too large for your GPU's VRAM by streaming losslessly compressed weights layer by layer, with speculative decoding, and an experimental benchmark suite showing it beats AirLLM and Hugging Face Accelerate at matched memory.
Research artifact for Memory-Sovereign Inference: Output-Exact Execution Beyond Full Residency
Quality-first local MiniMax H3 video-series generation and loopback API for dual RTX 4090 workstations, with native audio, references, P8/P9 continuity, and preserved artifacts.
Add a description, image, and links to the model-offloading topic page so that developers can more easily learn about it.
To associate your repository with the model-offloading topic, visit your repo's landing page and select "manage topics."