Skip to content
@moonmath-ai

MoonMath.ai

Private AI inference, engineered for performance.
MoonMath.ai

We are a small team of mathematicians and engineers building faster LLM inference through low-level algorithms, systems engineering, custom kernels, and hardware-aware optimization — with a focus on coding models and agentic workloads.

Zro: Fast inference for coding agents

Zro is MoonMath's private inference platform built for coding agents. It gives developers fast access to leading open-weight coding models through a single endpoint, optimized for long-context, multi-turn agentic workloads.

  • Built for coding agents: Connect Claude Code, Codex, OpenCode, Cline, and more in minutes.
  • Accelerated inference: Our serving stack is engineered for high-throughput, low-latency LLM inference.
  • Private by default: Zero request retention and no training on customer data.
  • Open models, one endpoint: Run capable open-weight coding models without managing GPUs, inference engines, or serving infrastructure yourself.
  • Performance from the stack down: MoonMath combines algorithmic optimization, custom kernels, compression, and hardware-aware deployment to push more performance from modern accelerators.

Start coding with Zro

MoonLite: LLM acceleration research & tooling

MoonLite is MoonMath's collection of low-level techniques and tools for making large models faster and more efficient:

  • LiteLinear: decomposed modules designed to replace standard FFN layers with more efficient alternatives.
  • BackLite: FlashAttention 3-based backward-pass acceleration using sparse gradient approximation.
  • LiteAttention: sparse attention techniques for accelerating large generative models.
  • LiteRunner: experiment infrastructure for benchmarking and developing model acceleration techniques, with local and Weights & Biases tracking.

Links

Pinned Loading

  1. LiteAttention LiteAttention Public

    Transforming Video Diffusion with Temporal Sparse Attention

    Python 56 5

  2. LiteLinear LiteLinear Public

    LiteLinear is a drop-in inference acceleration: compress nn.Linear layers via calibration-aware low-rank decomposition + quantization

    8 1

  3. LiteRunner LiteRunner Public

    MLOps-style tracking without touching the code.

    Python 3

  4. BackLite BackLite Public

    BackLite is a Hopper-optimized training kernel on top of FA3 that accelerates the backward pass of transformer attention layers

    Python 5

  5. HyperQuant HyperQuant Public

    Python 13 2

  6. amd-kernels amd-kernels Public

    Hand-tuned kernels for AMD Instinct

    C++ 5 2

Repositories

Showing 10 of 40 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…