You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
All benchmarks are batch size 1, single-stream decode, targeting local inference on consumer hardware. This is the llama.cpp/Ollama use case, not multi-tenant serving.
Hardware
Machine
GPU/Chip
Memory
Lucebox
NVIDIA RTX 3090
24GB VRAM
MacBook Pro
Apple M5 Max
36GB Unified
RTX 3090: pp520 tg128
Method
pp520 (tok/s)
tg128 (tok/s)
Megakernel
37,800
413
llama.cpp BF16
11,247
267
PyTorch HF
7,578
108
Speedups
vs llama.cpp
vs PyTorch
Decode (tg128)
1.55x
3.8x
Apple M5 Max
Method
tok/s
LM Studio (llama.cpp) BF16
229
Power Efficiency (DVFS)
Power Limit
Clock
Draw
tok/s
tok/J
vs Stock
420W (stock)
1980 MHz
314W
433
1.38
baseline
300W
1935 MHz
299W
432
1.44
99.8% speed, 5% less power
220W
1635 MHz
220W
411
1.87
95% speed, 30% less power
150W
405 MHz
150W
194
1.29
too aggressive
Sweet spot: 220W, 1.87 tok/J.
Methodology
Precision: BF16 weights and activations, FP32 accumulation. No quantization. All baselines (llama.cpp, PyTorch HF) also run BF16 for apples-to-apples comparison.
Power measurement: Accelerator power only via NVML energy counters (NVIDIA) and powermetrics (Apple Silicon), consistent with Hazy Research's Intelligence Per Watt methodology. Total system draw is higher for both platforms.
Correctness:bench_pp_tg.py includes an end-to-end correctness check, comparing megakernel output (prefill + decode) against a token-by-token reference decode path. Both must produce identical token sequences.
Warm-up: One warm-up run before timed measurements. Timing uses torch.cuda.synchronize() barriers with time.perf_counter().
llama.cpp version: Latest release at time of testing, BF16 mode, default settings.