Run models that do not fit in your RAM.
The always-read weights stay resident; the routed experts stream from disk, per token.
This file carries three things and nothing else: the progress bars, the document map, and tok/s for five fixed models with the machine that produced them named. Everything else lives one click away.
scripts/check-readme.shenforces it, and CI runs that.
Three platforms, one file each, from Releases.
| download | then | |
|---|---|---|
| Windows | Chaos-vX.Y.Z-windows-x86_64-Setup.exe |
double-click. Per-user, no admin rights. Opens the window |
| Linux | chaos_X.Y.Z_amd64.deb |
sudo apt install ./chaos_*.deb |
Chaos-vX.Y.Z-linux-x86_64.AppImage |
chmod +x and run — any distribution |
|
| macOS | Chaos-vX.Y.Z-macos-arm64.tar.gz |
sudo install -m 755 Chaos-*/chaos-* /usr/local/bin/ |
The window is Windows-only. Linux and macOS get every command-line tool, and
chaos-serve speaks both the OpenAI and the Anthropic APIs, so any client — an
editor, or Claude Code — works on all three.
No telemetry and no dependencies. Cargo.lock holds 22 packages and all 22 are
crates in this repository, so there is nothing third-party to audit. The one request
Chaos makes on its own is the window's update check against a static JSON file, which
CHAOS_NO_UPDATE_CHECK turns off; chaos update asks on demand instead.
chaos-serve <model.gguf> --port 8231 --context 16384
claude-chaos "read notes.txt and tell me what it says"Both ship with Chaos. On Windows there is a button for it — USE WITH CLAUDE CODE on the CHAOS page — which checks Claude Code is installed, offers to install it, asks which folder, and opens a terminal already wired up.
Pick the model on whether it calls tools, which is not the same as how good
it is at code: Qwen3-4B does, Qwen2.5-Coder-7B does not. A turn takes minutes on
a CPU machine. docs/CLAUDE-CODE.md is the whole path,
with the numbers.
chaos-pull --list # the catalogue, with the size that actually matters
chaos-pull qwen3-4b # ~2.5 GB, a good first one
chaos-run # lists the models you have
chaos-run qwen3 "The capital of France is" -n 16
chaos-run qwen3 "write a haiku" --auto # reads your machine and configures itself
chaos-probe # what can this machine run, and what should you close?
chaos-serve qwen3 --port 8080 # OpenAI API plus a chat page, both compiled in
chaos-draw "a red apple on a white table" -o apple.pngYou never type a path, and for a model split across five shards you never work out which
shard to open. Read the second column of chaos-pull --list, not the first: a 155 GB
Mixture-of-Experts model streams on a 16 GB machine; a 20 GB dense one does not, because
a dense container has no routed experts to leave on disk.
A model whose architecture has never been diffed against llama.cpp is refused by
name — a wrong forward pass produces fluent nonsense rather than an error, so the
default is refusal and --force is the override.
Five fixed models, measured 2026-09-01 in one session, on a machine with nothing
else running. The set is fixed in scripts/speed-five.tsv so this month's table can
be compared with last month's; regenerate the numbers with bash scripts/speed-five.sh
and never by hand.
Machine: i7-13650HX, 14 cores, 15.7 GiB RAM, SK Hynix NVMe (3.41 GiB/s at queue depth 8), CPU only, nothing else running.
model container resident tok/s what the number is
DeepSeek-V4-Flash 144.44 GiB 7.38 GiB 0.728 144 GB generating in 15.7 GiB of RAM
Qwen3-30B-A3B 17.28 GiB 17.28 GiB 4.41 a smaller MoE; --force, its diff has 1 FAIL
Qwen3-4B 2.33 GiB 2.33 GiB 8.27 dense, fits with room to spare
Falcon3-1B 0.98 GiB 0.98 GiB 22.31 dense, small
Qwen2-0.5B 0.37 GiB 0.37 GiB 32.00 the ceiling: nothing to stream, nothing to wait for
Three runs each, median reported; four agreed within 6%, and V4-Flash's 25% spread is a cold page cache on the first run after a build — at a fixed configuration it holds to 1%.
Your machine is not this machine, and Chaos will say so before you download
anything. chaos-model-info <model> --budget <GiB> predicts tok/s from the resident
set your machine would carry; the law behind it, across nine models spanning 23x in
size, is tok/s ~= 19 / resident GiB, with chaos-membench for the constant.
Against llama.cpp: 1.20–1.27x behind on the dense path hand-tuned, 1.23x ahead out
of the box because Chaos measures your machine and llama.cpp uses a fixed default. Both
engines alternating in one session, command lines recorded:
docs/graph/research/where-we-stand-vs-llamacpp-2026-08-16.md.
On the 144 GB streaming model no comparison is published, deliberately: llama.cpp
ranged 0.16–0.47 tok/s over eight runs of one command line here while Chaos held to
1%, so best-of would claim 1.70x and worst-of 4.35x and no ratio is honest —
docs/graph/research/the-v4flash-parity-cell-does-not-reproduce-2026-09-01.md.
If you want the fastest local inference today, use llama.cpp. Chaos is worth your time if you want an engine that owns residency and tells you the truth about your machine.
The road to v0.0.30, the release built to LTS standard. Nothing is tagged until its gate is green — 23 releases went out in 21 days once, and none of them got a stabilisation period.
v0.0.24 One truth 100% [####################] merged
v0.0.25 Guard the binary 100% [####################] merged
v0.0.26 Measure before optimising 100% [####################] merged
v0.0.27 Quality harness, then levers 100% [####################] 2 passed, 1 refused
v0.0.28 Any machine, any model 80% [################....] needs other machines
v0.0.29 Every platform, actually run 70% [##############......] macOS is what is left
v0.0.30 LTS 100% [####################] 19 of 19 cells measured
Coverage against llama.cpp. Every bar is a ratio of two counted things, both named;
filled cells are floored. The first four are test-enforced; the last is scripts/check-docs.sh,
and a node is in order when INDEX.md lists it and declares its topic — ten were in none.
CLI flags 91% [##################..] 165 of llama.cpp's 182
Chat templates 96% [###################.] 52 of its 54 names
Tokenizers 83% [################....] 5 of 6 families
Samplers 80% [################....] 16 of 20
Architectures 10% [##..................] 14 of the 141 it declares
GPU backends 20% [####................] 1 of 5, Vulkan only
Browser UI 33% [######..............] 2 of 6 — tracked, not a priority
V4-Flash speed 14% [##..................] 0.728 of 5 tok/s
Documents 69% [#############.......] 102 of 146 graph nodes in order
"Verified" here means diffed, and every one of the 14 architectures was diffed against llama.cpp at 8 prompts. The Vulkan device path is bound but not verified: it fails 1 of those 8 where the CPU path fails none.
STATUS.md |
the scoreboard. Read it before quoting any number |
CHECKLIST.md |
the tick-list: what is done, what is not, and why |
CHANGELOG.md |
every release, including the retractions |
CONTRIBUTING.md |
how to build and test, and the citation rule |
SECURITY.md |
what a running node exposes, measured |
SUPPORT.md |
what gets fixed, what stays stable, and what is not supported |
docs/graph/INDEX.md |
every working note, one line each — start here |
docs/graph/research/ |
measurements. Many refuted the idea that motivated them |
docs/graph/backlog/ |
what is planned, with a definition of done for each |
docs/graph/decisions/ |
the choices that are settled, and what they superseded |
docs/graph/reference/hard-won-facts.md |
read before proposing any optimisation |
docs/graph/history/ |
superseded, kept rather than deleted |
The code, by what a crate is for: core/ the engine — containers, ggml, I/O, residency,
architectures · cli/ the front door and the runner · network/ the server and the worker ·
gui/ the window and the installer · scripts/ CI's checks.
Apache-2.0 — see LICENSE and NOTICE. Chaos distributes no model weights; models are yours to obtain, under their own licences.
