Strata runs one model, Qwen3.8-Flash-Next, in several sizes (the same model, compressed more or less) and several versions (the original, a coding version, a fine-tune). The installer recommends one for your PC; this page explains the choice. Back to the README.
On this page: Pick by RAM · Speed · The sizes · Will it fit? · The versions · Adding another model
| Your RAM | Take | Why |
|---|---|---|
| 32 GB | Coder | it fits 32 GB, and it is made for code (with a 24 GB card, Q2_0 and IQ2_XS run too: low-RAM mode) |
| 48 GB | IQ2_XS (or Q2_0, the fastest) | the larger sizes do not fit |
| 64 GB | IQ2_XS (recommended), or IQ3_XXS / IQ3_S | every size fits (IQ3_S with little else open) |
| 96 GB or more | IQ3_S, or Unsloth's 4-bit (experimental) | room for the largest sizes |
Not sure? Take IQ2_XS - or the Coder if you mainly write code, or have 32-48 GB of RAM.
Measured on an RTX 5070 (12 GB), a Ryzen 5 7600 and 64 GB of RAM:
| Size | Writes answers (short chat) | Writes answers (128K context) | Reads your prompt |
|---|---|---|---|
| Q2_0 | 93 tokens/s | 74 tokens/s | 2,170 tokens/s |
| IQ2_XS | 79 tokens/s | 63 tokens/s | 2,090 tokens/s |
| IQ3_XXS | 62 tokens/s | 49 tokens/s | 1,750 tokens/s |
| IQ3_S | 53 tokens/s | 46 tokens/s | 1,620 tokens/s |
| Coder (IQ1_M) | 55 tokens/s | 43 tokens/s | 2,180 tokens/s |
Measured on an AMD RX 9070 XT (16 GB), a Ryzen 9 3900X and 47 GB of RAM (Linux, setup's own install; Q2_0 is
above setup's RAM estimate for 47 GB and was installed with --model Q2_0 --yes):
| Size | Writes answers (short chat) | Writes answers (128K context) | Reads your prompt |
|---|---|---|---|
| Q2_0 | 60 tokens/s | 48 tokens/s | 1,160 tokens/s |
| IQ2_XS | 52 tokens/s | 36 tokens/s | 1,110 tokens/s |
| Coder (IQ1_M) | 44 tokens/s | 33 tokens/s | 1,420 tokens/s |
- Writes answers = how fast the reply appears (tokens per second; a token is about ¾ of a word).
- Reads your prompt = how fast it takes in what you send (long documents, code, chat history), measured on a 32K-token prompt; a 4K prompt reads at 910-1,580 tokens/s. A 32K prompt takes about 15 seconds with Q2_0.
A card with more VRAM is faster, because more of the model fits on the GPU: an RTX 3090 (24 GB) should do roughly 100-140 tokens per second. All measurements, long-context numbers and estimates for other cards are in the details. The AMD measurements per card (RX 9070 XT, Radeon AI PRO R9700, RX 7800 XT, RX 9060 XT, RX 6900 XT) are in AMD_HIP.md.
Every PC is different: START-HERE.bat --calibrate measures a few engine settings on yours and keeps the fastest
(about 5-10 minutes; on the PC above it made the Coder 7% faster; NVIDIA cards for now). Measured Strata on your own
PC? See Community benchmark results for a report template and how to share your results
in a pull request.
| Model | RAM+VRAM Requirements | Speed | Quality |
|---|---|---|---|
| Q2_0 | 37.6 GB | fastest | good |
| IQ2_XS | 39.2 GB | fast | better (recommended) |
| IQ3_XXS | 47.0 GB | slower | great |
| IQ3_S | 54.8 GB | slowest | best: matches the full model on the published tests (original model only) |
The download is 66-76 GB for the three smaller sizes (details); the first start also fetches the MTP draft layer (~6 GB, +1 GB with images).
Shard 1 is the part of the model that gets loaded when it starts: its experts go into your RAM, the rest onto your graphics card (the second shard, a 29 GB lookup table, stays on the SSD). So it fits when your RAM is at least shard 1 + about 10 GB for Windows and your other programs. With 64 GB of RAM every size fits (IQ3_S with little else open); with 48 GB, Q2_0 and IQ2_XS. A bigger graphics card makes it faster, but it doesn't lower the RAM needed
- except in the low-RAM mode below.
On a PC whose RAM cannot hold the experts beside the system, setup maps them from the model's files instead and keeps
in RAM only what the graphics card does not hold (the low-RAM mode, chosen by setup). For
example, a 32 GB PC with a 24 GB GPU runs Q2_0, IQ2_XS and the Coder this way, and a 32 GB PC with a 12-16 GB GPU the
Coder. With a small card most experts then come from the SSD and it is much slower (setup says so).
START-HERE.bat --setup --low-ram on|off overrides the choice.
The original, in all four sizes. With images, and with the experimental speed projection as an option.
Coder - ISTA-DASLab's coding version: half of the experts removed, keeping the ones that code, tool use and images need (91% of the full model's SWE-bench Verified score, 99% of LiveCodeBench, by its authors). One size (IQ1_M: its experts stored like IQ3_S): shard 1 is 29.6 GB, so it fits a PC with 32 GB of RAM, runs 262K context on 64 GB, and reads long prompts the fastest of all. It is weaker outside code, and that includes Chinese and other CJK text (#438: Chinese answers came out wrong or looping where English was fine). For general chat or CJK text, take a size that keeps every expert: Q2_0, IQ2_XS or IQ3_S. More: details.
START-HERE.bat --setup --family coder
Swift 1.5 - a fine-tune by UkisAI that thinks much shorter before answering, so you get the answer sooner, with about the same quality. Same speed per token, and about the same RAM as the same size of the original (no IQ3_S). Its own license applies (see its page). More: details.
START-HERE.bat --setup --family swift --model IQ2_XS
Unsloth's 4-bit UD-Q4_K_XL (experimental) is the fourth version in setup's menu (--family unsloth): the closest
to the full model, but a 111 GB download whose 77 GB of experts do not fit in RAM. Strata keeps your RAM minus 24 GB
of them in RAM and reads the rest from the SSD while it answers: 7-8.5 tokens/s on a 64 GB PC with a 12 GB GPU, several
times slower than the sizes above, and long prompts are slow. It needs 48 GB of RAM or more, an NVMe SSD and one
NVIDIA GPU (no images yet). Details and measurements: UD-Q4_K_XL.
START-HERE.bat --setup --family unsloth --model UD-Q4_K_XL
For OrcaRouter's Flash-Next Uncensored IQ3_XXS, see the manual compatibility setup. It needs an explicit packing conversion and is not an installer menu option.
You can add another model any time with SETUP.bat (the same as START-HERE.bat --setup; on Linux
./setup.sh --setup). Files another model shares are not downloaded again (the Coder uses the original's shard 2 and
vision encoder). With more than one model installed, START-HERE.bat asks which one to start; run-<model>.bat
(Linux: run-<model>.sh) starts one directly.