Skip to content

Latest commit

 

History

History
139 lines (103 loc) · 7.62 KB

File metadata and controls

139 lines (103 loc) · 7.62 KB

Which model? Sizes, versions and what fits

Strata runs one model, Qwen3.8-Flash-Next, in several sizes (the same model, compressed more or less) and several versions (the original, a coding version, a fine-tune). The installer recommends one for your PC; this page explains the choice. Back to the README.

On this page: Pick by RAM · Speed · The sizes · Will it fit? · The versions · Adding another model

Pick by RAM

Your RAM Take Why
32 GB Coder it fits 32 GB, and it is made for code (with a 24 GB card, Q2_0 and IQ2_XS run too: low-RAM mode)
48 GB IQ2_XS (or Q2_0, the fastest) the larger sizes do not fit
64 GB IQ2_XS (recommended), or IQ3_XXS / IQ3_S every size fits (IQ3_S with little else open)
96 GB or more IQ3_S, or Unsloth's 4-bit (experimental) room for the largest sizes

Not sure? Take IQ2_XS - or the Coder if you mainly write code, or have 32-48 GB of RAM.

How fast is each size

Measured on an RTX 5070 (12 GB), a Ryzen 5 7600 and 64 GB of RAM:

Size Writes answers (short chat) Writes answers (128K context) Reads your prompt
Q2_0 93 tokens/s 74 tokens/s 2,170 tokens/s
IQ2_XS 79 tokens/s 63 tokens/s 2,090 tokens/s
IQ3_XXS 62 tokens/s 49 tokens/s 1,750 tokens/s
IQ3_S 53 tokens/s 46 tokens/s 1,620 tokens/s
Coder (IQ1_M) 55 tokens/s 43 tokens/s 2,180 tokens/s

Measured on an AMD RX 9070 XT (16 GB), a Ryzen 9 3900X and 47 GB of RAM (Linux, setup's own install; Q2_0 is above setup's RAM estimate for 47 GB and was installed with --model Q2_0 --yes):

Size Writes answers (short chat) Writes answers (128K context) Reads your prompt
Q2_0 60 tokens/s 48 tokens/s 1,160 tokens/s
IQ2_XS 52 tokens/s 36 tokens/s 1,110 tokens/s
Coder (IQ1_M) 44 tokens/s 33 tokens/s 1,420 tokens/s
  • Writes answers = how fast the reply appears (tokens per second; a token is about ¾ of a word).
  • Reads your prompt = how fast it takes in what you send (long documents, code, chat history), measured on a 32K-token prompt; a 4K prompt reads at 910-1,580 tokens/s. A 32K prompt takes about 15 seconds with Q2_0.

A card with more VRAM is faster, because more of the model fits on the GPU: an RTX 3090 (24 GB) should do roughly 100-140 tokens per second. All measurements, long-context numbers and estimates for other cards are in the details. The AMD measurements per card (RX 9070 XT, Radeon AI PRO R9700, RX 7800 XT, RX 9060 XT, RX 6900 XT) are in AMD_HIP.md.

Every PC is different: START-HERE.bat --calibrate measures a few engine settings on yours and keeps the fastest (about 5-10 minutes; on the PC above it made the Coder 7% faster; NVIDIA cards for now). Measured Strata on your own PC? See Community benchmark results for a report template and how to share your results in a pull request.

The sizes

Model RAM+VRAM Requirements Speed Quality
Q2_0 37.6 GB fastest good
IQ2_XS 39.2 GB fast better (recommended)
IQ3_XXS 47.0 GB slower great
IQ3_S 54.8 GB slowest best: matches the full model on the published tests (original model only)

The download is 66-76 GB for the three smaller sizes (details); the first start also fetches the MTP draft layer (~6 GB, +1 GB with images).

Will it fit?

Shard 1 is the part of the model that gets loaded when it starts: its experts go into your RAM, the rest onto your graphics card (the second shard, a 29 GB lookup table, stays on the SSD). So it fits when your RAM is at least shard 1 + about 10 GB for Windows and your other programs. With 64 GB of RAM every size fits (IQ3_S with little else open); with 48 GB, Q2_0 and IQ2_XS. A bigger graphics card makes it faster, but it doesn't lower the RAM needed

  • except in the low-RAM mode below.

A big graphics card and little RAM

On a PC whose RAM cannot hold the experts beside the system, setup maps them from the model's files instead and keeps in RAM only what the graphics card does not hold (the low-RAM mode, chosen by setup). For example, a 32 GB PC with a 24 GB GPU runs Q2_0, IQ2_XS and the Coder this way, and a 32 GB PC with a 12-16 GB GPU the Coder. With a small card most experts then come from the SSD and it is much slower (setup says so). START-HERE.bat --setup --low-ram on|off overrides the choice.

The versions

Qwen3.8-Flash-Next

The original, in all four sizes. With images, and with the experimental speed projection as an option.

Coder

Coder - ISTA-DASLab's coding version: half of the experts removed, keeping the ones that code, tool use and images need (91% of the full model's SWE-bench Verified score, 99% of LiveCodeBench, by its authors). One size (IQ1_M: its experts stored like IQ3_S): shard 1 is 29.6 GB, so it fits a PC with 32 GB of RAM, runs 262K context on 64 GB, and reads long prompts the fastest of all. It is weaker outside code, and that includes Chinese and other CJK text (#438: Chinese answers came out wrong or looping where English was fine). For general chat or CJK text, take a size that keeps every expert: Q2_0, IQ2_XS or IQ3_S. More: details.

START-HERE.bat --setup --family coder

Swift 1.5

Swift 1.5 - a fine-tune by UkisAI that thinks much shorter before answering, so you get the answer sooner, with about the same quality. Same speed per token, and about the same RAM as the same size of the original (no IQ3_S). Its own license applies (see its page). More: details.

START-HERE.bat --setup --family swift --model IQ2_XS

Unsloth UD-Q4_K_XL (experimental)

Unsloth's 4-bit UD-Q4_K_XL (experimental) is the fourth version in setup's menu (--family unsloth): the closest to the full model, but a 111 GB download whose 77 GB of experts do not fit in RAM. Strata keeps your RAM minus 24 GB of them in RAM and reads the rest from the SSD while it answers: 7-8.5 tokens/s on a 64 GB PC with a 12 GB GPU, several times slower than the sizes above, and long prompts are slow. It needs 48 GB of RAM or more, an NVMe SSD and one NVIDIA GPU (no images yet). Details and measurements: UD-Q4_K_XL.

START-HERE.bat --setup --family unsloth --model UD-Q4_K_XL

OrcaRouter Uncensored IQ3_XXS

For OrcaRouter's Flash-Next Uncensored IQ3_XXS, see the manual compatibility setup. It needs an explicit packing conversion and is not an installer menu option.

Adding or switching models

You can add another model any time with SETUP.bat (the same as START-HERE.bat --setup; on Linux ./setup.sh --setup). Files another model shares are not downloaded again (the Coder uses the original's shard 2 and vision encoder). With more than one model installed, START-HERE.bat asks which one to start; run-<model>.bat (Linux: run-<model>.sh) starts one directly.