One model can run across several NVIDIA cards in one PC. The layers are split into contiguous ranges, one per GPU: the first card runs layers 0 to K-1, the next card runs K onward, and so on; the last card also runs the output head and the draft (MTP) layer. Each card keeps an expert cache for its own layers only, so two cards hold about twice the experts one card holds - for the Coder model on a 16 GB + 24 GB pair, nearly all of them, which is where the speed comes from (decode then barely touches the CPU pool).
This is pipeline (layer) parallelism, not tensor parallelism: a token crosses from one card to the next once per
verify window (a few hundred KB through pinned RAM), not twice per layer. No NVLink or peer-to-peer access is
needed; cards on x4 or x1 slots work, and the PCIe share of each card is probed on its own link. A --pcie-frac you give
is every card's share and skips those probes; there is no per-card setting yet.
Nothing to type. START-HERE.bat (Linux: ./setup.sh) lists your NVIDIA cards and says for each one whether
Strata can use it:
Your NVIDIA GPUs:
GPU 0: NVIDIA GeForce RTX 5080, 16 GB VRAM - can be used
GPU 1: NVIDIA GeForce GTX 1080 Ti, 11 GB VRAM - not supported - older than the RTX 20 series (compute capability 6.1; Strata needs 7.5 or newer)
GPU 2: NVIDIA GeForce RTX 3090, 24 GB VRAM - can be used
...
1) GPU 0 (NVIDIA GeForce RTX 5080, 16 GB) + GPU 2 (NVIDIA GeForce RTX 3090, 24 GB) together (recommended)
2) GPU 2 (NVIDIA GeForce RTX 3090, 24 GB) only
3) GPU 0 (NVIDIA GeForce RTX 5080, 16 GB) only
Which GPUs? [1]:
When two or more cards can share the model, the two best together are recommended (the newest generation first: it becomes the main card). A model installed on one card asks once, at its next start, whether to use both from now on; the answer is kept.
Choosing yourself (at setup or at any start):
--gpus 0,2 these cards together, as nvidia-smi numbers them; the first is the main one. Remembered.
--gpus all every card that can share the model
--gpu 0 one card (at a start: for that start only)
--layer-split auto (default) or the first layer of each later card, e.g. 18 or 16,32
Not supported (setup says so and names the cards that can be used instead):
- a card older than the RTX 20 series (compute capability below 7.5: GTX 10 and older);
- a card with less than 8 GB of VRAM, together with others (each card holds a copy of the dense weights and its
own prompt buffers) - unless you name it with
--gpus: then setup says the risk and asks (--yeswith the named cards goes ahead); - Intel GPUs, and a mix of NVIDIA and AMD cards. (AMD cards share a model among themselves:
./setup.sh --backend hip --gpus 1,0, see AMD_HIP.md.)
Or edit an existing config (strata-*.json), then restart:
"gpu": [0, 2],
"layer_split": "auto""layer_split" is "auto" (placed by each card's free VRAM) or the first layer of each later card: one rising
number per card after the first, not a count of layers per card. With 4 cards and a 48-layer model, "24,36,42"
(or [24, 36, 42]) puts layers 0-23 on the first card, 24-35 on the second, 36-41 on the third and 42-47 on the last.
The server checks it before the start and says what is wrong (0.1.39, #644).
Skip the split when the first card holds everything (opt-in, 0.1.31): "split_skip_if_fits": true in the config
(engine flag --split-skip-if-fits, with --layer-split auto) runs on the first card alone when it holds every
profiled expert plus the context's KV, the draft layer and the reserve, and says so in the log; otherwise the split
stays. On an R9700 32 GB + RX 9070 XT the R9700 holds all of the Coder's experts: with the flag 4K prompts read at
1,776 tok/s instead of 1,244 (split) and decode runs at ~60 tok/s instead of ~51 (16K prompts ~5% slower than split).
Short prompts on a split (0.1.32, #340). 0.1.30 gave each card's prompt path a loan from its own expert cache,
refilled after every request; on cards that hold nearly all their experts that cost short prompts up to a third of
their speed. 0.1.32 refills all cards at once, uses a smaller streaming ring on a split, and lets a card with free
VRAM keep its own prompt buffers - the same output as 0.1.31, measured on an R9700 + RX 9070 XT: 2K prompts 993 ->
1,265 tok/s, 16K 1,852 -> 1,950, decode unchanged. STRATA_SPLIT_OWN=1 (opt-in) gives every card its own buffers:
2K 1,450 and 16K 2,227 tok/s there, but a full card then keeps a different set of experts resident, so the output
differs from the default's (stable and coherent); STRATA_SPLIT_OWN=auto does that only where the buffers are at
most 12% of each card's VRAM.
The idle card can help one-chunk prompts (opt-in, STRATA_PREFILL_HELP=1). A prompt that fits one chunk runs the stages one after the other, so while
one card reads its layers the other idles. With it on, each stage hands a share of its streamed experts to the idle card: it
streams them over its own PCIe link into its own (lent) prompt buffers, computes their rows on the MMQ path and sends
them back - --peer-device's peer streaming, without P2P (the activations and the rows go through mapped host memory,
read by copy kernels, so they do not queue behind the expert blobs on either card's copy engine). The share falls with
the prompt (0.41 of the streamed experts at 1.5K tokens, 0.32 at 3K) and is off from ~3.3K tokens, where the stages
overlap anyway and no share paid. Measured on 2x RTX 3090 (UD-Q4_K_XL, no P2P), prompt tok/s without / with: 1.1K 496 /
723, 1.5K 674 / 833, 2K 878 / 1,128, 2.5K 1,017 / 1,296, 3K 1,260 / 1,440; 4K and 8K unchanged, decode unchanged. It
costs no VRAM (the idle card's own prompt buffers) and ~110 KB of mapped host memory per token of the largest chunk it
helped (~360 MB at 3.3K tokens). The rows it computes round like a different MMQ grouping, so the output is not
bit-identical to the default's (it is repeatable: same prompt, same output), which is why it is opt-in:
STRATA_PREFILL_HELP=1 turns it on. Native packs on the MMQ prompt path only (not with the fused prompt kernels,
STRATA_PF_FUSED=1, nor with --peer-device). With it on, STRATA_PREFILL_HELP_FRAC=f fixes the share.
The engine flags behind it: --layer-split K1[,K2..]|auto and --split-device D1[,D2..] (the later stages'
devices; default the next visible ones). --layer-split K --split-device 0 runs both stages on one card sharing
everything - the bit-exact check of the hand-off, not a speed mode.
Each card loads only its own layers' dense weights (0.1.39, PR #639) with explicit split points (--layer-split 27, not auto): every card used to keep a full copy (~3.4 GB for the Coder) though its stage reads only its own
layers, and the VRAM it frees goes to that card's expert cache (2x MI50 16 GB, Coder: 8,819 -> 10,626 experts in
VRAM, decode 39.2 -> 41.7 tok/s). Opt-in for now, on AMD and NVIDIA alike: STRATA_STAGE_TRIM=1 (please report how
it goes).
A card holding more experts can change which experts run on the GPU, so the output can differ slightly from a run
without it.
The resident RAM mode works on a split (--resident-experts, the low-RAM mode). The RAM copy of the experts
leaves out the ones every card's cache holds, not only the first card's, and when the rest does not fit whole it keeps
the hottest by the expert profile over all the layers. An adaptive swap copies the evicted expert back into RAM from
the card that owns its layer. Before, --resident-experts with a split ran as --mmap-experts, and setup recommended
one card in the low-RAM mode; with an engine that has it (RESIDENT_SPLIT_ENGINE in setup.py), setup keeps the cards
together and the experts no card holds in RAM. Swift 1.5 IQ3_XXS at
160K (q4_0 KV, --prefill 4096, --spec 4 with the stock draft layer), RTX 4060 Ti (layers 0-19) + RTX 5080 (20-47),
i9-14900KF, 32 GB of RAM, Windows 11, four greedy prompts at a time, decode tok/s:
| first four prompts | after three more rounds | |
|---|---|---|
5080 alone, --resident-experts |
25 | 29 |
split, --mmap-experts (what --resident-experts became on a split) |
32 | 64 |
split, --resident-experts (22 GiB of experts locked in RAM) |
71 | 69 |
The split with --mmap-experts catches up once the OS file cache holds the experts, on a PC with nothing else
running; the resident copy is there from the first request and stays locked when other programs need the RAM. With
--pcie-frac 0 --adapt-every 0 the split's greedy output is the same with either mode.
A separate VRAM reserve for the later cards: --vram-reserve-later-mib N (default: --vram-reserve-mib's value).
The card that drives the monitors needs more headroom than one that drives none; with the display on the last card,
--vram-reserve-mib 300 --vram-reserve-later-mib 1800 gives the first card's cache that VRAM.
auto tries every placement (all of them for two, three or, since 0.1.40, four cards; proportional to the free VRAM beyond that) and keeps the one whose caches would hold the most of the expert profile, hottest pairs weighted most; ties go to the placement that leaves the fullest card the most room. The startup log prints the choice:
strata generate: layer split auto: K=19 - the caches hold 11767 of 12288 profiled pairs (fullest device 100%)
strata serve: layer split: layers 0-18 (CUDA0), 19-47 (CUDA1), one hand-off per window
- every card: a copy of the dense weights (~3.4 GB for the Coder), its own session state (the KV cache of the full context), its verify window and its prompt-path buffers, and an expert cache for its layers filled from the profile;
- the last card: also the output head and the draft layer (~0.8 GB);
- host RAM: the expert arena once, shared by all cards (the CPU pool computes whatever no card holds).
Prompts are read in chunks that flow through the cards in turn; while a later card reads chunk c, the first card already reads chunk c+1. Conversation checkpoints save and restore every card's state; the adaptive expert swaps copy into the card that owns the layer.
- Works across cards (bench/results/2026-09-29-layer-split-limits):
- images (
--vision): each card keeps its own image-position table; - control vectors and the experimental speed projection: each card holds the vector's tables, switched on and off per request on all of them;
- KV streaming (
--kv-resident): each card streams the KV of its own session; - mid-prompt checkpoints (
--prompt-cache-every): each card saves its part of a checkpoint when it has read that chunk; - the older helper-GPU caches (
--expert-cache-remote, docs/SECOND_GPU.md): they take the visible GPUs no stage runs on, and hold only experts no stage's cache holds. On the test rig, a 2080 Ti helper made decoding slower, as it did without a split: its per-layer round trip costs more than the CPU pool needs for those experts.
- images (
--mmap-expertsneeds a canonical pack (experts.bin), with or without a split; a native (IQ) pack says so at start.- The prompt path has its own buffers on every card (1.5 GB each at the default 2048-token chunk;
--prefill 1024halves that) instead of borrowing cache slots as one card does. An explicit--expert-cacheon the first card is capped to leave room for them. - Under WDDM (Windows, and WSL2) only 8 GiB of the expert arena is pinned (more, mapped into two GPU contexts,
leaves WDDM refusing allocations); the rest streams through the pinned staging ring. A Linux driver has no such
limit, so there the whole arena is pinned (since 0.1.31; the cap cost a 4090 + 3060 split two thirds of its
prompt speed, #253).
STRATA_ARENA_PIN_GIB=Npins at most N GiB,0the whole arena, on any OS. - Every card needs compute capability 7.5 (RTX 20 or newer). The pre-sm_80 QSA scorer path is fp32 FMAs, so a Turing card runs the same kernels instead of the tensor-core prompt attention.
The Coder on an RTX 5080 + RTX 3090 (Ryzen 9 9950X3D), 32K context; details in
bench/results/2026-09-29-layer-split/:
| Prompt 16K / 28K tok/s | Decode story / code tok/s | |
|---|---|---|
| 5080 alone | 1,726-2,017 / 1,970 | 83-87 / 88-105 |
| 5080 + 3090, best split (K=26) | 2,039 / 2,357 | 84 / 110 |
| 5080 + 3090, auto (K=22) | 2,037 / 2,073 | 80 / 109 |
- Prompts gain the most (+18-20%): each card reads its own layers of the chunk while the other reads the next.
- Decode is on par with the faster card alone, and ahead on code. Once both caches hold nearly every routed expert, the per-layer GPU time decides.
- Correctness: one GPU is byte-identical to 0.1.20, and the hand-off itself is bit-exact.
Which cards and in what order:
- Put the fastest card first; auto gives it as many layers as its cache allows.
- Leave out a much slower card when two already hold the model. An RTX 2080 Ti as a third card made the 5080 + 3090 pair slower (68 / 90 tok/s decode): every extra card costs its own round per window.
- More cards pay off when the model's routed experts do not fit the faster ones.
With a split the cards take turns on a verify window: the first card runs its layers and hands off, then waits while
the last card runs the rest, the head and the draft. --pipeline-windows 2 lets the first card start the next
window while the last card still verifies this one. The next window is a guess: that this window is accepted whole
and that its bonus token is the one the draft layer predicts (the draft layer is run on through the drafts of the
window in flight). When the guess holds, half of the next window is already done; when it does not, the first card
puts its state back (a copy of its recurrent state taken while the window ran) and the next window is built from the
real tokens. A guessed window is only started when the draft layer's estimate says it is likely to be kept. Windows
copied from earlier context (--suffix-draft) are guessed past as well, which is where edits that copy text gain
most. --pipeline-windows 1 overlaps only the short prompt reads that go through the verify windows
(--short-read).
Two cards, exactly two stages, --serve. In the config:
"args": [ ..., "--pipeline-windows", "2" ],
"layer_split": "20"
- Cost: a second verify window on each card, 160 MiB more kept out of each card's expert cache, plus two copies
of the first card's recurrent state (about 3 MiB per GDN layer it runs) on the first card with
2. - Same text: the last card only ever runs windows that are verified, and every window row computes what it
would in any other window, so the tokens are the serial loop's. With
STRATA_IQ_MT_MIN=1 --pcie-frac 0 --adapt-every 0the greedy output is identical bit for bit to the serial loop's with the same expert caches. The pipeline keeps its VRAM out of the caches, so against a run without the flag a few experts move from a card to the CPU, which rounds them differently, and a near-tie can flip (a serial run given the same caches through--vram-reserve-mibmatches it exactly). - Off, with one line in the log saying why, with
--batchslots,--peer-device, the helper caches (--expert-cache-device1..3,--remote-expert-opt), a split into three or more stages or onto one GPU (--split-device 0), or no draft layer. A request with repetition penalties (penalty_last_n) or coupled draft sampling decodes serially. - Measured (Swift 1.5 IQ3_XXS, 160K context, q4_0 KV, the stock draft layer, RTX 4060 Ti (layers 0-19) +
RTX 5080 (20-47), i9-14900KF, 32 GB of RAM with the resident RAM mode on the split (#848); greedy, 500 tokens, two
interleaved pairs of three rounds, decode tok/s): Python code 90.4 -> 103.1, C code 76.6 -> 81.6, English prose
63.0 -> 71.1, Italian prose 41.6 -> 46.3, a copy-heavy edit (a 5 KB file back with a rename) 90.5 -> 117.9; mean
72.4 -> 84.0 (+16%). It pays when the windows are GPU-bound: with the experts read through the OS file cache
(
--mmap-expertson that 32 GB PC) the file reads dominate and it measured no faster.
The STRATA_PIPELINE_* tuning and test variables (THETA, FORCE_MISS, SWITCH, LOG, TRACE and the like) are read only with
STRATA_PIPELINE_DEBUG=1. --pipeline-windows and --adapt-async 1 exclude each other (the engine says so and keeps the
pipeline).
With a layer split, --batch N --batch-groups G --trim-stage-weights decodes several conversations together and
pipelines them through the cards: see BATCHING.md.