Skip to content

Sync upstream Strata v0.1.39 into Lanes - #24

Merged
rhgo1749 merged 221 commits into
mainfrom
sync/upstream-0.1.39
Oct 5, 2026
Merged

rhgo1749 merged 221 commits into
mainfrom
sync/upstream-0.1.39

Conversation

@rhgo1749

@rhgo1749 rhgo1749 commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

Summary

Merge upstream Strata v0.1.39 (6f32ec070f23ced9f50e704d854d775da52591ab) onto the Lanes 0.1.38 baseline while preserving the independent whole-GPU lane runtime.

Key integration decisions:

  • keep the Lanes shared-arena leader/follower readiness contract on top of upstream 0.1.39 expert-source/pinned-arena changes;
  • use upstream 0.1.39 server.py and native conversation-cache telemetry, avoiding the old Lanes DONE field extension that collided with upstream's offloaded-expert counter;
  • preserve Lanes tool_choice / malformed tool-call handling while retaining upstream frontend/reasoning changes;
  • strip inherited parallel, --batch, --slots, --batch-groups, and --trim-stage-weights from ordinary lanes so upstream internal batching cannot silently change the Lanes concurrency unit;
  • route /v1/responses through the normal Lanes generation admission path.

Validation

Reference host, CUDA 13.4.92, sm_120:

  • python3 -m unittest discover -s serve -p 'test_*.py': 333 passed, 7 skipped.
  • Fresh Release strata build with STRATA_BUILD_TESTS=OFF: pass.
  • Fresh Release strata build with STRATA_BUILD_TESTS=ON: pass.
  • Focused CTests: pinned_shared_test, expert_profile_save_test, file_expert_source_test: 3/3 pass.
  • Live IQ3_XXS single-lane generation smoke: pass.
  • Live 3-lane shared-arena smoke on GPUs 0/1/2: three concurrent requests routed one per lane and all returned HTTP 200; leader populated the ~39.97 GiB arena once and both followers logged shared population ready; source load skipped.
  • Live POST /v1/responses through the Lanes supervisor: HTTP 200, X-Strata-Lane-Index: 0, output text OK.

The first test-enabled build attempt showed an unresolved cuGetErrorString, but this was traced to a CMake directory created by an initial configure with no CUDA compiler: its cached CMAKE_CUDA_FLAGS_RELEASE was empty. A fresh configure with /usr/local/cuda-13.4/bin/nvcc from the first invocation restored -O3 -DNDEBUG and the full test-enabled executable linked successfully.

The compatibility/promotion gate above remains distinct from historical 0.1.30/0.1.31/0.1.38 benchmark evidence; those measurements keep their original engine labels.

0.1.39 topology recheck

After the compatibility gate, the topology comparison was reduced to the decision-bearing points and rerun only where evidence was missing.

Fixed decode scorecard:

  • M=1: one independent lane 73.63 ± 1.67 tok/s vs retained same-binary three-GPU layer-split control 120.62 ± 1.24 (M=1 was not rerun under the exact batch3/groups3/trim config);
  • M=2: two independent lanes 143.19 ± 3.06 tok/s vs the fixed explicit 18,34 pipeline with --batch 3 --batch-groups 3 --trim-stage-weights and shared expert arena 147.84 ± 2.26 tok/s (+3.25%);
  • M=3: three independent lanes 192.16 ± 4.11 tok/s vs the same fixed pipelined layer split 209.66 ± 4.02 tok/s (+9.11%).

Cold-prefill crossover:

  • three ~15K cold prompts: lanes 5901.34 ± 55.16 tok/s vs pipelined split 3289.19 ± 18.83;
  • three ~110K cold prompts: pipelined split 6028.09 ± 14.50 tok/s vs lanes 5822.71 ± 4.49.

All retained PP requests report cache_n=0. The 15K/110K nonce text is workload/length matched rather than byte-identical across topology arms. The ordinary FIFO and batch-only runs are retained only as controls; broad oversubscription, heterogeneity, workload-sensitivity and PSS probes remain raw/supporting evidence rather than headline topology benchmarks.

This reopens the always-on three-GPU layer-split server as a serious default-topology candidate, but production promotion still depends on mixed prompt/output lengths, session behavior, queue/tail latency, power, failure isolation and operational flexibility.

maxfridbe and others added 30 commits September 29, 2026 11:09
The Strata engine is CUDA. On a GPU without it - written for and measured on an Intel Arc Pro
B70 - the same GGUF weights run through llama.cpp's SYCL backend. serve/engine_llama.py puts
llama-server behind Engine.generate(ids, max_new, sampling, cancel), so the OpenAI and
Anthropic endpoints, streaming, tool calls, MCP and the web app are unchanged above it.

It sends the prompt as token ids (the server's own tokenizer, from the same GGUF) with
return_tokens, streams ids back with llama-server's timings mapped onto Strata's `last`,
heartbeats while a long prompt is read, and stops the server by closing the connection when a
client goes away. Attach to a running server ("llama": {"url"}) or spawn one ("llama": {"exe"}).

Not yet: images (Strata's Vision produces embeddings for the CUDA engine; llama-server takes
the image itself) - start without "vision" and image requests are refused cleanly.

Tests: a scripted llama-server (the wire format observed on b29c606), plus a live round trip
when STRATA_LLAMA_URL is set. Measured on the B70 with Coder IQ1_M: 23-25 tok/s decode,
~150 tok/s prefill, correct code, GPU-bound.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…erver itself

Without nvidia-smi, setup now looks for an Intel discrete GPU in sysfs (vendor 0x8086 under xe or
i915, named by PCI id - lspci's database does not know an Arc Pro B70 yet) and takes a path with no
compiler and no CUDA toolkit: download the model, export the tokenizer, find a llama-server
(--llama-server PATH, one on PATH, or the B70 container image via docker), write a config with a
"llama" block and a start-llama.sh, and a run script that uses --engine llama.

The CUDA engine's rules that do not apply are stepped around rather than deleted: the compute
capability and driver-version checks, the experts-in-RAM gates (on llama.cpp the fit is VRAM,
shard 1 = the download minus the 28.8 GB lookup table), the MTP draft download (5 GB the llama
engine cannot use), and the engine's own argument list.

engine_from_config runs the config's start script when nothing answers at the url, so the run
script is one click and closing the server stops the container.

Verified on an Arc Pro B70: setup runs steps 1-7 clean against the downloaded Coder IQ1_M.
docs/INTEL.md has the whole picture and the measurements.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…webui here)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
About 20 KiB of KV per token (only 12 of 48 layers are full attention). 32k, 64k, 96k and 128k all
load with the whole model on the card, 128k with 1.6 GB to spare; 160k is under the safety margin
and 256k does not fit - asking for it spilled VRAM into host RAM and took the machine down until its
watchdog reset it, so the doc says so in bold.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…4 tok/s

Prompt reading scales up with prompt size (150 tok/s at 2.7k, 424 tok/s at 105k), so the full
window costs about five minutes, not fifteen. 30.3 GB in use with 1.5 GB free at 131,072.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
… kernel parity tests pass on the Arc Pro B70

The port lives in sycl/ and changes no upstream file: sycl/src and sycl/include hold only the files the
migration touched, sycl/CMakeLists.txt takes everything else from the original tree.

How it was made, reproducibly (sycl/tools/): a dev image (the llama.cpp SYCL image + SYCLomatic 2025.3 +
CUDA 12.8 headers from the pip wheels), migrate.sh (compilation database + dpct over 86 translation
units), fixups.py (every hand fix with its reason, idempotent), build.sh (icpx, ninja).

What dpct got wrong and the script repairs: CUDA's null stream became a null sycl::queue* (strata::q_of
maps it to the default queue); __ldg((const float*) p) lost its cast and read one byte; __fadd_rn lost
parentheses; ggml's lookup tables were threaded through kernel parameters with the wrong table per
template; helper headers renamed since 2025.3; cudaGraphUpload, graph introspection and %globaltimer
have no equivalents. Two compiler flags are load-bearing: -fp-model=precise, and the device compiler's
-cl-fp32-correctly-rounded-divide-sqrt - the Arc's fp32 divide is not correctly rounded by default, and
quantize_act_parity went from 303k mismatches to byte-exact with it.

Not ported: the three kernels with inline PTX (mma.sync, ldmatrix, cp.async) take their existing
pre-sm_80 fallback; the launchers refuse the device. The XMX joint_matrix versions are next.

The engine itself starts on the card, loads the pack, the projections, the PLE table and the 23.4 GiB
expert arena, and is then stopped by the test rig's memory guard: Strata keeps the experts as a resident
host copy, and the rig has 23 GiB of RAM (upstream asks for 32). A streaming fill of the VRAM cache
from the GGUF is the next piece; the whole model fits in the 32 GB card.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
… Pro B70

--stream-experts (GgufExpertSource): every expert goes straight from the GGUF into the VRAM cache through a
small ring of staging buffers; no resident host arena, so the engine runs on a 23 GiB machine.

What it took to get from "loads" to "<think>\nThe user wants a Python function":
- the doorbell handshake through host-mapped memory: volatile device loads do not bypass the caches on
  Intel; system-scope atomics (strata::sys_load/sys_store) do. A device-side spin is bounded (kSpinMax): an
  unbounded one wedged the GT twice when its process died.
- STRATA_VERIFY_NO_HOST=1 + STRATA_VERIFY_DEVICE_PLAN=1: with every expert resident the GPU plans each
  layer itself and the host waits for the whole window graph; per-layer ring visibility inside a graph is
  not reliable on this platform (measured with sycl/probe).
- -fsycl-default-sub-group-size=32: the kernels are written for 32-lane warps; dpct pinned only 134 of 289.
- cudaMemcpy is synchronous; dpct's default-queue memcpy had no wait, and the streaming ring reuses its
  buffers - the expert cache filled from overwritten memory (non-deterministic residuals, NaN by layer 5).
- native_expert_parity hand-ported (dpct cannot parse it without ggml-cpu.h): the GPU native expert
  kernel matches the float reference on real IQ1_M rows.
- verify.cpp debug readbacks (STRATA_VERIFY_DEBUG): the per-layer residual ladder that found all of it.

Not yet: speed (0.17 tok/s - the spin budgets), the segfault at exit, the tensor-core kernels.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…, eager stage profiler, docs

Measured on the B70: 64 tokens in 52.0 s, of which ~47 s is the first window (the runtime JIT-compiles
every kernel on first use: 6 tokens take 47.5 s, 64 take 52.0 s). The output is the model's real answer.

Debug tooling kept in the port: STRATA_VERIFY_EAGER=1 runs the window on the queue instead of as a graph
and turns the stage stamps into host clocks, so STRATA_VERIFY_PROFILE=1 prints ms per stage without a
device timer; STRATA_VERIFY_DEBUG=1 prints the residual after every layer and the round's phase times.
build.sh takes BUILD_DIR and AOT (an ahead-of-time bmg-g31 build removes the JIT cost).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…st token in 0.85 s, no JIT

The -Xs device flag was mis-quoted; fixed. Measured on the B70 with the same 19-token prompt, 32 tokens:
decode 15.17 tok/s, prefill 28.79 tok/s, a 4-token window 73 ms on the card. The persistent JIT cache
(SYCL_CACHE_PERSISTENT=1) is the alternative for the JIT build: 15 MB, second start 15.6 tok/s.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…er on iq3_xxs layers

A window routes the same expert for up to 6 tokens; the grouped kernels recomputed the dequantised weights
(grid lookups, sign unpacking, the byte-wise compare/xor/subtract emulations) once per token. Multi<TY>
splits each of the five formats this model uses (iq2_s, iq3_xxs, iq3_s gate/up; iq4_nl, q2_0 down) into
prep (once) and acts + finish (per token); the kernels take entries four at a time. Lanes per row is a
compile-time constant (STRATA_EXPERT_LANES, 8): 32 lanes on an 80-call row was mostly reduction.

Measured with the new NATIVE_BENCH=1 mode of native_expert_parity (10 experts x 4 entries, one layer):
iq3_xxs/iq4_nl 1.467 -> 0.649 ms, iq2_s/q2_0 ~0.68 -> 0.579 ms, iq3_s/iq4_nl 0.750 ms; parity unchanged.
strata::dp4a wraps dpct::dp4a (sycl::ext::oneapi::dot_acc is the same emulation with a header that breaks
the link across translation units).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…tok/s decode

route<NE> templated on the expert count; native_router_top10_multi_ne dispatches 256/512. Both call sites
(the verify window and the per-token moe_route) take it for the Coder. 64 tokens: 16.25 -> 16.95 tok/s,
same output. Also: Dockerfile.unitrace (per-kernel Level Zero profiling image).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…: the mmvq kernels run at 110-130 GB/s warm

- exit segfault: dpct's static global_memory/constant_memory objects destructed after the default queue
  (queue::get_device in the core); heap-allocated, never freed. Exit code 0; unitrace can write its table.
- Dockerfile.unitrace (pti-gpu unitrace in the dev image) + sycl/rank_kernels.py: per-kernel device time.
- Q6KTraits: the column-independent unpack (shifts, masks, saturating byte subtract) moved into load();
  acts()/dot() split; a row-blocked kernel (native_mmvq_rowwarp_kernel, STRATA_MMVQ_RPW) kept but off:
  with a 300 ms warm-up the shared kernel already does 2560x2560x4 in 48 us (110 GB/s) and the variants tie.
- mmvq_bench: the dense kernel timing harness (warm-up + 400 iterations; a 5 ms run measures the clock ramp).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…a thread-safe 512-slot keyed cache

The prompt path collects hundreds of blob pointers per layer and copies from them on other threads later;
the 8-slot ring handed back pointers that were overwritten by then, and its index was not thread-safe.
Slots are now keyed by (layer, expert), reused oldest-first among 512 (~1.4 GB), under a mutex.

Long prompts also hung in the prompt path's streamed-expert handshake (lent cache slots make experts
non-resident, the host supplies them behind a flag the GPU wait kernel does not see). --no-prefill-borrow
keeps every expert resident and removes the handshake: 2,000 tokens 693 tok/s, 2,185 tokens 560 tok/s,
300 tokens 180 tok/s (llama.cpp on the same card: 150 short, 424 at 105k).

Rig: watchdog reports and kills a hung run at the expected time (EXPECT_S/QUIET_S); --spec 2 default
(20.75 tok/s decode vs 16.95 at spec 4: the suffix drafter accepts 9%).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…rrectness check; bandwidth probe

bw probe (incompressible data): 4-byte loads stream at 395 GB/s, 16-byte at 596. The Q6_K matvec sits at
130-160 GB/s in every variant tried (lanes per row, rows per warp, unroll, 16-byte loads, software
pipelining): latency-bound on a dependent load->unpack->dp4a chain, not ALU or bandwidth. The 16-byte
kernel is kept for n_out >= 4096 (+13% there, slower on small shapes); rel 5e-8 against the shared one.
Decode 20.80 tok/s (spec 2). The drafter is the lever now: MTP fetch + pack running.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
… vs 138-319 tok/s, decode 20.8 vs 24-26

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…cepted) - 1.8x llama.cpp on the same card

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…75 tok/s in the same run)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…), AOT divide rounding, hybrid K8V4 KV

The merge of upstream main (a790805) left 28 migrated copies behind. They are refreshed by re-migration:
dpct over the old and the merged tree, sycl/tools/normalize.sh (dpct's names, unchanged files, message
serials, fixups.py) on both, then a 3-way merge per file. Copies dpct never produced (verify.cpp, mtp.cpp)
take upstream's K8V4 diff by hand. The PTX gate fixup follows upstream's new __HIPCC__ guard (#elif).

The AOT build never received -cl-fp32-correctly-rounded-divide-sqrt (it went to the spir64 backend, which
an AOT build does not have): quantize_act_parity had its 303k mismatches back. The option now reaches
ocloc through -device "... -options ...". 19 of 22 parity tests pass again; kv_hybrid_parity (new) passes
all but its tensor-core step and is built, not registered.

B70, merged engine, AOT, --stream-experts --no-prefill-borrow --spec 4 --mtp: 2,184-token prompt 566.5
tok/s (was 475), decode 38.3 tok/s at that context (was 42.7; 77% draft acceptance on this text), MTP
prompt pass 38.6 ms (E-9 batched).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…k/s decode, 566 tok/s prompt)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…r-store dequant, XMX kernels (opt-in), profiles

Measured on the B70, outputs identical before and after each change (docs/INTEL.md "Speed work"):
- every window graph and the drafter's graphs captured at load (STRATA_WARM_GRAPHS=0 restores first-use capture)
- the commit graph is left running while the drafter's round runs on its own queue: decode 38.3 -> 43.1 tok/s at
  the 2,184-token context, 46.5 -> 49.7 on a short prompt
- the draft layer's batched prompt pass takes prompts under 64 rows (106 -> 16 ms on 19 tokens)
- one device module per kernel: the first launch in the prompt path 245 -> 1 ms
- the expert dequant writes each thread's run as one vector store: 0.085 -> 0.030 ms per expert; the 2,184-token
  prompt 566 -> 682 tok/s, 8,000 tokens 720 -> 841
- SWAR sign compare/subtract in the expert dots, a local-memory resident-plan kernel, a split-K fused down kernel
  (bit-identical to the single-token kernel, gr_parity): parity-clean, no measurable decode change

XMX: oneMKL's FP16 GEMMs already run on the XMX units (30-60 TFLOP/s). Two joint_matrix kernels are in the tree,
correct and opt-in because both lose to the paths they would replace on this card: xmx_gemm_iq (fused dequant +
GEMM from the quantized rows; xmx_gemm_bench) and qsa_prompt_attn_xmx (the mma.sync prompt attention's port;
qsa_prompt_attn_parity passes at 1e-6 of scale; STRATA_PROMPT_ATTN_XMX=1).

Profiling: unitrace -q needs Intel's metrics libraries (Dockerfile.metrics) and dev.xe.observation_paranoid=0;
the per-kernel table is in INTEL.md. A trap recorded there: --prefill-until with a native pack drops the prompt
tail (the token loop is skipped), so a "windows for short prompts" routing was invalid and is not in this commit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…t timers around the prompt path's per-layer grouping

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…(2,184 tokens 711 -> 765 tok/s, 8,000 857 -> 986)

The rows of a 2,184-token prompt are 27k random 4 KB O_DIRECT reads (114 MB for ~7 MB of rows), all before the
first kernel: 466 ms with the reader's 16 blocking threads. 64 threads (the port's default now, a copy of
src/platform/direct_file.cpp; STRATA_IO_THREADS still overrides) give 330 ms and 128 / 256 the same - the drive
tops out near 85k IOPS. A 256-token first chunk (STRATA_PREFILL_FIRST, 0: off) lets the GPU start while the
rest of the rows are read behind it: 2,184 tokens 3,075 -> 2,860 ms, time to first token ~150 ms less; 8,000
tokens 986 tok/s. The chunk boundary moves rounding in the expert GEMMs, so a long greedy continuation can
diverge late (token 45 of 64 on the test prompt); short prompts are bit-identical.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…streaming per chunk, where the 77 s go

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…he SSD is the floor of the streamed experts

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…mpt path (stale ring pointers, stream-all hang)

- Re-migration merge of upstream a790805..d6708a4 (GDN recurrence pipelining, block scores read once per window,
  the GPU-wide sampler, the draft head's VRAM reserve, correctness fixes); verify.cpp/mtp.cpp ported by hand; two
  new fixups (a cudaStream_t cast in the sampler, a const sycl::free in native_head). Parity unchanged (19 of 21).
  B70: 2,184-token prompt 765 -> 792 tok/s, decode at that context 43.1 -> 45.2 tok/s, outputs identical.
- The prompt path's stream plan kept GgufExpertSource::blob() pointers for ~1,900 experts per chunk against a
  512-slot ring, so blobs were overwritten before they were copied (the first 80k measurement was partly wrong
  experts). The stager threads now read (layer, expert) themselves via GgufExpertSource::read_into.
- The stream-all walk hangs in its first large chunk on this card (copy engine stuck on a barrier, any ring size,
  either issuer; the old read path too). The port defaults to the per-layer routed-only walk (upstream cf0d12c's
  STRATA_PREFILL_RING=8, included); STRATA_PREFILL_STREAM_ALL=1 restores it for debugging.
- The PLE reader copy keeps its 64-deep queue. 80,000 tokens now complete: 101 s = 790 tok/s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…0 tokens at 1,062 tok/s, decode 35-40 after

Without borrowing, the VRAM reserve for the KV state and the chunk buffers evicts ~1,900 experts, the prompt
streams them (790 tok/s at 80,000 tokens) and decode after it runs at 2 tok/s. Borrowing lends ~940 slots and
refills them in ~1 s: 1,062 tok/s and 35-40 tok/s decode. It costs ~1 s per prompt at short contexts (2,184
tokens 610 vs 792 tok/s), so the default is by context: borrow above 32K. --prefill-borrow / --no-prefill-borrow
decide explicitly; layer splits keep their own buffers as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
Niko1221 and others added 25 commits October 4, 2026 02:13
…X2, what was measured

INSTALL.md gets the section (the build setup compiles, what differs, the forced-path numbers from a
Ryzen 5 7600 + RTX 5070, what is untested); the CPU rows of INSTALL.md / DETAILS.md and AI_SETUP.md
point to it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…h; docs: BATCHING.md for one GPU, measured on the RTX 5070

server: a request decoding in a slot whose neighbours finished (no other slot busy, nobody waiting, at least 32
tokens allowed) is BSTOPped and continued on the solo path with its MTP drafts - the engine copies the slot's
sessions back (INFO slot_cache=1; at most twice per request; STRATA_PARALLEL_SOLO=0 keeps it in the slot).
engine: INFO batch_slots=N slot_cache=0|1.
tools/batch_interleave_test.py: the back-to-solo path, and a next turn from a slot's turn checkpoint (a client
that drops the reply's thinking).
docs: "parallel": N, what a slot costs and what setup recommends, the server's handling (interleaved prompt reads,
BYIELD, slot conversation cache, back to solo), exactness and its settings, the RTX 5070 measurements (the slots
buy waiting time, not speed, on a 12 GB card), the protocol; DETAILS / README / AI_SETUP mention the option.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…er sends a lone slot request back to the solo path only when the slots keep their conversations)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…s own slot count

The ring was 384 slots - a number measured on Q2_0, where a slot is one
expert blob of 1,382,400 B, so 506 MiB.  A slot is a whole blob, so on a
pack with bigger blobs the same 384 slots are a different amount of
memory: 1,912 MiB on Q8_0, more than half of an 8 GB card's expert cache,
and what kept that card at a 1024-token chunk when its own buffers would
have fitted 6144.  Measured on the 4-way rig, one env var and nothing
else: chunk 1024 -> 6144, prefill 87 -> 402 tok/s.

So the budget is the constant and the slot count is derived from the
pack's max blob: 384 on Q2_0 (pinned) or 96 (not), exactly as before, and
the pack it was tuned on does not move.  The fused rule (1024 slots, Niko1221#136)
is preserved the same way, and `ring_cap()` still bounds the result.  This
is also what lets the auto chunk scan treat "the ring is full" as a plain
byte comparison.

Also exports the three accessors the scan needs: bytes_needed_no_ring
(the chunk's own buffers, without the ring), ring_max_slots (the cap in
slots) and ring_slots_for (what the ring resolves to for a chunk, after
the override, the env var and the pinned-share rule - what `init` lays out
and what the INFO line reports).

And one fix on the way: `bytes_needed` counted `carve`'s buffers from
memory and got two wrong - it took `xn` unconditionally, though `carve`
takes it only under STRATA_GR_UNFUSED, and it never took `grs`.  Net
T*(D-HC)*4 bytes too many: 42 MB at a 1024-token chunk.  The direction is
safe - the prompt path was told it had less room than it did - but it
under-sizes every loan.

Co-Authored-By: Claude Code <noreply@anthropic.com>
…point exactness check

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…gh to overlap the others on a busy PC

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…y fit in VRAM

On a card whose experts mostly run on the CPU the slots cost speed per request (RTX 5070, Q2_0: 11-24 %), so setup
now recommends slots only when the expert cache (every card of a split counted) still holds half of the model's
experts beside them (Q2_0: 24 GB and up, or a split); elsewhere it stays at one at a time and says "parallel N
reduces waiting for several users but costs about 10-25% speed per request on this card". --parallel N is
honoured as asked, as before. BATCHING.md: the table with the change against one at a time.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ted count opt-in (STRATA_RING_BYTES=1)

Default: 0.1.39's 384/1024-slot ring and bytes_needed count, so the default loan and output stay as they were.
The auto chunk scan (the PR's second commit) is not taken.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…the better buy

`--prefill auto` walked a fixed list of chunk sizes - 32768, 16384, 8192,
6144, ... - and took the first that fitted, with the ring a fixed 384 slots
taken off the top before the chunk was ever considered.  Both halves of
that were wrong on a rig like this one.

The list is coarse exactly where it matters.  `bytes_needed` is a sum of
(T x positive constant) terms plus a max of such sums, so it rises
monotonically with T, and so does every test the scan applies - which makes
the largest chunk that fits a bisection on the 256-token grid the prompt
path already works on.  Seven probes against the list's ten, at one
`bytes_needed` each.  It matters because this rig affords ~8,700 tokens and
was handed 8192; 8704 is not on the list.

The order was wrong because chunk and ring spend the same borrowed VRAM,
and a ring slot is worth far more than a chunk token.  Measured here (2x
RTX 3060 + 2x RTX 5060, CUDA3 lending at its 90% cap, 120K prompt): 8960
tokens with the 17-slot ring that leaves reads at 963 tok/s, 8192/130 at
1,008, 7168 with the ring full at its 199-slot byte budget at 1,000, 5632/199
at 915.  A ring slot is worth ~0.53 tok/s and a chunk token ~0.05, so the 69
slots between a full ring and 8192's 130 are worth more than the 512 chunk
tokens they cost - and once the ring IS full, a smaller chunk buys nothing.
So: the largest chunk that still leaves the ring full, and only a rig where
no chunk can afford one falls back to the old rule with a kRingMin floor.

The room is counted in BYTES, not slots.  A ring slot is max_blob, but a
cache slot holds its own layer's blob - 2.15 MiB against a 2.54 MiB max_blob
here.  In slots the ring looked 18% cheaper than it is, a chunk whose ring
fitted only after that discount was refused outright, and the scan stopped a
step short.  `bytes_from_slots` is the inverse of the existing
`slots_from_bytes`, and `cache_slots_for` factors the latter out for the
serve path's per-participant version.  Both inverses clamp at k <= 0: the
lend budget is `min(slots - 128, pct * slots / 100)` and goes negative on a
cache under 128 slots, where `slots - k` is off_[slots] - the end of the
prefix sum, so that cache lends 0 bytes rather than reading past its offsets.

Upstream's per-participant `fits`/`fits_one` and the Niko1221#448 diagnostic are kept:
the ring room is taken over the smallest participant, `alone` is the chunk
CUDA0 alone would read in, and the small-card warning is unchanged.

Verified on the rig: the scan picks 7680/199 against the old rule's 8960/17,
1018.8 and 1021.1 tok/s on two cold 120K prompts against 963.1 - +5.9%.
Q8_0 picks 5632/101, its 506 MiB budget over a 5,222,400 B blob.

Co-Authored-By: Claude Code <noreply@anthropic.com>

0.1.39b (Niko1221): behind STRATA_RING_BYTES=1 with the byte-budget ring; the default keeps 0.1.39's chunk list
(plan_lend and the serve scan), its log line gains the ring. The INFO fields (the Monitor commit) are not taken.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…mbers and quality check

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
's fused_gr carveout on HIP)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Added where a config goes to several GPUs (setup --gpus / a multi-GPU install, start --gpus, "use both" at
start). Single-GPU configs are not touched. Opt-out: setup --no-remote-expert-opt, or "remote_expert_opt": false
in the config (a user key, kept on re-runs). The engine uses it only with a helper cache (--expert-cache-device1..3),
so a layer split runs as before. tools/test_setup_remote_opt.py; the 2x16 GB golden configs gain the flag.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…S.md with the measured quality

Changes default output on long prompts of native packs against 0.1.39 (expected; gate notes in the S40 report).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…8, Linux 525); the wheel pins match the 12.9.1 toolkit

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… keep 0.1.39's ring

- a prompt that fits 0.1.39's auto chunk keeps 0.1.39's ring (Prefill::set_ring_budget's small_max): 4K prompts
  were 1-17% slower with the smaller ring (Coder -15%, IQ3_S -5%, IQ3_XXS -2%; 5 interleaved pairs each);
- ring slots are given up for a bigger chunk only where 0.1.39's rule held the chunk under 6144 (ring_cap_for);
  from 6144 on the ring keeps its 0.1.39 size (the Coder 6144/384 -> 7936/199 was -12% at 32K);
- the unpinned arm's cap stays 96 (the byte budget made it ~49 on IQ3_S);
- the scan sets its own ring budget instead of set_ring_override, so a layer split's 96-slot override (Niko1221#340) is no
  longer clobbered by the scan.
RTX 5070 vs STRATA_RING_BYTES=0: 4K -0.7..+1.1% (bit-identical); 32K IQ3_XXS +18.5% (cache 1500) / +9.4%, Coder
+3.1%, IQ3_S +0.2%.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJJwUv9g3cD2ztX6KWS8rk
# Conflicts:
#	src/core/expert_source.cpp
#	src/kernels/cuda/iq_kernels.cu
# Conflicts:
#	src/core/verify.cpp
#	src/program/generate.cpp
…l); UD-Q4_K_XL stays experimental

- "experimental" is per model now: MODELS["UD-Q4_K_XL"] keeps it, the unsloth family and UD-IQ4_XS drop it. The
  first menu lists the family without [experimental], with a one-line about (UD-IQ4_XS: a 94 GB download; under
  ~80 GB of RAM part of its experts are read from the SSD). Its size menu lists UD-IQ4_XS first (the default) and
  UD-Q4_K_XL second; an experimental size always sorts last. --model names are unchanged.
- UD-IQ4_XS: no EXPERIMENTAL warning, --check verdict without it. Images are asked (off by default), per model: the
  engine's image path has no restriction for this pack, which uses the original model's image encoder; not yet run
  with images. UD-Q4_K_XL keeps images off.
- MCP: the fallback tables mirror setup's (per-model experimental/vision), sizes in setup's order, per-size
  "experimental" and "images", the vision check per model.
- Docs: README and its six translations, MODELS, UNSLOTH_Q4, AI_SETUP, MCP_SERVER.
- Tests: the menus, the default, --model by name, images, --check, the pack/tokenizer paths; the golden configs
  are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJJwUv9g3cD2ztX6KWS8rk
…e-mib 1000 if a request stalls

With the image encoder on the GPU, the 700 MiB reserve can leave ~200 MiB free on a 12 GB card (UD-IQ4_XS
end-to-end check on an RTX 5070). Setup prints a one-line recommendation; the written config is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJJwUv9g3cD2ztX6KWS8rk
@rhgo1749
rhgo1749 marked this pull request as ready for review October 5, 2026 00:15
@rhgo1749
rhgo1749 merged commit 9ae0839 into main Oct 5, 2026
1 check passed
@rhgo1749
rhgo1749 deleted the sync/upstream-0.1.39 branch October 5, 2026 03:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.