Repository navigation
Sync upstream Strata v0.1.39 into Lanes - #24
Merged
Merged
Conversation
The Strata engine is CUDA. On a GPU without it - written for and measured on an Intel Arc Pro
B70 - the same GGUF weights run through llama.cpp's SYCL backend. serve/engine_llama.py puts
llama-server behind Engine.generate(ids, max_new, sampling, cancel), so the OpenAI and
Anthropic endpoints, streaming, tool calls, MCP and the web app are unchanged above it.
It sends the prompt as token ids (the server's own tokenizer, from the same GGUF) with
return_tokens, streams ids back with llama-server's timings mapped onto Strata's `last`,
heartbeats while a long prompt is read, and stops the server by closing the connection when a
client goes away. Attach to a running server ("llama": {"url"}) or spawn one ("llama": {"exe"}).
Not yet: images (Strata's Vision produces embeddings for the CUDA engine; llama-server takes
the image itself) - start without "vision" and image requests are refused cleanly.
Tests: a scripted llama-server (the wire format observed on b29c606), plus a live round trip
when STRATA_LLAMA_URL is set. Measured on the B70 with Coder IQ1_M: 23-25 tok/s decode,
~150 tok/s prefill, correct code, GPU-bound.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…erver itself Without nvidia-smi, setup now looks for an Intel discrete GPU in sysfs (vendor 0x8086 under xe or i915, named by PCI id - lspci's database does not know an Arc Pro B70 yet) and takes a path with no compiler and no CUDA toolkit: download the model, export the tokenizer, find a llama-server (--llama-server PATH, one on PATH, or the B70 container image via docker), write a config with a "llama" block and a start-llama.sh, and a run script that uses --engine llama. The CUDA engine's rules that do not apply are stepped around rather than deleted: the compute capability and driver-version checks, the experts-in-RAM gates (on llama.cpp the fit is VRAM, shard 1 = the download minus the 28.8 GB lookup table), the MTP draft download (5 GB the llama engine cannot use), and the engine's own argument list. engine_from_config runs the config's start script when nothing answers at the url, so the run script is one click and closing the server stops the container. Verified on an Arc Pro B70: setup runs steps 1-7 clean against the downloaded Coder IQ1_M. docs/INTEL.md has the whole picture and the measurements. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…webui here) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
About 20 KiB of KV per token (only 12 of 48 layers are full attention). 32k, 64k, 96k and 128k all load with the whole model on the card, 128k with 1.6 GB to spare; 160k is under the safety margin and 256k does not fit - asking for it spilled VRAM into host RAM and took the machine down until its watchdog reset it, so the doc says so in bold. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…4 tok/s Prompt reading scales up with prompt size (150 tok/s at 2.7k, 424 tok/s at 105k), so the full window costs about five minutes, not fifteen. 30.3 GB in use with 1.5 GB free at 131,072. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
… kernel parity tests pass on the Arc Pro B70 The port lives in sycl/ and changes no upstream file: sycl/src and sycl/include hold only the files the migration touched, sycl/CMakeLists.txt takes everything else from the original tree. How it was made, reproducibly (sycl/tools/): a dev image (the llama.cpp SYCL image + SYCLomatic 2025.3 + CUDA 12.8 headers from the pip wheels), migrate.sh (compilation database + dpct over 86 translation units), fixups.py (every hand fix with its reason, idempotent), build.sh (icpx, ninja). What dpct got wrong and the script repairs: CUDA's null stream became a null sycl::queue* (strata::q_of maps it to the default queue); __ldg((const float*) p) lost its cast and read one byte; __fadd_rn lost parentheses; ggml's lookup tables were threaded through kernel parameters with the wrong table per template; helper headers renamed since 2025.3; cudaGraphUpload, graph introspection and %globaltimer have no equivalents. Two compiler flags are load-bearing: -fp-model=precise, and the device compiler's -cl-fp32-correctly-rounded-divide-sqrt - the Arc's fp32 divide is not correctly rounded by default, and quantize_act_parity went from 303k mismatches to byte-exact with it. Not ported: the three kernels with inline PTX (mma.sync, ldmatrix, cp.async) take their existing pre-sm_80 fallback; the launchers refuse the device. The XMX joint_matrix versions are next. The engine itself starts on the card, loads the pack, the projections, the PLE table and the 23.4 GiB expert arena, and is then stopped by the test rig's memory guard: Strata keeps the experts as a resident host copy, and the rig has 23 GiB of RAM (upstream asks for 32). A streaming fill of the VRAM cache from the GGUF is the next piece; the whole model fits in the 32 GB card. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
… Pro B70 --stream-experts (GgufExpertSource): every expert goes straight from the GGUF into the VRAM cache through a small ring of staging buffers; no resident host arena, so the engine runs on a 23 GiB machine. What it took to get from "loads" to "<think>\nThe user wants a Python function": - the doorbell handshake through host-mapped memory: volatile device loads do not bypass the caches on Intel; system-scope atomics (strata::sys_load/sys_store) do. A device-side spin is bounded (kSpinMax): an unbounded one wedged the GT twice when its process died. - STRATA_VERIFY_NO_HOST=1 + STRATA_VERIFY_DEVICE_PLAN=1: with every expert resident the GPU plans each layer itself and the host waits for the whole window graph; per-layer ring visibility inside a graph is not reliable on this platform (measured with sycl/probe). - -fsycl-default-sub-group-size=32: the kernels are written for 32-lane warps; dpct pinned only 134 of 289. - cudaMemcpy is synchronous; dpct's default-queue memcpy had no wait, and the streaming ring reuses its buffers - the expert cache filled from overwritten memory (non-deterministic residuals, NaN by layer 5). - native_expert_parity hand-ported (dpct cannot parse it without ggml-cpu.h): the GPU native expert kernel matches the float reference on real IQ1_M rows. - verify.cpp debug readbacks (STRATA_VERIFY_DEBUG): the per-layer residual ladder that found all of it. Not yet: speed (0.17 tok/s - the spin budgets), the segfault at exit, the tensor-core kernels. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…, eager stage profiler, docs Measured on the B70: 64 tokens in 52.0 s, of which ~47 s is the first window (the runtime JIT-compiles every kernel on first use: 6 tokens take 47.5 s, 64 take 52.0 s). The output is the model's real answer. Debug tooling kept in the port: STRATA_VERIFY_EAGER=1 runs the window on the queue instead of as a graph and turns the stage stamps into host clocks, so STRATA_VERIFY_PROFILE=1 prints ms per stage without a device timer; STRATA_VERIFY_DEBUG=1 prints the residual after every layer and the round's phase times. build.sh takes BUILD_DIR and AOT (an ahead-of-time bmg-g31 build removes the JIT cost). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…st token in 0.85 s, no JIT The -Xs device flag was mis-quoted; fixed. Measured on the B70 with the same 19-token prompt, 32 tokens: decode 15.17 tok/s, prefill 28.79 tok/s, a 4-token window 73 ms on the card. The persistent JIT cache (SYCL_CACHE_PERSISTENT=1) is the alternative for the JIT build: 15 MB, second start 15.6 tok/s. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…er on iq3_xxs layers A window routes the same expert for up to 6 tokens; the grouped kernels recomputed the dequantised weights (grid lookups, sign unpacking, the byte-wise compare/xor/subtract emulations) once per token. Multi<TY> splits each of the five formats this model uses (iq2_s, iq3_xxs, iq3_s gate/up; iq4_nl, q2_0 down) into prep (once) and acts + finish (per token); the kernels take entries four at a time. Lanes per row is a compile-time constant (STRATA_EXPERT_LANES, 8): 32 lanes on an 80-call row was mostly reduction. Measured with the new NATIVE_BENCH=1 mode of native_expert_parity (10 experts x 4 entries, one layer): iq3_xxs/iq4_nl 1.467 -> 0.649 ms, iq2_s/q2_0 ~0.68 -> 0.579 ms, iq3_s/iq4_nl 0.750 ms; parity unchanged. strata::dp4a wraps dpct::dp4a (sycl::ext::oneapi::dot_acc is the same emulation with a header that breaks the link across translation units). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…tok/s decode route<NE> templated on the expert count; native_router_top10_multi_ne dispatches 256/512. Both call sites (the verify window and the per-token moe_route) take it for the Coder. 64 tokens: 16.25 -> 16.95 tok/s, same output. Also: Dockerfile.unitrace (per-kernel Level Zero profiling image). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…: the mmvq kernels run at 110-130 GB/s warm - exit segfault: dpct's static global_memory/constant_memory objects destructed after the default queue (queue::get_device in the core); heap-allocated, never freed. Exit code 0; unitrace can write its table. - Dockerfile.unitrace (pti-gpu unitrace in the dev image) + sycl/rank_kernels.py: per-kernel device time. - Q6KTraits: the column-independent unpack (shifts, masks, saturating byte subtract) moved into load(); acts()/dot() split; a row-blocked kernel (native_mmvq_rowwarp_kernel, STRATA_MMVQ_RPW) kept but off: with a 300 ms warm-up the shared kernel already does 2560x2560x4 in 48 us (110 GB/s) and the variants tie. - mmvq_bench: the dense kernel timing harness (warm-up + 400 iterations; a 5 ms run measures the clock ramp). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…a thread-safe 512-slot keyed cache The prompt path collects hundreds of blob pointers per layer and copies from them on other threads later; the 8-slot ring handed back pointers that were overwritten by then, and its index was not thread-safe. Slots are now keyed by (layer, expert), reused oldest-first among 512 (~1.4 GB), under a mutex. Long prompts also hung in the prompt path's streamed-expert handshake (lent cache slots make experts non-resident, the host supplies them behind a flag the GPU wait kernel does not see). --no-prefill-borrow keeps every expert resident and removes the handshake: 2,000 tokens 693 tok/s, 2,185 tokens 560 tok/s, 300 tokens 180 tok/s (llama.cpp on the same card: 150 short, 424 at 105k). Rig: watchdog reports and kills a hung run at the expected time (EXPECT_S/QUIET_S); --spec 2 default (20.75 tok/s decode vs 16.95 at spec 4: the suffix drafter accepts 9%). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…rrectness check; bandwidth probe bw probe (incompressible data): 4-byte loads stream at 395 GB/s, 16-byte at 596. The Q6_K matvec sits at 130-160 GB/s in every variant tried (lanes per row, rows per warp, unroll, 16-byte loads, software pipelining): latency-bound on a dependent load->unpack->dp4a chain, not ALU or bandwidth. The 16-byte kernel is kept for n_out >= 4096 (+13% there, slower on small shapes); rel 5e-8 against the shared one. Decode 20.80 tok/s (spec 2). The drafter is the lever now: MTP fetch + pack running. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
… vs 138-319 tok/s, decode 20.8 vs 24-26 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…cepted) - 1.8x llama.cpp on the same card Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…75 tok/s in the same run) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
# Conflicts: # setup.py
…), AOT divide rounding, hybrid K8V4 KV The merge of upstream main (a790805) left 28 migrated copies behind. They are refreshed by re-migration: dpct over the old and the merged tree, sycl/tools/normalize.sh (dpct's names, unchanged files, message serials, fixups.py) on both, then a 3-way merge per file. Copies dpct never produced (verify.cpp, mtp.cpp) take upstream's K8V4 diff by hand. The PTX gate fixup follows upstream's new __HIPCC__ guard (#elif). The AOT build never received -cl-fp32-correctly-rounded-divide-sqrt (it went to the spir64 backend, which an AOT build does not have): quantize_act_parity had its 303k mismatches back. The option now reaches ocloc through -device "... -options ...". 19 of 22 parity tests pass again; kv_hybrid_parity (new) passes all but its tensor-core step and is built, not registered. B70, merged engine, AOT, --stream-experts --no-prefill-borrow --spec 4 --mtp: 2,184-token prompt 566.5 tok/s (was 475), decode 38.3 tok/s at that context (was 42.7; 77% draft acceptance on this text), MTP prompt pass 38.6 ms (E-9 batched). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…k/s decode, 566 tok/s prompt) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…r-store dequant, XMX kernels (opt-in), profiles Measured on the B70, outputs identical before and after each change (docs/INTEL.md "Speed work"): - every window graph and the drafter's graphs captured at load (STRATA_WARM_GRAPHS=0 restores first-use capture) - the commit graph is left running while the drafter's round runs on its own queue: decode 38.3 -> 43.1 tok/s at the 2,184-token context, 46.5 -> 49.7 on a short prompt - the draft layer's batched prompt pass takes prompts under 64 rows (106 -> 16 ms on 19 tokens) - one device module per kernel: the first launch in the prompt path 245 -> 1 ms - the expert dequant writes each thread's run as one vector store: 0.085 -> 0.030 ms per expert; the 2,184-token prompt 566 -> 682 tok/s, 8,000 tokens 720 -> 841 - SWAR sign compare/subtract in the expert dots, a local-memory resident-plan kernel, a split-K fused down kernel (bit-identical to the single-token kernel, gr_parity): parity-clean, no measurable decode change XMX: oneMKL's FP16 GEMMs already run on the XMX units (30-60 TFLOP/s). Two joint_matrix kernels are in the tree, correct and opt-in because both lose to the paths they would replace on this card: xmx_gemm_iq (fused dequant + GEMM from the quantized rows; xmx_gemm_bench) and qsa_prompt_attn_xmx (the mma.sync prompt attention's port; qsa_prompt_attn_parity passes at 1e-6 of scale; STRATA_PROMPT_ATTN_XMX=1). Profiling: unitrace -q needs Intel's metrics libraries (Dockerfile.metrics) and dev.xe.observation_paranoid=0; the per-kernel table is in INTEL.md. A trap recorded there: --prefill-until with a native pack drops the prompt tail (the token loop is skipped), so a "windows for short prompts" routing was invalid and is not in this commit. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…t timers around the prompt path's per-layer grouping Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…(2,184 tokens 711 -> 765 tok/s, 8,000 857 -> 986) The rows of a 2,184-token prompt are 27k random 4 KB O_DIRECT reads (114 MB for ~7 MB of rows), all before the first kernel: 466 ms with the reader's 16 blocking threads. 64 threads (the port's default now, a copy of src/platform/direct_file.cpp; STRATA_IO_THREADS still overrides) give 330 ms and 128 / 256 the same - the drive tops out near 85k IOPS. A 256-token first chunk (STRATA_PREFILL_FIRST, 0: off) lets the GPU start while the rest of the rows are read behind it: 2,184 tokens 3,075 -> 2,860 ms, time to first token ~150 ms less; 8,000 tokens 986 tok/s. The chunk boundary moves rounding in the expert GEMMs, so a long greedy continuation can diverge late (token 45 of 64 on the test prompt); short prompts are bit-identical. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…streaming per chunk, where the 77 s go Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…he SSD is the floor of the streamed experts Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
# Conflicts: # setup.py
…mpt path (stale ring pointers, stream-all hang) - Re-migration merge of upstream a790805..d6708a4 (GDN recurrence pipelining, block scores read once per window, the GPU-wide sampler, the draft head's VRAM reserve, correctness fixes); verify.cpp/mtp.cpp ported by hand; two new fixups (a cudaStream_t cast in the sampler, a const sycl::free in native_head). Parity unchanged (19 of 21). B70: 2,184-token prompt 765 -> 792 tok/s, decode at that context 43.1 -> 45.2 tok/s, outputs identical. - The prompt path's stream plan kept GgufExpertSource::blob() pointers for ~1,900 experts per chunk against a 512-slot ring, so blobs were overwritten before they were copied (the first 80k measurement was partly wrong experts). The stager threads now read (layer, expert) themselves via GgufExpertSource::read_into. - The stream-all walk hangs in its first large chunk on this card (copy engine stuck on a barrier, any ring size, either issuer; the old read path too). The port defaults to the per-layer routed-only walk (upstream cf0d12c's STRATA_PREFILL_RING=8, included); STRATA_PREFILL_STREAM_ALL=1 restores it for debugging. - The PLE reader copy keeps its 64-deep queue. 80,000 tokens now complete: 101 s = 790 tok/s. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…0 tokens at 1,062 tok/s, decode 35-40 after Without borrowing, the VRAM reserve for the KV state and the chunk buffers evicts ~1,900 experts, the prompt streams them (790 tok/s at 80,000 tokens) and decode after it runs at 2 tok/s. Borrowing lends ~940 slots and refills them in ~1 s: 1,062 tok/s and 35-40 tok/s decode. It costs ~1 s per prompt at short contexts (2,184 tokens 610 vs 792 tok/s), so the default is by context: borrow above 32K. --prefill-borrow / --no-prefill-borrow decide explicitly; layer splits keep their own buffers as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Rb4SRxFuLpffQAAXBf4HxA
…X2, what was measured INSTALL.md gets the section (the build setup compiles, what differs, the forced-path numbers from a Ryzen 5 7600 + RTX 5070, what is untested); the CPU rows of INSTALL.md / DETAILS.md and AI_SETUP.md point to it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…h; docs: BATCHING.md for one GPU, measured on the RTX 5070 server: a request decoding in a slot whose neighbours finished (no other slot busy, nobody waiting, at least 32 tokens allowed) is BSTOPped and continued on the solo path with its MTP drafts - the engine copies the slot's sessions back (INFO slot_cache=1; at most twice per request; STRATA_PARALLEL_SOLO=0 keeps it in the slot). engine: INFO batch_slots=N slot_cache=0|1. tools/batch_interleave_test.py: the back-to-solo path, and a next turn from a slot's turn checkpoint (a client that drops the reply's thinking). docs: "parallel": N, what a slot costs and what setup recommends, the server's handling (interleaved prompt reads, BYIELD, slot conversation cache, back to solo), exactness and its settings, the RTX 5070 measurements (the slots buy waiting time, not speed, on a 12 GB card), the protocol; DETAILS / README / AI_SETUP mention the option. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…er sends a lone slot request back to the solo path only when the slots keep their conversations) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…s own slot count The ring was 384 slots - a number measured on Q2_0, where a slot is one expert blob of 1,382,400 B, so 506 MiB. A slot is a whole blob, so on a pack with bigger blobs the same 384 slots are a different amount of memory: 1,912 MiB on Q8_0, more than half of an 8 GB card's expert cache, and what kept that card at a 1024-token chunk when its own buffers would have fitted 6144. Measured on the 4-way rig, one env var and nothing else: chunk 1024 -> 6144, prefill 87 -> 402 tok/s. So the budget is the constant and the slot count is derived from the pack's max blob: 384 on Q2_0 (pinned) or 96 (not), exactly as before, and the pack it was tuned on does not move. The fused rule (1024 slots, Niko1221#136) is preserved the same way, and `ring_cap()` still bounds the result. This is also what lets the auto chunk scan treat "the ring is full" as a plain byte comparison. Also exports the three accessors the scan needs: bytes_needed_no_ring (the chunk's own buffers, without the ring), ring_max_slots (the cap in slots) and ring_slots_for (what the ring resolves to for a chunk, after the override, the env var and the pinned-share rule - what `init` lays out and what the INFO line reports). And one fix on the way: `bytes_needed` counted `carve`'s buffers from memory and got two wrong - it took `xn` unconditionally, though `carve` takes it only under STRATA_GR_UNFUSED, and it never took `grs`. Net T*(D-HC)*4 bytes too many: 42 MB at a 1024-token chunk. The direction is safe - the prompt path was told it had less room than it did - but it under-sizes every loan. Co-Authored-By: Claude Code <noreply@anthropic.com>
…point exactness check Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…gh to overlap the others on a busy PC Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…y fit in VRAM On a card whose experts mostly run on the CPU the slots cost speed per request (RTX 5070, Q2_0: 11-24 %), so setup now recommends slots only when the expert cache (every card of a split counted) still holds half of the model's experts beside them (Q2_0: 24 GB and up, or a split); elsewhere it stays at one at a time and says "parallel N reduces waiting for several users but costs about 10-25% speed per request on this card". --parallel N is honoured as asked, as before. BATCHING.md: the table with the change against one at a time. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ted count opt-in (STRATA_RING_BYTES=1) Default: 0.1.39's 384/1024-slot ring and bytes_needed count, so the default loan and output stay as they were. The auto chunk scan (the PR's second commit) is not taken. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…the better buy `--prefill auto` walked a fixed list of chunk sizes - 32768, 16384, 8192, 6144, ... - and took the first that fitted, with the ring a fixed 384 slots taken off the top before the chunk was ever considered. Both halves of that were wrong on a rig like this one. The list is coarse exactly where it matters. `bytes_needed` is a sum of (T x positive constant) terms plus a max of such sums, so it rises monotonically with T, and so does every test the scan applies - which makes the largest chunk that fits a bisection on the 256-token grid the prompt path already works on. Seven probes against the list's ten, at one `bytes_needed` each. It matters because this rig affords ~8,700 tokens and was handed 8192; 8704 is not on the list. The order was wrong because chunk and ring spend the same borrowed VRAM, and a ring slot is worth far more than a chunk token. Measured here (2x RTX 3060 + 2x RTX 5060, CUDA3 lending at its 90% cap, 120K prompt): 8960 tokens with the 17-slot ring that leaves reads at 963 tok/s, 8192/130 at 1,008, 7168 with the ring full at its 199-slot byte budget at 1,000, 5632/199 at 915. A ring slot is worth ~0.53 tok/s and a chunk token ~0.05, so the 69 slots between a full ring and 8192's 130 are worth more than the 512 chunk tokens they cost - and once the ring IS full, a smaller chunk buys nothing. So: the largest chunk that still leaves the ring full, and only a rig where no chunk can afford one falls back to the old rule with a kRingMin floor. The room is counted in BYTES, not slots. A ring slot is max_blob, but a cache slot holds its own layer's blob - 2.15 MiB against a 2.54 MiB max_blob here. In slots the ring looked 18% cheaper than it is, a chunk whose ring fitted only after that discount was refused outright, and the scan stopped a step short. `bytes_from_slots` is the inverse of the existing `slots_from_bytes`, and `cache_slots_for` factors the latter out for the serve path's per-participant version. Both inverses clamp at k <= 0: the lend budget is `min(slots - 128, pct * slots / 100)` and goes negative on a cache under 128 slots, where `slots - k` is off_[slots] - the end of the prefix sum, so that cache lends 0 bytes rather than reading past its offsets. Upstream's per-participant `fits`/`fits_one` and the Niko1221#448 diagnostic are kept: the ring room is taken over the smallest participant, `alone` is the chunk CUDA0 alone would read in, and the small-card warning is unchanged. Verified on the rig: the scan picks 7680/199 against the old rule's 8960/17, 1018.8 and 1021.1 tok/s on two cold 120K prompts against 963.1 - +5.9%. Q8_0 picks 5632/101, its 506 MiB budget over a 5,222,400 B blob. Co-Authored-By: Claude Code <noreply@anthropic.com> 0.1.39b (Niko1221): behind STRATA_RING_BYTES=1 with the byte-budget ring; the default keeps 0.1.39's chunk list (plan_lend and the serve scan), its log line gains the ring. The INFO fields (the Monitor commit) are not taken. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…mbers and quality check Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Added where a config goes to several GPUs (setup --gpus / a multi-GPU install, start --gpus, "use both" at start). Single-GPU configs are not touched. Opt-out: setup --no-remote-expert-opt, or "remote_expert_opt": false in the config (a user key, kept on re-runs). The engine uses it only with a helper cache (--expert-cache-device1..3), so a layer split runs as before. tools/test_setup_remote_opt.py; the 2x16 GB golden configs gain the flag. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…S.md with the measured quality Changes default output on long prompts of native packs against 0.1.39 (expected; gate notes in the S40 report). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…8, Linux 525); the wheel pins match the 12.9.1 toolkit Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… keep 0.1.39's ring - a prompt that fits 0.1.39's auto chunk keeps 0.1.39's ring (Prefill::set_ring_budget's small_max): 4K prompts were 1-17% slower with the smaller ring (Coder -15%, IQ3_S -5%, IQ3_XXS -2%; 5 interleaved pairs each); - ring slots are given up for a bigger chunk only where 0.1.39's rule held the chunk under 6144 (ring_cap_for); from 6144 on the ring keeps its 0.1.39 size (the Coder 6144/384 -> 7936/199 was -12% at 32K); - the unpinned arm's cap stays 96 (the byte budget made it ~49 on IQ3_S); - the scan sets its own ring budget instead of set_ring_override, so a layer split's 96-slot override (Niko1221#340) is no longer clobbered by the scan. RTX 5070 vs STRATA_RING_BYTES=0: 4K -0.7..+1.1% (bit-identical); 32K IQ3_XXS +18.5% (cache 1500) / +9.4%, Coder +3.1%, IQ3_S +0.2%. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJJwUv9g3cD2ztX6KWS8rk
# Conflicts: # setup.py
# Conflicts: # src/core/expert_source.cpp # src/kernels/cuda/iq_kernels.cu
# Conflicts: # src/core/verify.cpp # src/program/generate.cpp
…l); UD-Q4_K_XL stays experimental - "experimental" is per model now: MODELS["UD-Q4_K_XL"] keeps it, the unsloth family and UD-IQ4_XS drop it. The first menu lists the family without [experimental], with a one-line about (UD-IQ4_XS: a 94 GB download; under ~80 GB of RAM part of its experts are read from the SSD). Its size menu lists UD-IQ4_XS first (the default) and UD-Q4_K_XL second; an experimental size always sorts last. --model names are unchanged. - UD-IQ4_XS: no EXPERIMENTAL warning, --check verdict without it. Images are asked (off by default), per model: the engine's image path has no restriction for this pack, which uses the original model's image encoder; not yet run with images. UD-Q4_K_XL keeps images off. - MCP: the fallback tables mirror setup's (per-model experimental/vision), sizes in setup's order, per-size "experimental" and "images", the vision check per model. - Docs: README and its six translations, MODELS, UNSLOTH_Q4, AI_SETUP, MCP_SERVER. - Tests: the menus, the default, --model by name, images, --check, the pack/tokenizer paths; the golden configs are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJJwUv9g3cD2ztX6KWS8rk
…e-mib 1000 if a request stalls With the image encoder on the GPU, the 700 MiB reserve can leave ~200 MiB free on a 12 GB card (UD-IQ4_XS end-to-end check on an RTX 5070). Setup prints a one-line recommendation; the written config is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJJwUv9g3cD2ztX6KWS8rk
rhgo1749
marked this pull request as ready for review
October 5, 2026 00:15
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Merge upstream Strata v0.1.39 (
6f32ec070f23ced9f50e704d854d775da52591ab) onto the Lanes 0.1.38 baseline while preserving the independent whole-GPU lane runtime.Key integration decisions:
server.pyand native conversation-cache telemetry, avoiding the old LanesDONEfield extension that collided with upstream's offloaded-expert counter;tool_choice/ malformed tool-call handling while retaining upstream frontend/reasoning changes;parallel,--batch,--slots,--batch-groups, and--trim-stage-weightsfrom ordinary lanes so upstream internal batching cannot silently change the Lanes concurrency unit;/v1/responsesthrough the normal Lanes generation admission path.Validation
Reference host, CUDA 13.4.92, sm_120:
python3 -m unittest discover -s serve -p 'test_*.py': 333 passed, 7 skipped.stratabuild withSTRATA_BUILD_TESTS=OFF: pass.stratabuild withSTRATA_BUILD_TESTS=ON: pass.pinned_shared_test,expert_profile_save_test,file_expert_source_test: 3/3 pass.shared population ready; source load skipped.POST /v1/responsesthrough the Lanes supervisor: HTTP 200,X-Strata-Lane-Index: 0, output textOK.The first test-enabled build attempt showed an unresolved
cuGetErrorString, but this was traced to a CMake directory created by an initial configure with no CUDA compiler: its cachedCMAKE_CUDA_FLAGS_RELEASEwas empty. A fresh configure with/usr/local/cuda-13.4/bin/nvccfrom the first invocation restored-O3 -DNDEBUGand the full test-enabled executable linked successfully.The compatibility/promotion gate above remains distinct from historical 0.1.30/0.1.31/0.1.38 benchmark evidence; those measurements keep their original engine labels.
0.1.39 topology recheck
After the compatibility gate, the topology comparison was reduced to the decision-bearing points and rerun only where evidence was missing.
Fixed decode scorecard:
18,34pipeline with--batch 3 --batch-groups 3 --trim-stage-weightsand shared expert arena 147.84 ± 2.26 tok/s (+3.25%);Cold-prefill crossover:
All retained PP requests report
cache_n=0. The 15K/110K nonce text is workload/length matched rather than byte-identical across topology arms. The ordinary FIFO and batch-only runs are retained only as controls; broad oversubscription, heterogeneity, workload-sensitivity and PSS probes remain raw/supporting evidence rather than headline topology benchmarks.This reopens the always-on three-GPU layer-split server as a serious default-topology candidate, but production promotion still depends on mixed prompt/output lengths, session behavior, queue/tail latency, power, failure isolation and operational flexibility.