Faster on more than one GPU, faster short prompts on NVIDIA, about twice as fast prompts on Windows with little RAM, and a long list of fixes. Two defaults change: --batch on a layer split now runs one pipeline group per GPU, and short prompt chunks on one NVIDIA GPU get help from the CPU (this one changes the last bits of an answer; one setting turns it off). Update with UPDATE.bat (Linux: ./update.sh); setup replaces the engine with 0.1.41. This release includes the Pascal decode fix from 0.1.40.4.
Speed
Against 0.1.40.4, default settings, interleaved pairs, medians (prompt = time to read it; higher is better in every column):
| Machine, model | Prompt 512 / 1K tokens | Prompt 4K / 20K | Decode |
|---|---|---|---|
| RTX 3060 12 GB, Linux, IQ3_XXS | +28% / +22% | equal | equal |
| Tesla P100 16 GB, Linux, IQ3_XXS | +33% / +25% | equal | equal |
| RTX 5070 12 GB, Windows, Q2_0 | equal* | equal | +0.4% (story +4.9%) |
| RTX 5070, Windows, IQ3_XXS / IQ3_S with a RAM budget | +16% / +23-25% | +3% | |
| 1x Radeon AI PRO R9700, Linux, Q2_0 | equal | equal | equal |
| Arc Pro B70 32 GB, Linux, IQ3_S | +13% / +12% | +10% / +10% | equal |
| Arc A750 8 GB, Linux, IQ3_XXS | equal | equal | equal |
* On the RTX 5070 with Q2_0, setup's automatic expert cache already holds about 3,700 experts on the card, so the CPU has almost nothing to take during a short prompt; with a smaller cache (--expert-cache 1500) the same prompts are 15-27% faster. Answers in every row are identical to 0.1.40.4 with STRATA_PREFILL_CPU_SHARE=0 (Q2_0, IQ3_XXS, Coder and IQ3_S on the 5070, including 4K/20K-token prompts; IQ3_XXS on the 3060 and P100; Q2_0 on the R9700; the 10-prompt check on both Intel cards).
-
Several clients on several GPUs (default;
--batchwith a layer split).--batch-groups autois now the default: every GPU stage works on its own group of requests at the same time, instead of one request group walking through the stages while the other cards wait. Total tokens per second at 8 clients: +32% to +68% (2, 3 and 4 GPUs; 4x Radeon AI PRO R9700 in the larger cases), with a better p95 as well. The answers are the same as with one group (the 8-request identity test passes on 2, 3 and 4 GPUs).--batch-groups 1turns it off. The first version of this slowed one long request that ran next to short ones by 39-47%; that is fixed and measured. The server also listens with a backlog of 256 now (STRATA_HTTP_BACKLOG): with 30 to 40 clients opening connections at once, some got "connection reset"; at 40 clients there are now 0 resets (134 tok/s in total on 4x R9700). The engine still serves 8 slots at a time (--batchabove 8 is capped), so with 16 or more clients the rest queue.4x R9700, IQ3_S, --batch 80.1.40.4 0.1.41 8 clients, total tok/s 88 165 (+87%) 8 clients, finish p50 17.4 s 9.3 s 40 clients, total tok/s 88 153 (+75%) 40 clients, connection resets per run 14-21 0 -
Short prompts up to a third faster on one NVIDIA GPU (default, changes bits). The CPU, idle while a prompt is read, takes part of the experts the GPU would otherwise stream, for chunks below 1,024 tokens (an agent's tool result, a short follow-up). It helps as much as the GPU has to stream: 0.1.41 against 0.1.40.4 reads a 512-token prompt 28% faster on an RTX 3060 and 33% faster on a Tesla P100 (1,000 tokens: 22% and 25%). On an RTX 5070 with Q2_0, setup's automatic cache already holds most of what a short prompt needs and the gain is 0-2%; with a smaller cache it is 15-27%. It is never slower in our runs. Longer prompts are unchanged. The answers can differ slightly from 0.1.40.4: first-token KL against the share off averages 0.004 (max 0.025).
STRATA_PREFILL_CPU_SHARE=0turns it off and gives the exact 0.1.40.x answers. Only one NVIDIA GPU without--batchslots; layer splits, batch and AMD are as before. Thanks to sergqwer, whose measured on/off check this is. -
Windows, model bigger than RAM: prompts about twice as fast (#1323). The prompt reader asks the drive for several experts in one request, and gives up the memory-mapped view of
experts.binonce its reads bypass the cache (while it is mapped, Windows serves those reads one at a time). On the RTX 5070 PC the prompt time fell 45-52%, with the same tokens in 18 of 18 runs; in the release check, IQ3_XXS and IQ3_S with a RAM budget read 4K-20K prompts 16-25% faster on the same PC. On Linux the gain is small (3-4% with unbuffered reads, none through the page cache) and never a loss. Thanks to adonizm, whose #833 this is a port of, and to malloc32 for bringing it back. -
A prompt whose experts are already in RAM (#1353). When the GGUF experts read in place are in the page cache (90% or more), the prompt reader uses its few-thread profile instead of the 32-thread SSD one, which had made RAM-resident prompts 4-7x slower for the reporter. On our 4x R9700 box (fast storage, 251 GB RAM) the old path was not slow to begin with, so we measured no change there (identical output); please tell us in #1353 what it does on yours. Where the engine cannot tell (Windows), it behaves as before;
STRATA_STAGER_SSD=0/1forces either. -
Intel Arc Pro B70: the expert dequant writes in whole 16-byte runs. 200 -> 423 GB/s of writes in the micro-benchmark, prompt +6%, output bit-identical.
-
Intel Arc A750: decode 9.8 -> 12.0 tok/s (+22%). The sampler no longer needs FP64, so the A-series FP64 emulation settings are not needed (docs/INTEL.md).
-
Tokenizer cache. A resent agent prompt (89K tokens) is encoded in 0.07 s instead of 0.32 s, the same ids.
-
Automatic card order. With an auto layer split, the faster card (more SMs x clock; on Linux AMD from the KFD topology) is put last, the order that measured faster in the reports that asked for it (#1352).
"gpu_order": "as_given"keeps your order. Setup no longer writes--remote-expert-optnext to--pipeline-windows(the first turns the second off). The auto placement also sees each stage's PCIe link (#794, from shy). -
Prompt batches fill holes (#793): a request takes a free slot below a busy one of its group first, so a request left alone goes back to the plain single-request path.
Fixes
- A deadlock with
parallel: 2fixed. Long, long, short requests could hang forever: a request waiting for a slot gave back nothing while it waited, and the slot a paused read held needed those lines. - A full conversation cache evicts old entries to pass the RAM check instead of dropping the snapshot (#1347; a conversation that holds a pinned shared prefix is never evicted this way). Thanks to alanthinker.
- Hung engines (#1317, #1407). A request whose engine prints nothing for 90 s, uses no CPU, moves no disk bytes and has an idle GPU is ended and restarted (
STRATA_ENGINE_STALL_S, 0 turns it off; needspsutil, which setup installs). An engine that is silent but working is never ended by this. The prompt-stall watchdog now waits, up to 10 times its limit, while the OS is still reading the expert file (STRATA_WATCHDOG_IO_S, 0 turns it off); a real hang stops at the old limit. CUDA_LAUNCH_BLOCKING=1hangs every verify window (#1341, #1383): the engine now says so at start and in the stall report, with the variable named. Thanks to ischencheng.--api-keytakes several keys separated by commas, as llama.cpp does (#1344).setup --calibratekeeps its earlier measurements when a later engine start fails; it drops that candidate and goes on (#1337). The PCIe sweep no longer ends at 1.0 and visits each value once per round (#1332, #1345, from aly8246).- A config with a
visionsection but no--visionin its args now starts the engine with images (#1322). - Tool calls written
<function= NAME>are read as the call NAME (#1430). The vision encoder's start error now names the busy or unavailable GPU, and setup warns about a non-Default compute mode (#1445). - q8_1 scale overflow at the three remaining emit sites is clamped like the others (#1448); default answers unchanged.
- Intel B70, two cards (#1440): the host mirror keeps to the first stage's layers and the free-memory estimate no longer subtracts the same VRAM twice; a single card is unchanged (11,338 cache slots as before). The doorbell spin bound is chosen per device at run time (#1397; no gain measured on a Linux B70, Windows keeps 2,000,000 reads).
- Prompt buffers (#1454, #1465): the half output shares the embedding storage (320 MiB less at 32,768 rows) and the hyper-connection norm products stay in existing scratch. Default answers are identical;
STRATA_EMB_REUSE_ACCOUNT=1lets the planner use the saved bytes (below). - Batch MTP with a shared draft head (#1329); a gfx906 build that compiles again (#1396); the engine links against the toolkit whose
nvccbuilds it (#1359). - Server: a malformed body gets a 400 with the traceback in the log; bodies over 256 MiB (
STRATA_MAX_BODY_MIB) and a bad Content-Length are refused before they are read; the engine warns when--batch-groupsis used with--batch-mtp. - #795 (crash on a CPU without AVX-512): we could not find a software cause, so this is hardening, not a claimed fix. The AVX-512 probe and the "is this format supported" helpers are out of the wide-instruction files, and the tests now disassemble the kernels and run the CPU tests under Intel SDE, so a wide instruction on an AVX2 path fails the build.
- AMD: a warning when the long-prompt path's lend coverage is short, and the method behind a 925 -> 2,503 tok/s prompt on an RX 7900 XTX is in docs/AMD_HIP.md (#1389, from zhhhn); hipBLASLt tuning tables for gfx1151 (alexdns1, #1388) and gfx1150 / Radeon 890M (#1395, 1.5-1.7x prompts in the author's runs; we cannot measure either card);
STRATA_PREFILL_STREAM_MIN=128in the gfx1151 fast configuration (#1391). - Tooling:
strata --version; setup prints "checksum verified"; update scripts drop their backup branch when the move fails; fewer build warnings; test fixes;tools/serve_load.py, a load generator for the server. Docs: Windows "Lock pages in memory" takes effect at the next logon (#1412), the V100 notes, ten community benchmark folders and more.
New opt-in switches
None of these changes the default.
--peer-deviceprompt share without P2P (#1251). Most GeForce pairs have no P2P, and the prompt share used to refuse them. It now goes through pinned host memory, with per-token sums sent back in FP16. On 4x R9700 with P2P forced off, IQ3_S: prompt 690 -> 737 tok/s at 4K (+6.9%) and 1,353 -> 1,515 at 32K (+12%), but decode was 4-5% slower in those cases and the answers change bits, so it is off by default. First-token KL against the old path over 12 prompts: 0.009 mean with FP16 sums (11 of 12 top tokens the same), 0.005 with FP32, and 0 (bit-identical) when whole rows are sent instead of sums; with P2P available nothing changes. Thanks to Thig, whose work this is, built on Eddoursul's second-GPU prompt path.STRATA_GDN_CHUNKED=1(#1372). The DeltaNet recurrence of the prompt in 32-token chunks, for cards with 128 SMs (sergqwer measured 1.8x there). We have no such card; on 48 and 28 SMs it is slower (0.87x and 0.73x), so=1does nothing below 128 SMs. Other bits (KL 0.002-0.009 on the cards we forced it on).- Foresight swap space (
STRATA_FS_SLOTS, #1348). Per-layer VRAM slots filled with recently missed experts. In our runs it was 0.91-1.00 of the default decode speed on an RTX 3060, a Tesla P100 and an RTX 5070, so it stays off; it may pay on a single 24 GB card with many misses per layer, which we do not have. STRATA_ROUTE_RESIDENT(experimental). Routing that prefers experts already in the cache, in verify windows. It changes the answers (KL 0.07 at margin 0.5). RTX 5070 Q2_0: +4.1% decode (6 of 6 pairs). RTX 3060 IQ3_XXS: -0.1% and -1.6%, because the draft's acceptance on code drops. Try it on a 12-24 GB card with a Q2_0 pack, and measure; do not expect a win everywhere.STRATA_EMB_REUSE_ACCOUNT=1. The prompt planner counts the VRAM the shared buffer saved, so a VRAM-limited card gets a slightly larger chunk (RTX 3060 6,400 -> 6,656 tokens, about +3.9% prompt speed). It changes the prompt's rounding, so it is off.STRATA_ADAPT_LAG=2. The adaptive tier's copies are waited for one window later. Tesla P100 +3.5% (6 of 6), RTX 5070 +0.2%, RTX 3060 -1.3%; opt-in.STRATA_SM70_TABLE=1(#1401). The V100 decode kernels from rewin123 and NoxestNoz, with a per-architecture rows table. We have no V100; they become the default once measured on one.- The CPU share above 1,024 tokens, on a layer split and from mapped experts (#1414, from architectds). Documented in docs/DETAILS.md with its numbers; in our check the 3,072-token extension was 2% slower at 2K on the RTX 5070, which is why the default stays below 1,024 tokens.
STRATA_PREFILL_CPU_SHARE_MAX=3072turns the extension on. STRATA_STAGE_PIN=1(#1237, from ZhongUncle). The buffers the expert cache is filled from are page-locked. Tesla P100 with RAM capped at 24 GB and--mmap-experts: 13.9 -> 15.0 tok/s (+8%). It was meant to be on by default, but the release check found IQ3_S answers corrupted after a 4,096-token prompt with it on (most likely a buffer reused before its copy finished), so it is off until that is fixed. Use it only to measure.
Updating
UPDATE.bat (Linux: ./update.sh). Models and configs are not touched. If you want exactly the old answers on one NVIDIA GPU, put "STRATA_PREFILL_CPU_SHARE": "0" in the "env" block of your strata-<model>.json.
Thanks to everyone who reported, measured and sent patches; in particular sergqwer (the CPU share and the chunked recurrence), adonizm (batched reads for Windows), Thig (the no-P2P prompt share), paulhothersall and lineape (the bisect behind the Pascal fix), and mlfather (the engine review that started the multi-GPU concurrency work).
Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going: