About 13-16% faster decode on one NVIDIA GPU or a layer split, up to 30% more on a single card whose CPU outruns its PCIe link, a long list of fixes for multi-GPU splits, Windows AMD, Intel Arc and tight-RAM machines, and new opt-in speedups for AMD and multi-GPU cards. Two new NVIDIA defaults change answers a little (STRATA_ROUTE_TAIL_SKIP=7 and the measured PCIe share; each has a setting that turns it off). Update with UPDATE.bat (Linux: ./update.sh); setup replaces the engine with 0.1.42. The five fixes planned as 0.1.41.1 were never released and are part of this version.
Faster
-
Skipping the tail of missed experts, on by default on CUDA (
STRATA_ROUTE_TAIL_SKIP=7, from PR #1668 by LeGeRyChEeSe). In a verify window, an expert that is not in VRAM and that every token of the window ranks 7th or lower in its top 10 adds little to any token. It is now neither copied nor computed; the other experts keep the router's weights. Decode against=0, 5 interleaved rounds on fresh engines (story, code and a 6K-token document), median:Machine, model Decode Rounds faster RTX 3060 12 GB, IQ3_XXS +15.6% 5/5 RTX A4000 16 GB, Q2_0 +12.9% 5/5 Tesla P100 16 GB, IQ3_XXS +14.6% 5/5 2x RTX A4000, layer split (99.7% of experts resident) -0.2% no change A machine that holds nearly every expert gains nothing; prompt reading is untouched. It applies to CUDA serve with one GPU or a layer split, and is off when every expert is in VRAM, with
--batchslots, with a peer or helper GPU, and on AMD and Intel. It changes answers slightly (quality numbers under "Answers can differ");STRATA_ROUTE_TAIL_SKIP=0restores them. -
The PCIe share is set from what your CPU measures (CUDA, one GPU). Missed experts are split between the GPU (read over PCIe) and the CPU; the default share came from a start-up link probe alone and ignored the CPU. With no
--pcie-frac, the engine now times the CPU pool in the first decode windows (about 10 s) and sets the share to balance the two, in steps of 0.05 between 0.10 and 0.60, at most three times; each change prints onePCIe share:line. Decode against the old rule,STRATA_ROUTE_TAIL_SKIP=7on in both, interleaved pairs, median:Machine, model Share old -> new Decode Pairs faster RTX 3060 12 GB, IQ3_XXS 0.55 -> 0.20 +30.0% (47.7 -> 62.1 tok/s) 5/5 Tesla P100 16 GB, IQ3_XXS 0.29 -> 0.20-0.25 +1.1% 4/5 RTX A4000 16 GB, Q2_0 0.30 -> 0.30 no change (same share) --pcie-frac N(and--calibrate, which writes it) always wins, and a request's ownpcie_fracis kept. It applies to the native packs of the i-quants (Q2_0's own pack keeps its fixed share); off on a layer split, with--batch, a peer GPU, every expert in VRAM, and on AMD and Intel. A start on a busy CPU can raise the share; the cap is 0.60.STRATA_PCIE_FRAC_DEFAULT=oldrestores the 0.1.41 rule. -
gfx12 prompt switches on by default (#1478, report by jkuepker). Five prompt-path switches (
STRATA_GDN_HEAD,GDN_PP=2,GDN_CONVL2,GDN_NOY,CVEC_FUSE) are on for gfx1200/gfx1201 cards such as the Radeon AI PRO R9700; residual and logits are byte-identical to the old path at 4K and 20K tokens. R9700, IQ3_S, 10 pairs: prompt +0.8% at 4K, +1.6% at 16K, +1.8% at 32K (the reporter saw +8% and +4% on another config).STRATA_GFX12_DEFAULTS=0turns them off.
Final build against 0.1.41, default settings, 5 interleaved pairs, medians. The TAIL_SKIP gain comes on top of these for NVIDIA. A few percent on a shared box or from the draft acceptance rate is noise; no machine got consistently slower:
| Machine, model | Story decode | Code decode | Prompt 4K |
|---|---|---|---|
| RTX 3060, IQ3_XXS | +10.4% | +0.3% | +4.8% (shared box, 620-686 spread) |
| Tesla P100, IQ3_XXS (start-up PCIe probe, default flags) | -3.6% | -2.1% | 0.0% |
Tesla P100, same with --pcie-frac 0.29 (8 pairs)
| -0.5% | -0.1% | 0.0% |
| RTX A4000, Q2_0, 1 GPU | +4.3% | -0.1% | 0.0% |
| 2x RTX A4000, Q2_0, layer split | +3.7% | -0.2% | -0.1% |
| R9700, IQ3_S (two boxes) | +0.03% / +0.1% | +0.01% / +0.15% | +0.9% / +0.3% (16K +1.5%) |
| Arc Pro B70, IQ3_S, MTP | +0.9% | n/a | +0.1% |
Arc A750, IQ3_XXS, default / --pcie-frac 0
| -4.4% / +3.0% | n/a | -1.0% / equal in the fast runs |
On an RTX 5070 (Windows, Ryzen 5 7600) with Q2_0 and setup's own settings, 3 interleaved rounds: decode 77.0 -> 85.6 tok/s (+11.2%, 3 of 3 rounds faster, the tail skip's share), prompts at 4K 2,192 -> 2,193, 16K 2,778 -> 2,820, 32K 2,797 -> 2,833 tok/s.
The P100 and A750 default-flag swings come from the start-up PCIe probe (see "Answers can differ"); with the share fixed they are inside noise. Eight clients on 4x R9700: -0.7% total throughput, inside the run-to-run range, 0 errors.
Fixes
Crashes and regressions first.
- A card order you gave is kept (#1576, #1590, #1760). An automatic layer split puts the faster card last. On a pair with unequal VRAM that put the head, the draft layer and the verify on the card least able to hold them: the reporter's 12 GB card ended with a 3 GiB expert cache and decode fell from about 58 to 44 tok/s. The reorder is now skipped when it would put the last stage on a card with less total VRAM than your config chose, and cards whose speed scores are within 5% keep your config's order. Thanks to nekomario28.
- No one- or two-layer stages in the automatic split (#1616). A one-layer stage holds about 512 expert slots on a 12 GB RTX 3060 and clamps every card's prompt chunk. If the first pick has one, the engine searches again among placements with at least three layers per stage and prints both. Three is a rule of thumb from the reporter's numbers; we have no rig of that shape, so tell us if your split still looks odd. A near tie between two splits on a mixed pair now prints a warning (#1428).
- Windows AMD (HIP). A bare
--prefill autostops at 6,144 tokens (#1630): on an RX 7900 XTX the first long prompt at 8,192-token chunks charged about 70 GB of commit and the engine died (0xC0000409); it costs about 14% prompt speed there (645 against 748 tok/s). A value you pass is kept, with a warning from 8,192. Serve start prints the commit left and warns when planned host allocations plus 4 GiB do not fit, and the conversation cache uses the smaller of free RAM and free commit (#1607). The automatic expert cache keeps its 2,560 MiB VRAM reserve on 16 GB cards too (#1709). We have no Windows AMD box: these are guards that compile and pass their tests, not seen on a card. - AMD, 2+ GPUs: the default checkpoint save of a 104K-token prompt no longer fails. On two R9700 the save failed with "operation would make the legacy stream depend on a capturing blocking stream"; the snapshot copies ran on the legacy stream while another thread captured a graph. They now use a private non-blocking stream. Final check, 104K tokens, 2 cards, defaults, 5 interleaved pairs: 0.1.41 failed 3 of 5 runs, 0.1.42 passed 5 of 5 with the same output every run.
- CPU prompt share on an RTX 5090 (a user report). We could not reproduce the crash, so the cause is not proven. The 0.1.41 default share is now skipped when free RAM is below its buffers plus 3 GiB, when the Windows commit left is below buffers plus 4 GiB, or when its experts are file pages RAM cannot hold; the log says why. An explicit
STRATA_PREFILL_CPU_SHAREis kept and only warns. If the host cannot give the page-locked buffers, the share ends for that run instead of failing the chat. It is also not armed when every expert is in VRAM (#1595). - End of the context (#1700, midagedev). A 4,094-token prompt with
--max-context 4096 --max-new 2 --spec 4stopped with "ran out of context"; it runs now. - Intel. Without
--mtpthe engine could crash at exit (#1642): 6 of 6 runs on an Arc Pro B70, 0 of 3 now. After a batch window a mirrored-expert B70 answered!!!!to the next single request; fixed (#1721). A layer-split verifier no longer adopts the first card's queue (#1555, #1589), which also explains #1440's "graph associated with a different context". With--batch-mtpand a draft-vocab head, the slot drafters copy the shared head's type instead of failing with "unsupported native MMVQ GGML type" (#1121, #1790). We have one card each: two-card B70 and B60 cases are untested. - Pascal and Volta prompts with little free VRAM (#1757). The BF16 products run through FP32 conversion buffers; when those did not fit, the call fell through to a raw BF16 cuBLAS call that Pascal cannot run (status 14). The weight is now converted in tiles, and with no room for even a small tile the engine stops with a message naming the settings to change. On a Tesla P100 forced into the tiled path, a 2,495-token prompt gave the same answer word for word; we could not reproduce the original error on one card (it came from a layer split), so tell us in #1757 if it remains.
- A layer split stalled on Windows (#1740, #1741, architectds). The pipelined decode loop switched the CUDA device on every poll, millions of times a second; on an RTX 3060 + RTX 5070 Ti pair the host stalled inside one of those calls. It now returns before any CUDA call while nothing is waiting.
- cuBLAS FP16 fallback (#1650, #1659). A failing FP16 GEMM shape runs on a fixed algorithm. It needs a Turing plus Pascal pair or cuBLAS 12.4; we could not trigger it on two RTX A4000, so this is the reporter's fix, built and gated on CUDA 12.9 and 13.4.
- A cancelled request keeps the conversation it restored (#1620). The follow-up reused 1,834 of 1,851 prompt tokens; 0.1.41 reused 1,523. A cancel inside a batched or pipelined chunk behaves as before.
STRATA_STAGE_PIN=1corrupted answers (#1237). A page-locked buffer was reused before its copy finished (16 copies of token 0 after a 4,096-token prompt on IQ3_XXS). Fixed with an event per buffer; 4K and 20K prompts are now identical. It stays opt-in.- Parallel requests (#1603). A request could end up neither running nor queued and wait forever. We did not reproduce the exact hang; slot bookkeeping now tracks owners and reaps orphans. If it still happens, send
/metricsand the log in #1603. - Thinking budget (#984). A budget at or above what
max_tokensleaves could never fire, so the reply ended mid-thought. Thinking now closes atmax_tokensminus max(512, a quarter of it). - Memory guard in containers (#1601, noon-at-cgn). Parking and session SAVE read the RAM the engine really has: the smaller of
MemAvailable, the cgroup limit's room, and--memory-limit-mibminus use. Tested under a systemd memory cap; not with docker, cgroup v1 or Windows. - CPUs without AVX2 (#1701, #1697, #1699, bossman). 128-bit AVX kernels for the Q2_0 and i-quant rows. 0.1.41 died with SIGILL in two CPU tests under Sandy Bridge and Ivy Bridge emulation (qemu-user); 0.1.42 passes there. No real CPU of that age was tested.
- Smaller. Queued requests send keep-alives and give up their place when cancelled (#1619); a READY timeout ends with a 503 (#1527); overflow errors start with "request exceeds the context window:" (#1615);
--batchon a layer split takes a yield with no slot decoding (fixed by reading the code; the 2-GPUtools/batch_interleave_test.pyrun was not done); GCC 10 builds (#1488); a tensor-less first GGUF shard opens (#1611); the gfx906 build no longer arms the CUDA-only CPU prompt share (#1732); CUDA MMQ error reporting and the HIP stand-in targets link cleanly (#1812).
New opt-ins
None of these changes the default.
- Fused int8 prompt experts on gfx12 (
STRATA_PF_FUSED=1, #1570, jkuepker). R9700, IQ3_S: prompt 4K 761.6 -> 876.9 tok/s (+15.1%), 16K +17.4%, 32K +18.8%, 10 of 10 pairs each. First-token KL against the FP16 prompt path: IQ3_S fused 0.00284 against 0.00332 for the default MMQ path, but IQ3_XXS (24 prompts) fused 0.00388 against 0.00340. Because IQ3_XXS is not at or below MMQ, it stays opt-in. Q2_0 does not use these kernels. - Pipeline windows on 3+ layer-split stages (
--pipeline-windows 2, #1656, Cass67). 4x R9700, IQ3_S, 8 pairs, off -> on: code 96.4 -> 128.9 tok/s (+33.7%), 4K prompt +31.3%, 32K +22.8%, prose +2.2%, identical ids. On 3 cards code +29% but prose -0.9%; on 2 cards code -3.3%. Those losses keep it opt-in. --batch-mtpon a layer split (#1636, noon-at-cgn). Only when the experts do not fit in VRAM: on 2x R9700 with everything in VRAM total throughput fell 80.2 -> 62.7 tok/s (-21.9%, 0 of 5 pairs faster); answers identical.- Data-parallel replicas (
"replicas": 2,--replicas N,setup --replicas N). N engines on their own card groups and CPU cores behind one server; a conversation goes back to its replica. 4x R9700, IQ3_S,--batch 8, 2 replicas of 2 cards against one engine on 4 cards, 5 pairs: 8 clients 130.0 -> 160.6 tok/s (1.24x), 16 clients 1.45x, 40 clients 1.39x, 5 of 5 pairs; answers byte-identical to a solo run. Costs two copies of the model in RAM. Core pinning is Linux only (Alex Gorevski's #1436); no restart test on real hardware. - Multi-GPU extras (noon-at-cgn, Evan).
--aux-cpus(Linux): decode +1.7% / +1.8% on 2x R9700, prompt about -1%, no effect on a P100.STRATA_SPLIT_MTP_BATCH=1: prompt 32K +7.5% (10 of 10), decode flat, but the draft K/V rounds differently.STRATA_SPLIT_RING=384, 104K prompt: +2.3%, file reads 98.3 -> 34.5 GB (#1600, on #1190). - More routing switches (change answers).
STRATA_ROUTE_PRIOR=0.5on top of the new default (the other half of #1668): about +26-28% over 0.1.41 for both together, about three times the KL (0.017 mean on the RTX 3060).STRATA_ROUTE_RESIDENT=0.5: RTX 5070 Q2_0 with a 22 GiB RAM budget +16.5% decode (5 of 5), KL 0.0085 mean on the 3060;STRATA_ROUTE_RESIDENT_MTP=1adds no speed or KL change over it. - Intel A-series prompts (
STRATA_PF_XMX=1or2, #1711, aslater3). Arc A750, 4K prompt, 5 rounds: prefill 72.8 -> 88.8 tok/s (mode 1) and 105.2 (mode 2); decode within 2%; first-token KL 0.002-0.018. Not built for Battlemage. Also opt-in:STRATA_VERIFY_STEPPED=1(#1713, slower on an A750: decode 10.6 -> 9.7 tok/s),STRATA_SYCL_A770_FAST=1(#1710, duncanmcqueen),STRATA_DOORBELL_CHECK=1(#1602; the checksum wait cost 17% of A750 decode). The rest of #1602 is in with defaults scoped to what we could test, and an A750 decodes the same as on 0.1.41. We have no A770, so its claims are unverified. - Sessions.
--session-save-reclaim(#1573, tuandat3019): a SAVE that fails its RAM check frees retained K/V, old parked conversations and unpinned checkpoints, then retries (tested with 420 MiB free: refused without it, worked with it, identical ids).--conversation-cache-min-tokens N(#1597) parks only conversations of N tokens or more. - Other opt-ins.
STRATA_STAGE_PIN=1(now exact).STRATA_MMQ_RESIDENT_SORT_NE(#1660), helper-cache prefill (#1649), a pipelined commit guard (#1674): output identical at 4K and 20K, speed claims are the authors'. gfx906 (no MI50 here, so speed is the author's claim):STRATA_GFX906_ATTN_QUERY_SWIZZLE=1(#1661),STRATA_GFX906_ATTN_REDUCE12=1(#1718) from 0FL01; the build refuses a wave32 card (#1728).STRATA_HEAD_MIX_MULTI=1,STRATA_ONE_TOKEN_COMMIT=1(#1479): identical answers, no measurable gain. - Config keys and API.
"prevent_sleep": truekeeps Windows awake while a request runs or is queued (#1727). Checked on Windows: the server asks Windows to stay awake when a request starts and releases it when the work is done; with the key off nothing is asked."reasoning_effort"sets the default thinking level (#1641)./metrics/prometheusexposes counters behind the API key (#1635).logit_biasis applied in CUDA and HIP sampling, only for requests that send it (#1646): -100 is a hard ban. - Tools (measurement only).
STRATA_LOOKAHEAD_STATS=1andtools/replay_cache_policy.pyrecord and replay the expert-cache policy offline; they do not change decoding.
Setup and server
Setup: a damaged llama.cpp source archive is downloaded again instead of surviving behind its done mark (#1797, #1801); the SYCL setup reads the engine version from the source (#1794, #1817); a machine without CUDA finds Visual Studio again (#881, #1792); quant RAM budgets kept on a layer split (#1564); engine updates compile for the cards your models use (#1485); --gpus names AMD cards on AMD-only PCs (#1594); a broken cmake/ninja is skipped (#1605); CC/CXX reach CMake (#1645); the swift family takes IQ3_S (#1651); on a unified-memory APU setup writes an expert-cache number, not auto (#1715; conservative, no Strix Halo box here). Server: a reordered tool set keeps its first-seen order so the prompt cache survives (#1624); wrong-typed sampling fields are a 400 naming the field, numeric strings are converted (#1663); unread content parts are dropped, never a 400 (#1664); llama.cpp's repeat_penalty / repeat_last_n are accepted in the config's sampling block and in requests (#1819); effort_position "end" finds the option in the Intel engine behind its launcher script (#1781, #1782); with reasoning_loop_recovery on, a loop written as one long word, such as a cycling digit string, is caught (#1753, #1777). Web app and Monitor: a favicon, the PCIe tile for AMD cards on Linux (#1733), and sparklines no longer scaled by their own noise (#1731). Engine start: a PLE probe (under 0.6 s) switches --ple-io direct to mmap below 15,000 rows/s, with a message; STRATA_PLE_PROBE turns it off (#1425, #1549, #1629; same bits). On Windows the resident RAM mode builds its RAM copy with unbuffered reads (#1696, from sergiywith): on an RTX 5070 PC with Q2_0 it page-locked 28 GiB of experts in 11 s with at least 20 GB of RAM left free throughout, where 0.1.41 drove the same PC down to under 1 GB free while copying; answers are the same as without the resident mode. Twenty-one community benchmark folders are in bench/results, among them a Radeon Pro VII (gfx906), 2x Tesla P40 at 131K, an RX 9060 XT, an RX 9070 and an RTX 5060 Ti, plus an Arc A770 decode investigation in the docs (#1765).
Answers can differ
-
STRATA_ROUTE_TAIL_SKIP=7is on by default on CUDA. Teacher-forced KL of the real decode path against the same engine with it off (28 prompts of code, reasoning, chat, six languages and long-document summaries, 220 tokens each, about 6,000 scored tokens per set):RTX 3060, IQ3_XXS RTX A4000, Q2_0 Mean KL, greedy / sampled text 0.0058 / 0.0056 0.0032 / 0.0030 Top-1 token agreement 97.7% 98.2% Perplexity of the forced text x1.011 x1.005 Same engine run twice (the noise floor) 0.0007 0.0007 For scale, IQ3_XXS against Q2_0 on the same text differs by about 0.10. HumanEval and the reasoning set did not move beyond their own run-to-run noise, though at 82 and 30 tasks that noise is large.
STRATA_ROUTE_TAIL_SKIP=0gives the exact 0.1.41 answers, together withSTRATA_PREFILL_CPU_SHARE=0(with both set, the 10-prompt checks and 4K/20K residual dumps were identical on the RTX 3060, the A4000 alone and as a 2-GPU split, and the P100). -
The measured PCIe share (CUDA, one GPU, no
--pcie-frac). Where it moves the share, which experts the CPU computes changes. Teacher-forced KL on the RTX 3060 (share 0.55 -> 0.20), against the old share: mean 0.0015, against 0.0013 for the old share run twice; top-1 agreement 98.7% in both.STRATA_PCIE_FRAC_DEFAULT=oldor--pcie-frac Nkeeps the old behaviour. -
The CPU prompt share (0.1.41's change) is now skipped on machines short of RAM or commit; there the answers equal those with
STRATA_PREFILL_CPU_SHARE=0(first-token KL against the share averaged 0.004, max 0.025, in the 0.1.41 check). -
A thinking budget above
max_tokensless its reserve closes thinking early (#984). Nothing changes without a budget or with a small one. -
A tool set that comes back reordered keeps its first-seen order (#1624); the first request with a set is byte-for-byte what the client sent.
logit_biaschanges output only for requests that send it. -
Opt-ins marked "changes answers" (route prior, route-resident, fused prompt experts, XMX, split MTP batch) only when you turn them on.
Checked against 0.1.41 with STRATA_ROUTE_TAIL_SKIP=0, STRATA_PREFILL_CPU_SHARE=0 and fixed flags (final build): 10-prompt identity runs (and 4,096 / 20,000-token residual dumps where run) pass on the RTX 3060, Tesla P100, 1x and 2x RTX A4000, R9700 (one card and a 2-card split, gfx12 switches on and off), Arc Pro B70 and Arc A750. The A750 matches only with --pcie-frac 0 and equal expert-cache slots; its default PCIe share comes from a start-up bandwidth probe, which also makes a Tesla P100 alternate between two outputs on 0.1.41 itself. --pcie-frac N makes output repeatable. The A750 with MTP drafts differs from run to run on 0.1.41 as well (3 of 10 texts equal between two 0.1.41 runs), so there all 10 prompts passing on both versions is the check. On the RTX 5070 (Windows) the shipped CUDA 13 engine matches 0.1.41 on Q2_0, IQ3_XXS and the Coder: 10/10 identical answers each, and identical 4,096 / 20,000-token residual dumps.
Not fixed in this version
- Arc A750:
--batchfails with an out-of-host-memory error (also on 0.1.41). Arc Pro B70 withSTRATA_VERIFY_NO_HOST=0and adaptive swaps went from about 31 to 58 tok/s but hit a GPU hang once, so we do not recommend it. - Windows AMD with KV streaming (a 65,536 context): "unspecified launch failure" in the draft layer's prompt pass (#1644); we cannot test Windows AMD.
- Open, logs requested: #1723 (Windows crash while parking a very large conversation), #1682, #1683, #1686, #1703, #1648.
Updating
UPDATE.bat (Linux: ./update.sh). Models and configs are not touched. For exactly the old answers on NVIDIA, put "STRATA_ROUTE_TAIL_SKIP": "0", "STRATA_PCIE_FRAC_DEFAULT": "old" and "STRATA_PREFILL_CPU_SHARE": "0" in the "env" block of your strata-<model>.json.
Thanks to everyone who reported, measured and sent patches. Standouts this time: LeGeRyChEeSe (the tail-skip switch, now a default), noon-at-cgn (the memory guard, multi-GPU placement and batch work), duncanmcqueen (the Intel Arc series, eight commits), nekomario28 (the card-order fix, the layer-split queue fix and a run of serve fixes), jkuepker (the fused gfx12 prompt kernels and the R9700 default switches), Cass67 (pipeline windows on three or more cards) and bossman (AVX kernels for CPUs without AVX2). Thanks also to the community benchmark reporters, whose thirteen folders are in this release.
Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going: