Strix Halo (gfx1151) support, crash, NaN and security fixes, better multi-GPU, batching and resident modes, Q4_0 / Q4_1 and more PLE table formats, AMD fixes, a much faster Linux cold start, and a long list of opt-ins.
Strix Halo / Ryzen AI Max (gfx1151, experimental): the HIP engine now runs on the Radeon 8060S / 8050S, and setup recognizes the chip (PCI 1002:1586), sizes memory for it as one shared pool, recommends UD-IQ4_XS, and on Linux picks HIP by itself. The engine has RDNA3.5 matrix-core kernels for the prompt, a hipBLASLt table for ROCm 7.14.1, free-memory accounting for unified-memory chips (#409, a DGX Spark gets it too) and 18 speed switches that turn on by themselves on gfx1151 (STRATA_GFX1151_DEFAULTS=0 turns them off). All of them gave byte-identical output here. The Windows HIP zip now includes gfx1151 code. It is experimental, and we have not run it on Windows Strix Halo hardware. docs/STRIX_HALO.md has the toolchain (no root needed), a manual build and the opt-in prompt switches. The 8040S mapping in setup is a guess. Measured on a Ryzen AI Max+ 395 with 128 GB, UD-IQ4_XS, medians of interleaved runs:
| Context | Prompt tok/s | Output tok/s |
|---|---|---|
| 8K | 1,293 | 53.8 |
| 64K | 1,370 | 46.6 |
| 128K | 1,320 | 51.4 |
That is +6% prompt and +4% output over our earlier engine branch. UD-Q4_K_XL, IQ3_S and IQ3_XXS are in the doc. The table matches the final defaults, which keep the shared-expert stream fork on for gfx1151 (it is the 18th automatic switch): UD-IQ4_XS at 8K gave 54.07 output against 53.94 in the table. Thanks to StevenChenSE for the RDNA3 matrix-core GEMM kernels (#313).
Crash and NaN fixes:
- Non-finite values (#838, #550, #871, #606, #879): the fused SwiGLU q8_1 quantizers keep their scale finite. The residency-table upload waits for its own copy before the verifier reads it. The all-resident verify graph runs only while every expert is in VRAM, and a stale plan fails the window instead of printing
!!!!. A Stop sent during a prompt read reaches the engine at once. We could not reproduce #879 here, so it stays open for a retry on 0.1.40. --batchon a split (#776, #845): a fully resident split stage no longer waits for doorbells that are never rung.- A dead engine (#748): a fatal verify timeout is caught before batch routing, and the next request is refused.
- Parking a conversation (#752): a host out-of-memory skips the park instead of ending the server.
- HIP: the shared-expert stream fork is off by default on AMD, which brings decode back to 0.1.38 speed (#826, #816). The doorbell kernels fence after the ring store (#697), which may explain the batched-prompt stalls in #579 and #613.
- Windows: the RAM budget respects the commit limit (#749, #730), a large-page arena is no longer VirtualLocked (#779), and Smart App Control blocks get a clear message (#735).
Security:
- #553: network image paths are refused, URLs are capped at 32 MiB, and a page from another origin cannot name a local file as a picture. We added the same check to
/v1/responses, which the PR did not cover. - We did not merge #434 (a Manager page that any website can drive, plus a proxy that adds the API key) or #696 (the API key becomes a shell on a LAN server). The replies say what each needs.
More than one GPU, batching and resident modes:
- Resident RAM mode on a layer split (#848, thanks to Francesco Albano): about 70 tok/s against 32-64 on 32 GB machines. It also fixes a 0.1.39 bug where the adaptive swaps copied back from the wrong card's cache. Setup keeps the resident variant on a split only with an engine that has it, so setup now needs engine 0.1.40.
- A helper GPU's experts are not in the PCIe share (#854): 40 -> 66 tok/s on a helper rig, the author's number.
- Four GPUs (#887): the layer split scores every four-way placement instead of guessing. This changes the automatic placement on 4 GPUs. 2 and 3 GPUs keep 0.1.39's, and
STRATA_SPLIT_COVER_Basks for the new one at any count. - Stage weights (#880):
--layer-split autocan trim the later stages' weights (4 GPUs, prompt 470 -> 1,930 tok/s, the author's number).--vram-reserve-later-mibsets a separate reserve for the later cards. - Batching:
--batch-mtpkeeps MTP drafting in concurrent slots (#846, +31 to +39% with 2-4 clients, the author's number).serveruns without--mtp(#708). - Fusions ported from Eddoursul's fork: the draft round runs the draft layer's front once, a one-token window commits itself, the draft kernels use the main layer's forms, the final mixer is one fused read, and the argmax runs over many blocks.
--host-core first|lastand the--spec-follow/--window-hashestest tools come from the same fork. The #783 kernel work (QSA early exit, GDN decode kernels, vec4 router, batched KV append, multi-token GR) is merged with a kill switch for each. Its long-context QSA early exit gives +2.7% decode at 120K (6 pairs).STRATA_GR_DOWN_MAX4stays opt-in, because it gave no gain on the 5070. - Decode against 0.1.39 on NVIDIA (RTX 5070, 12 pairs, 200-token medians): Q2_0 code -0.3%, story +3.4%; IQ3_XXS code +2.9%, story +2.3%. Prompts are within 2%.
New formats:
- Q4_0 and Q4_1 experts (#599) on the GPU kernels, and a Q4_0 PLE table.
- PLE tables in Q8_0, Q5_1 and BF16 (#651, #586, #865): one table of PLE formats now covers these and Q5_0 and FP8, with bounds checks on the file. Against the BF16 table, Q8_0 is off by 0.53%, FP8 by 2.64% and IQ4_NL by 7.60%. The effect on answers is small (KL about 0.012 for any of them), so setup does not offer a bigger table.
--ple-io ramfaults the table in on 16 threads (8 GB cold in WSL: 15 s -> 6 s). - Unsloth UD-Q6_K_XL: the Q8_0 PLE table is read, and the Q6_K expert kernels are an opt-in build (
-DSTRATA_Q6K_EXPERTS=ON, #860). - IQ3_XXS and IQ4_XS PLE keys stay native, so those packs start again (#381).
--kv k8v4streams with--kv-resident(#711, #705), and setup offers it.
AMD:
- RDNA2 (#835, opt-in):
STRATA_HIP_PROMPT_F16=1runs the prompt's 16-bit GEMMs in FP16, and prompts read about 2x faster on an RX 6900 XT. It rounds differently, so it is off for now. The engine prints a tip on gfx103x. - hipBLASLt tables for gfx1100 (#766, #755, #565), RDNA3 / RDNA4 sibling arch names (#695), gfx1034 (#778) and a gfx906 build fix (#808).
- Windows HIP (#873, #654): one
hipFree(0)warm-up before the first memory query. - Opt-in:
STRATA_HIP_ADAPT_KERNEL_COPY=1does the adaptive swaps' copies with a kernel instead of SDMA (#884, for the 2x gfx1030 hang).STRATA_DENSE_MMQ=1(#820) gave no gain on gfx1151 and stays opt-in. - Older cards: the Q4_0 expert kernel overflowed on sm_60 (signed overflow). It is fixed, found on a new P100 test box.
- Intel Arc (#784, #866, #868, #870): the SYCL port compiles against 0.1.39's sources again, a ring wait over 2 ms is no longer reported as finished, and the Arc Pro B60 ids are known.
Linux cold start: weights, head, embedding, vision encoder and the RAM copy are now read ahead of use (#699). On an RTX 5090 with 32 GB, Q2_0 at 262K, ready in 70 s instead of about 920 s. STRATA_READ_AHEAD=0 turns it off. Linux also gets the unbuffered file tier (#773, O_DIRECT, same rule as Windows). On a 16 GB laptop with 30 GB RAM, decode went 4.4 -> 7.9 tok/s and startup 36 -> about 20 s (author's numbers). With transparent huge pages on always the arena no longer asks for more (#771, STRATA_NO_ARENA_THP=1).
Defaults that changed (the rest of the default output is unchanged):
- Empty assistant turns (#886): they are not rendered into the next prompt, so a client's history can't teach the model to answer with nothing.
STRATA_KEEP_EMPTY_TURNS=1restores the old rendering. tool_choice(#790):noneoffers no tools on both routes.requiredand a named function work.json_objectmode (#762): JSON is taken out of a code fence or the prose around it.- Stop strings (#454): OpenAI
stopand Anthropicstop_sequencesare honoured on every path. - Workers on hybrid CPUs (#798): P and E cores come from the CPU's own lists. A Core Ultra 7 270K Plus is now 8P + 16E, 15 workers.
- Windows (#691): the server is opted out of power throttling, so a minimized window no longer moves the engine to the E-cores.
- Conversation caches from 0.1.39 are rebuilt on first use: the cache's version key changed, so the first request after the update reads its prompt again.
- Setup's context menu has a 200K step (#608). It is never the recommended one.
Opt-ins (off by default, the output is unchanged):
- Server:
"strata_checkpoint": falsefor a one-shot request (#861),"return_progress": true(#837),"reasoning_loop_recovery": "stop" | "recover"(#869), session files with--slot-save-path(#668),--prompt-cache-tail(#614),STRATA_CACHE_MESSAGE_BOUNDARY=1(#734). - Memory:
--kv-grow, the K/V takes VRAM as the context grows (#378). - Speed:
STRATA_Q2_BITPLANE=1(#706),STRATA_PREFILL_EQUAL=1(#693),STRATA_KV_PREFETCH=1(#732, 4-5% slower on our 5070, so off),STRATA_IQ3S_MT1=1(#863),STRATA_DMA_BATCH=1|2(#807),STRATA_ADAPT_LAG=2(#764). - Multi-GPU:
--adapt-async 1(#876),STRATA_EXCHANGE_ROTATE=1(#864),--pipeline-windows(#859),STRATA_DISJOINT_ADAPT=1(#731). - Tools:
python setup.py --rollback-engineputs the previous engine back (#670), andstrata_pack build --forcerebuilds into a used folder (#634). - Turing prompts above about 90K cells take the wide top-k kernel (#743,
STRATA_TOPK_STREAM=0restores the old one). STRATA_DEFERRED_REGISTER=1(Windows, #285 part 2) registers the arena per layer on a thread ahead of the readers. It cost about 8% decode here, so it is off.
Testers wanted (we have one RTX 5070 and no second card):
- 2 GPUs: a 30-minute soak of the resident RAM mode on a split (#848, #642),
--batchon an all-resident stage (#776), and--pipeline-windowswith 10 or more interleaved pairs, with and without #848 (#859). - 4 GPUs: the new automatic placement (#887) and
--layer-split autowith the trim (#880). Tell us if the auto layout got worse. - Turing (RTX 20): a prompt at 90K, 131K and 155K, before and after (#743).
- Intel Arc: the SYCL port on an A- or B-series card (#867, #870).
- Older cards: Pascal and Volta owners with a result from the CUDA 12 engine. We now have a P100 for sm_60, but not Volta.
- Windows AMD: the HIP zip's gfx1151 code on a Strix Halo PC,
STRATA_HIP_PROMPT_F16=1on gfx103x (#835), the SDMA kernel copy on 2x gfx1030 (#884) and the HIP warm-up (#654). - Open reports where a retry helps: #879 (both reporters), #871 at 100% residency, #828 (
MALLOC_CHECK_=3) and #795 (BIOS microcode update).
Thanks to everyone who sent PRs, tests and reports. Special thanks to Eddoursul, whose fork shaped this release, to Francesco Albano for #848, to sergiywith for #838 and #550, and to StevenChenSE for the Strix Halo GEMM work.
Checked before the release:
- The same answers as 0.1.39 on all four quants (Q2_0, IQ3_XXS, IQ3_S, the Coder), 10/10 each with a fixed cache, on the final tip. The prompt path's internal state is identical (exact_pp). A Tesla P100 (sm_60) gave 10/10 on 3 quants.
- Speed: the decode numbers above, interleaved against 0.1.39 on the RTX 5070.
- Strix Halo: output is byte-identical to the earlier branch.
- AMD and Intel: the HIP engine compiles, and the Windows HIP zip builds with gfx1151 (598 MiB). Nothing here ran on an AMD or Arc card except gfx1151 on Linux.
Updating: run UPDATE.bat (Linux: ./update.sh). Setup installs engine 0.1.40.
The ready-made Strata engines for Windows, which START-HERE.bat fetches by itself:
strata-windows-x64.zip: NVIDIA (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0, needs an NVIDIA driver 580 or newer. Contents:strata.exe,strata-vision.exe(the optional image encoder),BUILD.json.strata-windows-x64-cuda12.zip(experimental): older NVIDIA cards (Pascal and Volta: sm_60, sm_61, sm_70; it also has sm_75, sm_80, sm_86, sm_89 + PTX for a mixed PC), CUDA 12.9, needs an NVIDIA driver 528 or newer. Setup fetches it only for a model you put on such a card, or with--cuda 12. Contents:strata.exe,strata-vision.exe,BUILD.json.strata-windows-x64-hip.zip: AMD (gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030, gfx1151), ROCm 10.2.0a20260930 from AMD's TheRock builds, needs a current AMD driver. gfx1151 (Strix Halo) is in this zip too, experimental and untested on Windows hardware. Contents:strata.exe,strata-device.exe, the HIP runtime next to them,BUILD.json,rocm\(the ROCm libraries and their licenses).
Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going: