github Niko1221/Strata v0.1.40
Strata v0.1.40

latest release: v0.1.40.1
7 hours ago

Strix Halo (gfx1151) support, crash, NaN and security fixes, better multi-GPU, batching and resident modes, Q4_0 / Q4_1 and more PLE table formats, AMD fixes, a much faster Linux cold start, and a long list of opt-ins.

Strix Halo / Ryzen AI Max (gfx1151, experimental): the HIP engine now runs on the Radeon 8060S / 8050S, and setup recognizes the chip (PCI 1002:1586), sizes memory for it as one shared pool, recommends UD-IQ4_XS, and on Linux picks HIP by itself. The engine has RDNA3.5 matrix-core kernels for the prompt, a hipBLASLt table for ROCm 7.14.1, free-memory accounting for unified-memory chips (#409, a DGX Spark gets it too) and 18 speed switches that turn on by themselves on gfx1151 (STRATA_GFX1151_DEFAULTS=0 turns them off). All of them gave byte-identical output here. The Windows HIP zip now includes gfx1151 code. It is experimental, and we have not run it on Windows Strix Halo hardware. docs/STRIX_HALO.md has the toolchain (no root needed), a manual build and the opt-in prompt switches. The 8040S mapping in setup is a guess. Measured on a Ryzen AI Max+ 395 with 128 GB, UD-IQ4_XS, medians of interleaved runs:

Context Prompt tok/s Output tok/s
8K 1,293 53.8
64K 1,370 46.6
128K 1,320 51.4

That is +6% prompt and +4% output over our earlier engine branch. UD-Q4_K_XL, IQ3_S and IQ3_XXS are in the doc. The table matches the final defaults, which keep the shared-expert stream fork on for gfx1151 (it is the 18th automatic switch): UD-IQ4_XS at 8K gave 54.07 output against 53.94 in the table. Thanks to StevenChenSE for the RDNA3 matrix-core GEMM kernels (#313).

Crash and NaN fixes:

  • Non-finite values (#838, #550, #871, #606, #879): the fused SwiGLU q8_1 quantizers keep their scale finite. The residency-table upload waits for its own copy before the verifier reads it. The all-resident verify graph runs only while every expert is in VRAM, and a stale plan fails the window instead of printing !!!!. A Stop sent during a prompt read reaches the engine at once. We could not reproduce #879 here, so it stays open for a retry on 0.1.40.
  • --batch on a split (#776, #845): a fully resident split stage no longer waits for doorbells that are never rung.
  • A dead engine (#748): a fatal verify timeout is caught before batch routing, and the next request is refused.
  • Parking a conversation (#752): a host out-of-memory skips the park instead of ending the server.
  • HIP: the shared-expert stream fork is off by default on AMD, which brings decode back to 0.1.38 speed (#826, #816). The doorbell kernels fence after the ring store (#697), which may explain the batched-prompt stalls in #579 and #613.
  • Windows: the RAM budget respects the commit limit (#749, #730), a large-page arena is no longer VirtualLocked (#779), and Smart App Control blocks get a clear message (#735).

Security:

  • #553: network image paths are refused, URLs are capped at 32 MiB, and a page from another origin cannot name a local file as a picture. We added the same check to /v1/responses, which the PR did not cover.
  • We did not merge #434 (a Manager page that any website can drive, plus a proxy that adds the API key) or #696 (the API key becomes a shell on a LAN server). The replies say what each needs.

More than one GPU, batching and resident modes:

  • Resident RAM mode on a layer split (#848, thanks to Francesco Albano): about 70 tok/s against 32-64 on 32 GB machines. It also fixes a 0.1.39 bug where the adaptive swaps copied back from the wrong card's cache. Setup keeps the resident variant on a split only with an engine that has it, so setup now needs engine 0.1.40.
  • A helper GPU's experts are not in the PCIe share (#854): 40 -> 66 tok/s on a helper rig, the author's number.
  • Four GPUs (#887): the layer split scores every four-way placement instead of guessing. This changes the automatic placement on 4 GPUs. 2 and 3 GPUs keep 0.1.39's, and STRATA_SPLIT_COVER_B asks for the new one at any count.
  • Stage weights (#880): --layer-split auto can trim the later stages' weights (4 GPUs, prompt 470 -> 1,930 tok/s, the author's number). --vram-reserve-later-mib sets a separate reserve for the later cards.
  • Batching: --batch-mtp keeps MTP drafting in concurrent slots (#846, +31 to +39% with 2-4 clients, the author's number). serve runs without --mtp (#708).
  • Fusions ported from Eddoursul's fork: the draft round runs the draft layer's front once, a one-token window commits itself, the draft kernels use the main layer's forms, the final mixer is one fused read, and the argmax runs over many blocks. --host-core first|last and the --spec-follow / --window-hashes test tools come from the same fork. The #783 kernel work (QSA early exit, GDN decode kernels, vec4 router, batched KV append, multi-token GR) is merged with a kill switch for each. Its long-context QSA early exit gives +2.7% decode at 120K (6 pairs). STRATA_GR_DOWN_MAX4 stays opt-in, because it gave no gain on the 5070.
  • Decode against 0.1.39 on NVIDIA (RTX 5070, 12 pairs, 200-token medians): Q2_0 code -0.3%, story +3.4%; IQ3_XXS code +2.9%, story +2.3%. Prompts are within 2%.

New formats:

  • Q4_0 and Q4_1 experts (#599) on the GPU kernels, and a Q4_0 PLE table.
  • PLE tables in Q8_0, Q5_1 and BF16 (#651, #586, #865): one table of PLE formats now covers these and Q5_0 and FP8, with bounds checks on the file. Against the BF16 table, Q8_0 is off by 0.53%, FP8 by 2.64% and IQ4_NL by 7.60%. The effect on answers is small (KL about 0.012 for any of them), so setup does not offer a bigger table. --ple-io ram faults the table in on 16 threads (8 GB cold in WSL: 15 s -> 6 s).
  • Unsloth UD-Q6_K_XL: the Q8_0 PLE table is read, and the Q6_K expert kernels are an opt-in build (-DSTRATA_Q6K_EXPERTS=ON, #860).
  • IQ3_XXS and IQ4_XS PLE keys stay native, so those packs start again (#381).
  • --kv k8v4 streams with --kv-resident (#711, #705), and setup offers it.

AMD:

  • RDNA2 (#835, opt-in): STRATA_HIP_PROMPT_F16=1 runs the prompt's 16-bit GEMMs in FP16, and prompts read about 2x faster on an RX 6900 XT. It rounds differently, so it is off for now. The engine prints a tip on gfx103x.
  • hipBLASLt tables for gfx1100 (#766, #755, #565), RDNA3 / RDNA4 sibling arch names (#695), gfx1034 (#778) and a gfx906 build fix (#808).
  • Windows HIP (#873, #654): one hipFree(0) warm-up before the first memory query.
  • Opt-in: STRATA_HIP_ADAPT_KERNEL_COPY=1 does the adaptive swaps' copies with a kernel instead of SDMA (#884, for the 2x gfx1030 hang). STRATA_DENSE_MMQ=1 (#820) gave no gain on gfx1151 and stays opt-in.
  • Older cards: the Q4_0 expert kernel overflowed on sm_60 (signed overflow). It is fixed, found on a new P100 test box.
  • Intel Arc (#784, #866, #868, #870): the SYCL port compiles against 0.1.39's sources again, a ring wait over 2 ms is no longer reported as finished, and the Arc Pro B60 ids are known.

Linux cold start: weights, head, embedding, vision encoder and the RAM copy are now read ahead of use (#699). On an RTX 5090 with 32 GB, Q2_0 at 262K, ready in 70 s instead of about 920 s. STRATA_READ_AHEAD=0 turns it off. Linux also gets the unbuffered file tier (#773, O_DIRECT, same rule as Windows). On a 16 GB laptop with 30 GB RAM, decode went 4.4 -> 7.9 tok/s and startup 36 -> about 20 s (author's numbers). With transparent huge pages on always the arena no longer asks for more (#771, STRATA_NO_ARENA_THP=1).

Defaults that changed (the rest of the default output is unchanged):

  • Empty assistant turns (#886): they are not rendered into the next prompt, so a client's history can't teach the model to answer with nothing. STRATA_KEEP_EMPTY_TURNS=1 restores the old rendering.
  • tool_choice (#790): none offers no tools on both routes. required and a named function work.
  • json_object mode (#762): JSON is taken out of a code fence or the prose around it.
  • Stop strings (#454): OpenAI stop and Anthropic stop_sequences are honoured on every path.
  • Workers on hybrid CPUs (#798): P and E cores come from the CPU's own lists. A Core Ultra 7 270K Plus is now 8P + 16E, 15 workers.
  • Windows (#691): the server is opted out of power throttling, so a minimized window no longer moves the engine to the E-cores.
  • Conversation caches from 0.1.39 are rebuilt on first use: the cache's version key changed, so the first request after the update reads its prompt again.
  • Setup's context menu has a 200K step (#608). It is never the recommended one.

Opt-ins (off by default, the output is unchanged):

  • Server: "strata_checkpoint": false for a one-shot request (#861), "return_progress": true (#837), "reasoning_loop_recovery": "stop" | "recover" (#869), session files with --slot-save-path (#668), --prompt-cache-tail (#614), STRATA_CACHE_MESSAGE_BOUNDARY=1 (#734).
  • Memory: --kv-grow, the K/V takes VRAM as the context grows (#378).
  • Speed: STRATA_Q2_BITPLANE=1 (#706), STRATA_PREFILL_EQUAL=1 (#693), STRATA_KV_PREFETCH=1 (#732, 4-5% slower on our 5070, so off), STRATA_IQ3S_MT1=1 (#863), STRATA_DMA_BATCH=1|2 (#807), STRATA_ADAPT_LAG=2 (#764).
  • Multi-GPU: --adapt-async 1 (#876), STRATA_EXCHANGE_ROTATE=1 (#864), --pipeline-windows (#859), STRATA_DISJOINT_ADAPT=1 (#731).
  • Tools: python setup.py --rollback-engine puts the previous engine back (#670), and strata_pack build --force rebuilds into a used folder (#634).
  • Turing prompts above about 90K cells take the wide top-k kernel (#743, STRATA_TOPK_STREAM=0 restores the old one).
  • STRATA_DEFERRED_REGISTER=1 (Windows, #285 part 2) registers the arena per layer on a thread ahead of the readers. It cost about 8% decode here, so it is off.

Testers wanted (we have one RTX 5070 and no second card):

  • 2 GPUs: a 30-minute soak of the resident RAM mode on a split (#848, #642), --batch on an all-resident stage (#776), and --pipeline-windows with 10 or more interleaved pairs, with and without #848 (#859).
  • 4 GPUs: the new automatic placement (#887) and --layer-split auto with the trim (#880). Tell us if the auto layout got worse.
  • Turing (RTX 20): a prompt at 90K, 131K and 155K, before and after (#743).
  • Intel Arc: the SYCL port on an A- or B-series card (#867, #870).
  • Older cards: Pascal and Volta owners with a result from the CUDA 12 engine. We now have a P100 for sm_60, but not Volta.
  • Windows AMD: the HIP zip's gfx1151 code on a Strix Halo PC, STRATA_HIP_PROMPT_F16=1 on gfx103x (#835), the SDMA kernel copy on 2x gfx1030 (#884) and the HIP warm-up (#654).
  • Open reports where a retry helps: #879 (both reporters), #871 at 100% residency, #828 (MALLOC_CHECK_=3) and #795 (BIOS microcode update).

Thanks to everyone who sent PRs, tests and reports. Special thanks to Eddoursul, whose fork shaped this release, to Francesco Albano for #848, to sergiywith for #838 and #550, and to StevenChenSE for the Strix Halo GEMM work.

Checked before the release:

  • The same answers as 0.1.39 on all four quants (Q2_0, IQ3_XXS, IQ3_S, the Coder), 10/10 each with a fixed cache, on the final tip. The prompt path's internal state is identical (exact_pp). A Tesla P100 (sm_60) gave 10/10 on 3 quants.
  • Speed: the decode numbers above, interleaved against 0.1.39 on the RTX 5070.
  • Strix Halo: output is byte-identical to the earlier branch.
  • AMD and Intel: the HIP engine compiles, and the Windows HIP zip builds with gfx1151 (598 MiB). Nothing here ran on an AMD or Arc card except gfx1151 on Linux.

Updating: run UPDATE.bat (Linux: ./update.sh). Setup installs engine 0.1.40.

The ready-made Strata engines for Windows, which START-HERE.bat fetches by itself:

  • strata-windows-x64.zip: NVIDIA (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0, needs an NVIDIA driver 580 or newer. Contents: strata.exe, strata-vision.exe (the optional image encoder), BUILD.json.
  • strata-windows-x64-cuda12.zip (experimental): older NVIDIA cards (Pascal and Volta: sm_60, sm_61, sm_70; it also has sm_75, sm_80, sm_86, sm_89 + PTX for a mixed PC), CUDA 12.9, needs an NVIDIA driver 528 or newer. Setup fetches it only for a model you put on such a card, or with --cuda 12. Contents: strata.exe, strata-vision.exe, BUILD.json.
  • strata-windows-x64-hip.zip: AMD (gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030, gfx1151), ROCm 10.2.0a20260930 from AMD's TheRock builds, needs a current AMD driver. gfx1151 (Strix Halo) is in this zip too, experimental and untested on Windows hardware. Contents: strata.exe, strata-device.exe, the HIP runtime next to them, BUILD.json, rocm\ (the ROCm libraries and their licenses).

Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going:

Buy Me A Coffee

Don't miss a new Strata release

NewReleases is sending notifications on new releases.