github Niko1221/Strata v0.1.43
Strata v0.1.43

4 hours ago

Faster decode and prompts on NVIDIA, AMD and Intel, plus the fixes since 0.1.42. Decode is up to 38% faster on a 16 GB Radeon, 5-11% on RTX 30-series and Tesla P100 under Linux, up to 19% on four Radeons, and 14-21% on an Arc Pro B70, whose prompts read about twice as fast. Windows streaming from GGUF files decodes 1.35-1.85x faster. Answers are the same as 0.1.42 (byte-identical on every machine we tested) except where a default below says otherwise. Update with UPDATE.bat (Linux: ./update.sh); setup replaces the engine with 0.1.43.

Speed

All numbers are our own interleaved A/Bs against the previous build: 5 pairs (prompts on the B70: 3), medians, the same output tokens in both arms unless a line says so.

Card Model Decode Prompts
Radeon AI PRO R9700 with a 16 GB card's cache IQ3_XXS +21% to +38% 4K +7.5%, 32K +15.6%
Radeon AI PRO R9700 32 GB Q2_0 / IQ3_XXS +2% to +10% +12% to +14%
Radeon AI PRO R9700 32 GB IQ3_S +6% to +14% 32K +17%, 64K +27%
4x R9700 (cards at 225 W) IQ3_S story +3.8%, code +12.6%, a 4K-token document +18.7%
RTX A4000 (Linux) Q2_0 story +6.6%, code +5.5%, a 6K document +5.1%; the new Q3_K kernel adds +2.0-2.7% on top flat
Tesla P100 (Linux) IQ3_XXS story +7.0%, code +10.7%, a 6K document +8.0% flat
RTX 5070 (Windows) Q2_0 / IQ3_XXS the same as 0.1.42 by default (the two NVIDIA decode changes below are opt-in on Windows) flat
Arc Pro B70 IQ3_S +14% to +21% (new expert ranking) 4K 8.5 -> 3.3 s, 16K 27.4 -> 13.4 s (128K config)
  • AMD Radeon (gfx1200 / gfx1201, RDNA4). The two decode defaults NVIDIA got in 0.1.42 now apply here too, measured on the R9700: the route tail skip (STRATA_ROUTE_TAIL_SKIP=0 turns it off) and the measured PCIe share (STRATA_PCIE_FRAC_DEFAULT=old). Like on NVIDIA they change the answers slightly. The fused prompt kernels are the default for IQ3_S and IQ3_XXS (STRATA_PF_FUSED=0 turns them off), and long prompts read in 16K chunks on Linux (STRATA_GFX12_CHUNK=0). Fewer and fused graph nodes in the decode window give another 1.6-2.0%, with the same answers. On a layer split the tail skip stays off from 60% of the experts in VRAM. RDNA3 (gfx11) and older keep their 0.1.42 behaviour: we have no such card to measure on.
  • 3 or more Radeon cards. --pipeline-windows 2 is on by default for a gfx12 layer split of 3 or more cards in --serve (4x R9700: code +7-12%, 4K prompt +9-11%, 32K +8-19%). The pipelined windows hand the CPU differently sized batches of expert rows, so the CPU-side sums can round differently and answers can differ slightly from the serial order (on 4x R9700 with default settings 10 of 10 test prompts matched the serial order over 160 tokens). With 2 cards it stays opt-in: there it loses about 3% on code. --pipeline-windows 0 or STRATA_PIPELINE_AUTO=0 turns it off.
  • NVIDIA. On Linux the adaptive tier's copies overlap the next decode window (STRATA_ADAPT_OVERLAP=0 turns it off; on Windows it is opt-in with =1, since an RTX 5070 there decoded Q2_0 2.6% slower with it). The Tesla P100 gets its own kernel table (STRATA_SM60_TABLE=0). RTX 30-series cards run the Q2_0 pack's Q3_K weights on the interleaved kernels (STRATA_MMVQ_IL_Q3K=0). On Linux the canonical Q2_0 pack uses the measured PCIe share too (on Windows with STRATA_PCIE_AUTO_CANON=1). All of these give the same answers. The opt-in STRATA_ROUTE_RESIDENT runs one warp per token on NVIDIA and AMD (#1737, aly8246; measured on an R9700): +2% decode with a 32 GB cache, +6-8% with a 16 GB card's, same answers as before.
  • Intel Arc Pro B70 (Battlemage). The prompt path's expert copies run on the card's copy engine (STRATA_COPY_ENGINE=0), the host KV copy is written 4 bytes at a time, the A770's GEMM staging is skipped, and the prompt loan's experts refill from pinned memory. A new expert ranking (data/expert-profile-sycl.bin, written by setup for cards on the xe driver) holds 71% of a chat's experts instead of 60%. Same answers.

Fixes

  • Windows, GGUF in place: experts are read in parallel again (reported on X). With the model's GGUF shards left mapped in memory, NTFS runs the unbuffered reads of those files one at a time, so a low-memory setup decoded at 2.6 tok/s. The engine now closes every GGUF shard's view once the file tier reads unbuffered. STRATA_KEEP_MAPPING=1 brings the old behaviour back. Measured on an RTX 5070 (Windows, Ryzen 5 7600, 64 GB) with IQ3_XXS read from its GGUF files, 5 interleaved pairs each:

    Experts held in RAM Story Code Short prompt
    22 GiB 18.9 -> 25.5 tok/s (1.35x) 13.5 -> 21.1 (1.56x) 13.3 -> 18.2 (1.37x)
    12 GiB 9.6 -> 16.0 (1.67x) 6.5 -> 11.6 (1.78x) 4.6 -> 8.5 (1.85x)
  • Dual-socket and multi-node Linux hosts (reported on X). The page-locked expert copy is judged by the emptiest NUMA node too, which only makes it load in steps; it never refuses and never caps. The "not enough RAM" start failure now says how much of the RAM is file cache and prints the fix. New and off by default: STRATA_DROP_CACHE_ON_EXIT=1 makes strata serve give the model files' cache back when it stops.

  • The route tail skip stays off when nearly every expert is in VRAM (#1884, christopherrobertbrooks-tech). On a Tesla V100 32 GB holding 94.8% of the experts 0.1.42 decoded 1-2% slower. The default now stays off from 90% of the experts in the GPU cache (60% on an AMD layer split), and the start-up log says so; those cards get 0.1.41's answers back. STRATA_ROUTE_TAIL_SKIP=7 turns it on anyway; STRATA_TAIL_SKIP_MAX_CACHED=<percent> moves the line.

  • Timing numbers in the API (#1869, #1870, sdjger-xiaoniu). A request that went from a batch slot back to the single-request path now reports the whole request, and timings.prompt_n counts what the engine actually read, with the rest in cache_n.

  • A text prompt right after a picture could fault on RX 7000 cards (#1712, mantovaniluca91, who traced it). The prompt chunk's step records and token ids now go up from pinned memory, and a failed upload is reported instead of ignored.

  • Answers of only ! are caught (#879, #1815; the guard is #1892 by lask3802). When a decode window's logits come out all NaN, nothing is sent, the caches are dropped, the prompt is read again from token 0 and the server retries once; the engine logs non-finite logits (#879). Without a NaN the answers are byte-identical, and it costs at most 0.2% decode. STRATA_NAN_GUARD=0 turns it off. The cause on the reporters' machines is still open: we could not make it fire here (1,870 requests on an R9700).

New

  • BENCH.bat / ./bench.sh (python tools/strata_bench.py). Runs a fixed suite against your own install in about 3 to 10 minutes: short chat, 4K and 16K prompts, a five-turn agent session, optional --clients N and --long. It writes a report in the community-benchmark format, scrubs home folders, names and addresses, and --compare puts two runs side by side. Setup offers it at the end of an interactive install (default no). See docs/COMMUNITY_BENCHMARKS.md.

Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going:

Buy Me A Coffee

Don't miss a new Strata release

NewReleases is sending notifications on new releases.