github Niko1221/Strata v0.1.40.2
Strata v0.1.40.2

5 hours ago

Faster and fixed across the board: official Intel Arc support (Arc Pro B70 and A750 tested), bug fixes and speed-ups merged from the open issues and PRs, and big gains on multi-GPU, low-RAM Linux and long-document use. Default answers are byte-identical to 0.1.40 (checked on Q2_0, IQ3_XXS, Coder and IQ3_S). See Updating below if your copy is older than 0.1.40.1.

Speed

Measured against 0.1.40 in interleaved pairs (medians, 10 pairs unless noted):

Machine Model Decode (story / code) Prompt 4K / 20K
RTX 5070 12 GB, Windows IQ3_XXS 48.5 -> 50.3 (+3.9%, 10/10 pairs) / 53.1 -> 53.5 equal
RTX 5070 12 GB, Windows Q2_0 74.8 -> 75.6 / 79.3 -> 79.4 1377 -> 1387 / 1485 -> 1489
RTX 3060 12 GB, Linux IQ3_XXS 36.3 -> 37.0 (+1.7%) / 45.5 -> 47.5 (+4.3%) equal
Tesla P100 16 GB, Linux IQ3_XXS 27.6 -> 28.6 (+3.7%) / 31.5 -> 32.3 (+2.5%, 10/10 pairs) equal

Where it comes from, and the larger gains in particular setups:

  • 2-4 token verify windows read each weight block once (on by default, RTX 30 series and newer). The dense projections of a draft window decode a weight block once and apply it to every token, from an interleaved copy of the activations. Bit-identical output. RTX 3060 +4.2% decode (12/12 pairs), RTX 5070 +1.3 to +2.2%. STRATA_MMVQ_IL=0 turns it off. The idea comes from Eddoursul's fork.
  • Multi-GPU layer split: up to +60% prompt speed. When two placements scored the same, the auto split could pick the unbalanced one. On 4x Radeon AI PRO R9700 a 32K prompt now reads at 2,797 tok/s (3 GPUs: 1,970 -> 2,458).
  • Linux with less RAM than the model: +20 to +42% decode (#1194). The file tier now reads through the page cache when it can hold a useful share of the experts, instead of unbuffered reads. P100 box under a 32 GB limit: story 12.6 -> 15.2, code 9.3 -> 13.3 tok/s, with 3-7x less drive traffic. Same tokens. STRATA_UNBUFFERED_LOAD=1/0 forces either.
  • Many questions about one long document: first answer in 0.35-0.55 s instead of 48-149 s (opt-in pin=N). A pinned shared prefix is read once and kept; each question only reads its own part. See docs/RESEARCH_RUNS.md.
  • DGX Spark / Linux with --mmap-experts: prompts up to 10x faster (#1057). The prompt stager's threads sleep instead of spinning (IQ3_S 8K 91 -> 1,222 tok/s on a GB10). Default on Linux; Windows keeps the old wait (it was 1.2% faster there). STRATA_STAGER_SLEEP=1/0 forces either. Thanks to uncle daddy.
  • Short prompts 19-35% faster with the CPU helping (opt-in, STRATA_PREFILL_CPU_SHARE=auto) (#1282): for prompt chunks under 1,024 tokens (typical chat turns and tool calls), the CPU computes the least-routed experts while the GPU does the rest. 512-token prompt: RTX 3060 1.62 -> 1.16 s, P100 4.8 -> 3.1 s, RTX 5070 1.38 -> 1.02 s; 2K and longer unchanged. The answers are not byte-identical (CPU and GPU round differently: mean KL 0.006, the same top token in 19 of 22 prompts), so it is off by default. Thanks to sergiywith.
  • Opt-in SSD read-ahead on Linux (STRATA_IO_PREFETCH=1): reads the predicted experts of the next layers while the current one computes; STRATA_IO_PF_STAGE=1 measured +25% on an RTX 3060 under a 32 GB limit, but slower at 16 GB and on an iGPU, so it stays off by default. docs/DETAILS.md has the table. New per-request line: what the OS read from the drive against the page cache.
  • AMD: RX 7900 XTX hipBLASLt table rows for small prompt chunks (+32% prompt, measured by the author, #1289); docs for STRATA_HIP_WMMA=1 on the R9700 (+27-31% prompt, output bits change, so opt-in) and a start hint on gfx12 with --kv int8.
  • Opt-in sampled drafting: probabilistic draft acceptance (STRATA_SPEC_PROB, STRATA_SPEC_COUPLED) and Gumbel-max coupled drafts (STRATA_SPEC_GUMBEL=1, +5.7 to +9% decode on AMD in the reporters' runs; on an RTX 3060 acceptance rose about a point and decode stayed within noise).

Intel Arc: officially supported

Strata now officially supports Intel Arc graphics cards on Linux: its own engine ported to SYCL, behind the same server, APIs and web app. Install with python3 sycl/setup_intel.py (it runs the oneAPI build in a Docker image; see docs/INTEL.md). Tested on an Arc Pro B70 (32 GB, xe driver) and an Arc A750 (8 GB, i915 driver): the 10-prompt check with MTP passes on both (B70: 10/10 twice; A750: 6 of 6 runs), and neither card had a GPU hang.

Card Model Prompt (4K tokens) Decode (story / code)
Arc Pro B70 32 GB IQ3_S 980-1,000 tok/s 31 / 41 tok/s
Arc A750 8 GB IQ3_XXS about 58 tok/s about 16 tok/s

What it took (0.1.40 did not build for Intel at all):

  • Wrong answers on the B70: the xe driver could hand back a large GPU allocation with aliased pages (a write at one offset showed up 1 GiB further on), so cached experts held other experts' bytes. Every big allocation is now checked at start and re-allocated until clean (STRATA_ARENA_ALIAS_CHECK=0 turns it off).
  • B70 hang: without --stream-experts the GPU was handed pageable host memory, which the xe driver faults on. --pcie-frac now defaults to 0 in that case, with a warning; setup configures --stream-experts.
  • B70 prompt speed: setup asks a 24 GB+ Arc card for 4,096-token prompt chunks, and there is no short first chunk while part of the experts are in host memory (4K prompt 618 -> about 1,000 tok/s).
  • MTP drafts are accepted again (4% -> about 60% on the A750): a wait read a draft token before the GPU had written it.
  • The first-window hang on i915 cards (A-series) is fixed: the queues drain before each window.
  • Free VRAM on i915 is now this process's real use, and the sampler no longer uses FP64, which A-series cards emulate (A750 decode +6%).

Intel Arc is Linux only for now.

Fixes

  • Updating after the history cleanup (#1276): UPDATE.bat / update.sh now move an old clone to the new history by themselves when you have not edited Strata's files (your old commits stay in a branch pre-cleanup-backup). With edits, they print the commands and change nothing.
  • k8v4 KV with --kv-resident printed every number twice in long counts (#1188, #1135, #1169).
  • --batch slots hung during a prompt loan on fully resident multi-GPU setups (#1118).
  • STRATA_QFUSE with --batch gave garbage (#1139).
  • Host OOM at start on Linux when RAM was mostly file cache: the cache of the model files is given back first, then pinning is paced (#1250).
  • The server no longer waits for ever on an engine or image encoder that stopped answering (#1317): 900 s for the engine's start, 300 s for the encoder; STRATA_ENGINE_READY_S etc. change them (0 = no limit).
  • A restart could report the old engine's context size for a moment; fixed.
  • /v1/responses accepts a JSON text.format together with tools (Codex, #782).
  • Setup: ROCm wheels chosen per GPU family (gfx103X and gfx1151 get versions that work, #1103, #1267); CUDA 13.2.2 is accepted (13.2.0/13.2.1 still warned on RTX 50: they miscompile IQ3_S/IQ2_S); the downloaded engine is checked against GitHub's SHA-256 (refused on a mismatch, a warning when no digest is available, STRATA_SKIP_SHA256=1 skips it, #1218).
  • Also: --gpu LIST selects the visible GPUs (#852); Prometheus /metrics in the text format (#793); LRU graph eviction on VRAM OOM (#1185); docker REINSTALL=0 keeps the /data config (#1244); EXIF orientation of uploaded images (#1229); --kv-grow hold as an opt-in (STRATA_KV_GROW_HOLD=1, #1128); 14 community benchmark reports added to bench/results.

Release builds use CUDA 13.0 (as 0.1.40). The Windows AMD zip includes gfx1151 (Strix Halo).

Updating

From 0.1.40.1 or newer: UPDATE.bat (Linux: ./update.sh) as usual.

From 0.1.40 or older, once by hand (the old update script cannot follow the cleaned-up history). In your Strata folder:

git fetch origin
git branch pre-cleanup-backup
git checkout -B main origin/main

Your models, configs and the engine are not touched. Then run UPDATE.bat / ./update.sh to get the new engine.


Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going:

Buy Me A Coffee

Don't miss a new Strata release

NewReleases is sending notifications on new releases.