Faster and fixed across the board: official Intel Arc support (Arc Pro B70 and A750 tested), bug fixes and speed-ups merged from the open issues and PRs, and big gains on multi-GPU, low-RAM Linux and long-document use. Default answers are byte-identical to 0.1.40 (checked on Q2_0, IQ3_XXS, Coder and IQ3_S). See Updating below if your copy is older than 0.1.40.1.
Speed
Measured against 0.1.40 in interleaved pairs (medians, 10 pairs unless noted):
| Machine | Model | Decode (story / code) | Prompt 4K / 20K |
|---|---|---|---|
| RTX 5070 12 GB, Windows | IQ3_XXS | 48.5 -> 50.3 (+3.9%, 10/10 pairs) / 53.1 -> 53.5 | equal |
| RTX 5070 12 GB, Windows | Q2_0 | 74.8 -> 75.6 / 79.3 -> 79.4 | 1377 -> 1387 / 1485 -> 1489 |
| RTX 3060 12 GB, Linux | IQ3_XXS | 36.3 -> 37.0 (+1.7%) / 45.5 -> 47.5 (+4.3%) | equal |
| Tesla P100 16 GB, Linux | IQ3_XXS | 27.6 -> 28.6 (+3.7%) / 31.5 -> 32.3 (+2.5%, 10/10 pairs) | equal |
Where it comes from, and the larger gains in particular setups:
- 2-4 token verify windows read each weight block once (on by default, RTX 30 series and newer). The dense projections of a draft window decode a weight block once and apply it to every token, from an interleaved copy of the activations. Bit-identical output. RTX 3060 +4.2% decode (12/12 pairs), RTX 5070 +1.3 to +2.2%.
STRATA_MMVQ_IL=0turns it off. The idea comes from Eddoursul's fork. - Multi-GPU layer split: up to +60% prompt speed. When two placements scored the same, the auto split could pick the unbalanced one. On 4x Radeon AI PRO R9700 a 32K prompt now reads at 2,797 tok/s (3 GPUs: 1,970 -> 2,458).
- Linux with less RAM than the model: +20 to +42% decode (#1194). The file tier now reads through the page cache when it can hold a useful share of the experts, instead of unbuffered reads. P100 box under a 32 GB limit: story 12.6 -> 15.2, code 9.3 -> 13.3 tok/s, with 3-7x less drive traffic. Same tokens.
STRATA_UNBUFFERED_LOAD=1/0forces either. - Many questions about one long document: first answer in 0.35-0.55 s instead of 48-149 s (opt-in
pin=N). A pinned shared prefix is read once and kept; each question only reads its own part. See docs/RESEARCH_RUNS.md. - DGX Spark / Linux with --mmap-experts: prompts up to 10x faster (#1057). The prompt stager's threads sleep instead of spinning (IQ3_S 8K 91 -> 1,222 tok/s on a GB10). Default on Linux; Windows keeps the old wait (it was 1.2% faster there).
STRATA_STAGER_SLEEP=1/0forces either. Thanks to uncle daddy. - Short prompts 19-35% faster with the CPU helping (opt-in,
STRATA_PREFILL_CPU_SHARE=auto) (#1282): for prompt chunks under 1,024 tokens (typical chat turns and tool calls), the CPU computes the least-routed experts while the GPU does the rest. 512-token prompt: RTX 3060 1.62 -> 1.16 s, P100 4.8 -> 3.1 s, RTX 5070 1.38 -> 1.02 s; 2K and longer unchanged. The answers are not byte-identical (CPU and GPU round differently: mean KL 0.006, the same top token in 19 of 22 prompts), so it is off by default. Thanks to sergiywith. - Opt-in SSD read-ahead on Linux (
STRATA_IO_PREFETCH=1): reads the predicted experts of the next layers while the current one computes;STRATA_IO_PF_STAGE=1measured +25% on an RTX 3060 under a 32 GB limit, but slower at 16 GB and on an iGPU, so it stays off by default. docs/DETAILS.md has the table. New per-request line: what the OS read from the drive against the page cache. - AMD: RX 7900 XTX hipBLASLt table rows for small prompt chunks (+32% prompt, measured by the author, #1289); docs for
STRATA_HIP_WMMA=1on the R9700 (+27-31% prompt, output bits change, so opt-in) and a start hint on gfx12 with--kv int8. - Opt-in sampled drafting: probabilistic draft acceptance (
STRATA_SPEC_PROB,STRATA_SPEC_COUPLED) and Gumbel-max coupled drafts (STRATA_SPEC_GUMBEL=1, +5.7 to +9% decode on AMD in the reporters' runs; on an RTX 3060 acceptance rose about a point and decode stayed within noise).
Intel Arc: officially supported
Strata now officially supports Intel Arc graphics cards on Linux: its own engine ported to SYCL, behind the same server, APIs and web app. Install with python3 sycl/setup_intel.py (it runs the oneAPI build in a Docker image; see docs/INTEL.md). Tested on an Arc Pro B70 (32 GB, xe driver) and an Arc A750 (8 GB, i915 driver): the 10-prompt check with MTP passes on both (B70: 10/10 twice; A750: 6 of 6 runs), and neither card had a GPU hang.
| Card | Model | Prompt (4K tokens) | Decode (story / code) |
|---|---|---|---|
| Arc Pro B70 32 GB | IQ3_S | 980-1,000 tok/s | 31 / 41 tok/s |
| Arc A750 8 GB | IQ3_XXS | about 58 tok/s | about 16 tok/s |
What it took (0.1.40 did not build for Intel at all):
- Wrong answers on the B70: the xe driver could hand back a large GPU allocation with aliased pages (a write at one offset showed up 1 GiB further on), so cached experts held other experts' bytes. Every big allocation is now checked at start and re-allocated until clean (
STRATA_ARENA_ALIAS_CHECK=0turns it off). - B70 hang: without
--stream-expertsthe GPU was handed pageable host memory, which the xe driver faults on.--pcie-fracnow defaults to 0 in that case, with a warning; setup configures--stream-experts. - B70 prompt speed: setup asks a 24 GB+ Arc card for 4,096-token prompt chunks, and there is no short first chunk while part of the experts are in host memory (4K prompt 618 -> about 1,000 tok/s).
- MTP drafts are accepted again (4% -> about 60% on the A750): a wait read a draft token before the GPU had written it.
- The first-window hang on i915 cards (A-series) is fixed: the queues drain before each window.
- Free VRAM on i915 is now this process's real use, and the sampler no longer uses FP64, which A-series cards emulate (A750 decode +6%).
Intel Arc is Linux only for now.
Fixes
- Updating after the history cleanup (#1276):
UPDATE.bat/update.shnow move an old clone to the new history by themselves when you have not edited Strata's files (your old commits stay in a branchpre-cleanup-backup). With edits, they print the commands and change nothing. - k8v4 KV with --kv-resident printed every number twice in long counts (#1188, #1135, #1169).
- --batch slots hung during a prompt loan on fully resident multi-GPU setups (#1118).
- STRATA_QFUSE with --batch gave garbage (#1139).
- Host OOM at start on Linux when RAM was mostly file cache: the cache of the model files is given back first, then pinning is paced (#1250).
- The server no longer waits for ever on an engine or image encoder that stopped answering (#1317): 900 s for the engine's start, 300 s for the encoder;
STRATA_ENGINE_READY_Setc. change them (0 = no limit). - A restart could report the old engine's context size for a moment; fixed.
- /v1/responses accepts a JSON
text.formattogether with tools (Codex, #782). - Setup: ROCm wheels chosen per GPU family (gfx103X and gfx1151 get versions that work, #1103, #1267); CUDA 13.2.2 is accepted (13.2.0/13.2.1 still warned on RTX 50: they miscompile IQ3_S/IQ2_S); the downloaded engine is checked against GitHub's SHA-256 (refused on a mismatch, a warning when no digest is available,
STRATA_SKIP_SHA256=1skips it, #1218). - Also:
--gpu LISTselects the visible GPUs (#852); Prometheus/metricsin the text format (#793); LRU graph eviction on VRAM OOM (#1185); docker REINSTALL=0 keeps the /data config (#1244); EXIF orientation of uploaded images (#1229);--kv-growhold as an opt-in (STRATA_KV_GROW_HOLD=1, #1128); 14 community benchmark reports added to bench/results.
Release builds use CUDA 13.0 (as 0.1.40). The Windows AMD zip includes gfx1151 (Strix Halo).
Updating
From 0.1.40.1 or newer: UPDATE.bat (Linux: ./update.sh) as usual.
From 0.1.40 or older, once by hand (the old update script cannot follow the cleaned-up history). In your Strata folder:
git fetch origin
git branch pre-cleanup-backup
git checkout -B main origin/main
Your models, configs and the engine are not touched. Then run UPDATE.bat / ./update.sh to get the new engine.
Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going: