Faster decode and long prompts, several requests at once, older and other hardware as experimental opt-ins, the
OpenAI Responses API for Codex, a fix for slower prompts with a RAM budget, and a batch of fixes from your reports.
Faster (the same answers where nothing says otherwise):
- Decode (#646): the verify pass runs with fewer launches and host round trips (sub-warp expert packing, staged
inputs, the PLE and MTP steps batched). Measured here on one RTX 5070 against 0.1.38, 10 interleaved pairs, medians:
Q2_0 +6% (story 68.4 -> 72.7 tok/s, code 75.8 -> 80.0), IQ3_XXS +6% / +2.5% (46.1 -> 48.8, 50.7 -> 51.9); after
4K and 32K prompts +5% to +7%. The output is byte-identical to 0.1.38's (10/10 on all four quants with a fixed
cache). The zero-doorbell verify graph itself runs only when every expert of a layer is in VRAM. A 12 GB card never
gets there, so the gain on big cards was not measured here. - Long prompts (#583, on by default): the streamed expert ring is sized in bytes for the pack, and
--prefill auto
picks the largest chunk that keeps it full. Short prompts, and prompts that fit 0.1.38's chunk, keep 0.1.38's ring
and their exact bits (4K: identical). The gain depends on the free VRAM. On an RTX 5070 with 32K prompts, IQ3_XXS is
+18.5% with a 1,500-slot expert cache and unchanged with setup's default config. A long prompt's bits change against
0.1.38, because the experts go through a different mix of cached and streamed groups. Its quality stays in the same
band, teacher-forced over the next 2,001 tokens against the FP16 prompt path (IQ3_XXS): 8K prompt KL 0.042 (0.1.38:
0.054), top-1 93.5% (92.2%); 32K prompt KL 0.020 (0.019), top-1 95.3% (95.0%). It doesn't help every pack: the
Coder at a 32K context reads a 30K prompt about 8% slower (with a 64K context it was about 6% faster).
STRATA_RING_BYTES=0restores 0.1.38's ring. - Very long contexts: the attention's top-k past the register kernel's reach (262K-524K cells) runs with a
histogram per warp: a 243K-token prompt +26% on an RTX 3060 (#603). - More than one GPU (measured by their authors, not here: we have one GPU): setup now adds
--remote-expert-opt(#578) to a config with two or more GPUs. It skips the host's work for tokens whose experts
all run on the helper cards (dual RTX 4090: +63% mixed text, +132% code over the plain helper path), andsetup --no-remote-expert-optleaves it out. A layer split of 3+ GPUs overlaps its stages better (#598, 4x RTX 3060: +10
to 28% on prompts).STRATA_PREFILL_HELP=1lets a split's idle card stream a share of a one-chunk prompt's experts
(#663). It is opt-in because it rounds differently. With one or two GPUs the pipeline runs as before. - Also: the Linux expert arena on transparent huge pages (#650;
STRATA_NO_LARGEPAGES=1keeps 4 KB pages), an
opt-in AVX2 codebook gather for the IQ CPU kernels (STRATA_IQ256_GATHER=1, #622), and thread affinity on PCs with
more than 64 CPUs (#626).
Several requests at once (#465, opt-in): "parallel": N in strata-<model>.json (or setup --parallel N)
decodes up to N conversations together in the engine. More requests wait for a free slot, and a long prompt gives
way to a short one at a chunk boundary. A request left alone goes back to the one-at-a-time path, and /metrics
shows each slot. Every slot's greedy answer is the same as when it runs alone. On a 12 GB card it cuts waiting time
and costs speed. RTX 5070, Q2_0, four requests at once: the last one starts after 1.8 s instead of 11.2 s, but
together they decode 11% slower (63.1 against 70.7 tok/s). A request alone loses 11% with 2 slots and 22% with 4,
because each slot's session takes 0.56 GiB from the expert cache. So setup recommends it only where the experts
mostly fit in VRAM. A 4-GPU layer split (4x 16 GB, IQ3_S, PR #559) served 8 requests at 360 tok/s in total against
120 for one. docs/BATCHING.md has the numbers and the options.
Older and other hardware, experimental: three opt-in paths that community members wrote and measured on their
own machines. We have none of this hardware. Each is compile-checked and unit-tested here, and the ready-made
engines and their output are unchanged (checked byte-identical).
- Older NVIDIA cards (Pascal, Volta: P40, P100, GTX 10, V100, Titan V; #395 #600 #540 #655 #627): CUDA 13 cannot
compile for them, so setup keeps a second engine built with CUDA 12.9 (strata-windows-x64-cuda12.zipon
Windows, compiled on Linux). It is used only when you choose such a card: a PC with only Pascal / Volta cards, a
card named with--gpu N/--gpus, or--cuda 12.--cuda 12is also the way to run with an NVIDIA driver
older than 580 (528+ on Windows, 525+ on Linux). The choice is kept per model. The Volta prompt attention and the
BF16 path through FP16 / fp32 are compiled into this engine only. Untested here: every Pascal and Volta card, the
old drivers, and an RTX 50 card in that engine (it gets a warning; keep it on its own model). docs/OLDER_GPUS.md
has the reporters' numbers (V100: prompts 1,123-1,251 tok/s, UD-IQ4_XS). RTX 20 owners can try the same FP16
tensor-core path for the prompt's BF16 products withSTRATA_BF16_TC=1(+15-18% on an RTX 2080 Ti, #655). Its
sums are not bitwise cuBLAS's.
Older AMD cards are built by hand: gfx906 (Instinct MI50 / MI60, Radeon VII;-DSTRATA_HIP_GFX906=ON, #638 #677:
2x MI50, the Coder, 50 tok/s at 4K) and gfx1012 (RX 5500 XT, #442). The RX 6700 XT (gfx1031, #524) goes through
setup. - Intel Arc (#423): maxfridbe's SYCL port of the engine, built from source on Linux with Intel oneAPI:
./setup.sh --backend sycl. Reported on the Arc Pro B70 / B50 and the B580 (Coder IQ1_M 70-78 tok/s on a B70) on
earlier versions. Here it compiles (oneAPI 2026.1) and its kernel tests run on a CPU device. Untested: the 0.1.39
port on an Arc, Windows (no build path yet; setup points at Linux), WSL2, the A-series and integrated Arc GPUs, and
images. There is no ready-made Intel engine. docs/INTEL_ARC.md. - Older CPUs without AVX2 (#394 #595): AVX-only (Sandy / Ivy Bridge, Xeon E5 v1/v2, Bulldozer) and
SSE4.2-only CPUs (Nehalem / Westmere). Run setup as usual: it warns and compiles the engine for that CPU
(STRATA_ISA_FLOOR=avxornone, 10-20 minutes once) instead of stopping. Only the i-quant models run there, and
the CPU's share is slow (forced on our Ryzen: 11-17 tok/s with the AVX build, 3.5-3.8 with SSE4.2, against 26).
Untested here on a real old CPU; contributors ran earlier versions of it on Xeon E5-2680 / E5-2687W / X5690.
docs/INSTALL.md#older-cpus-experimental.
Prompts with a RAM budget are fast again (#577): 0.1.38 decided too early whether to read the experts past the
file cache, and it compared against the size of every model file. On a 96 GB PC with UD-Q4_K_XL and a 72 GiB budget
it chose wrong, so every refill after a prompt read the drive (prompts 15-40% slower than 0.1.34). The choice now
counts only the expert bytes outside the RAM copy, and is made again once that copy is built. The tokens are the same
either way.
The OpenAI Responses API (#451): POST /v1/responses, so Codex CLI works with Strata (tested with Codex 0.160.0,
a tool loop included; later turns reused ~96% of the prompt from the cache). It is stateless, as Codex uses it, and
covers function tools, tool results, reasoning effort, JSON schemas and the streaming events. Not supported:
previous_response_id (nothing is stored), hosted tools and reasoning summaries. The Codex config.toml is in
docs/DETAILS.md.
Security: SECURITY.md says how to report a problem privately (GitHub's private vulnerability reporting) and what
the server exposes: 127.0.0.1 by default, the API key, the Host and Origin checks of 0.1.38, CORS, and the opt-in MCP
tools and request monitor.
Fixes from your reports:
- A reply stuck on one token is ended (#606): a reply that repeats one token 256 times ends there
("repeat_stop_tokens"sets the length,0turns it off), and the q8_1 activations stay finite. - Starting on a tight card (#620): the native head and the logits are loaded before the expert arena, and a failed
allocation names the free VRAM. - The RAM check before the arena (#633): a container whose memory limit is below the arena is refused with the
numbers instead of being killed during the load. Less RAM available only warns. - A rotational disk (#605):
--ple-io directwarns at start, setup keeps the n-gram table in RAM there, and the
stall report counts disk waits. layer_splitin the config (#644) is checked before the start, takes a JSON list, and says the format.- Linux CUDA toolkits (#601):
STRATA_NVCCpicks the toolkit andCUDA_HOMEis honoured. - Windows Pascal/Volta source builds (#585) link the shared CUDA runtime.
- The hit rate (#588) names the PCIe share beside it.
- The server: a malformed
toolsvalue is a 400 instead of a dropped connection (#592); a literal<think>in a
message is read as text (#537);/loadand/unloadread the request body before replying (#630). - Setup: running it again keeps the run config's other keys (#629);
--no-browser/"open_browser": false(#609
#631);--vision-tokens N(#625);--draft-vocab fr, the English/code subset plus French (#597); a calibration on
Linux HIP is kept for its card (#566); Unsloth's UD-IQ4_XS as an experimental choice (#621); on hybrid CPUs with
more E-cores than P-cores, setup writes a recommended--pool-workersyou can edit (#642). - The web page: a Model settings card (#564) and a Conversation cache card in the Monitor tab (#596).
- Docs: the decode window profiler (#610) and the five reading errors #604 found.
New options (off by default; the default output is unchanged):
"effort_position": "end"(#458): a request that only changes the effort reuses the cached conversation."vram_elastic": true(#533): give VRAM back to other programs while the model runs, and take it again
(POST /v1/vram).STRATA_ARENA_MMAP=1(Linux, PR #640): the expert arena as a read-only mapped file, for machines with little RAM
whose GPUs hold most experts.STRATA_STAGE_TRIM=1(PR #639): with explicit--layer-splitpoints, each GPU loads only its own layers' dense
weights (2x MI50, the Coder: 8,819 -> 10,626 experts in VRAM). Opt-in for now on AMD and NVIDIA alike.tools/make_profile.py --reorder(#589) ranks a routing trace ahead of the base profile.- Diagnosis switches for the gfx1201 prompt stalls and the gfx1030 verify timeouts (#579 #613 #541 #649).
Also: a dead engine is restarted for real, with the start retried and the last known context kept (#637).
Thanks to everyone who sent PRs, tests and reports, and especially to JeanP00l for the multi-GPU, gfx906 and restart
work, to maxfridbe for the Intel Arc engine, and to Stuart Chapin for the decode work in #646.
Checked before the release:
- The same answers as 0.1.38 on all four quants (Q2_0, IQ3_XXS, IQ3_S, the Coder), 10/10 each with
a fixed cache and no failures with the default settings. The prompt path's internal state is identical at 4K and
20K tokens (Q2_0) and at 4K (IQ3_XXS, the Coder). Long prompts on a native pack differ by design (#583, above). - Speed, interleaved runs against 0.1.38, medians: decode +2.5% to +6% (10 pairs each, Q2_0 and IQ3_XXS), prompts
at 4K and 32K the same to +2% with setup's default configs (5 pairs each), decode after them +5% to +7%. - Parallel requests (Q2_0): 4 slots of 128 tokens, each identical to the same request alone. These also matched
their solo tokens: a prompt read beside two decoding slots, a prompt that gives way and goes on, a next turn from
its slot, a slot request back on the solo path, and a turn checkpoint. - Kernel parity tests (11), real use at a 59K-token prompt (Q2_0, the Coder), Linux (WSL) Q2_0
identical 10/10 to 0.1.38. - The CUDA 12 engine:
strata-windows-x64-cuda12.zipbuilds (sm_60 to sm_89 + PTX) and gives the same answers as
the CUDA 13 engine on the RTX 5070 (through its PTX), 10/10. We have no Pascal or Volta card, so this checks the
build, not those cards. - AMD on Windows: the HIP zip builds; the AMD changes are untested on an AMD card here (we have none).
- Intel Arc: the SYCL project configures in WSL (oneAPI 2026.1) and compiled on its branch.
- Tests: tools/setup 376, server 268, each run twice. The README speed table was not re-measured for this
release (its IQ2_XS files are not on the test PC).
Updating: run UPDATE.bat (Linux: ./update.sh). Setup installs engine 0.1.39.
The ready-made Strata engines for Windows, which START-HERE.bat fetches by itself:
strata-windows-x64.zip: NVIDIA (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0, needs an
NVIDIA driver 580 or newer. Contents:strata.exe,strata-vision.exe(the optional image encoder),BUILD.json.strata-windows-x64-cuda12.zip(experimental): older NVIDIA cards (Pascal and Volta: sm_60, sm_61, sm_70; it
also has sm_75, sm_80, sm_86, sm_89 + PTX for a mixed PC), CUDA 12.9, needs an NVIDIA driver 528 or newer. Setup
fetches it only for a model you put on such a card, or with--cuda 12. Contents:strata.exe,strata-vision.exe,
BUILD.json.strata-windows-x64-hip.zip: AMD (gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030), ROCm 10.2.0a20260930
from AMD's TheRock builds, needs a current AMD driver. Contents:strata.exe,strata-device.exe, the HIP runtime
next to them,BUILD.json,rocm\(the ROCm libraries and their licenses).
Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going: