github Niko1221/Strata v0.1.39
Strata v0.1.39

2 hours ago

Faster decode and long prompts, several requests at once, older and other hardware as experimental opt-ins, the
OpenAI Responses API for Codex, a fix for slower prompts with a RAM budget, and a batch of fixes from your reports.

Faster (the same answers where nothing says otherwise):

  • Decode (#646): the verify pass runs with fewer launches and host round trips (sub-warp expert packing, staged
    inputs, the PLE and MTP steps batched). Measured here on one RTX 5070 against 0.1.38, 10 interleaved pairs, medians:
    Q2_0 +6% (story 68.4 -> 72.7 tok/s, code 75.8 -> 80.0), IQ3_XXS +6% / +2.5% (46.1 -> 48.8, 50.7 -> 51.9); after
    4K and 32K prompts +5% to +7%. The output is byte-identical to 0.1.38's (10/10 on all four quants with a fixed
    cache). The zero-doorbell verify graph itself runs only when every expert of a layer is in VRAM. A 12 GB card never
    gets there, so the gain on big cards was not measured here.
  • Long prompts (#583, on by default): the streamed expert ring is sized in bytes for the pack, and --prefill auto
    picks the largest chunk that keeps it full. Short prompts, and prompts that fit 0.1.38's chunk, keep 0.1.38's ring
    and their exact bits (4K: identical). The gain depends on the free VRAM. On an RTX 5070 with 32K prompts, IQ3_XXS is
    +18.5% with a 1,500-slot expert cache and unchanged with setup's default config. A long prompt's bits change against
    0.1.38, because the experts go through a different mix of cached and streamed groups. Its quality stays in the same
    band, teacher-forced over the next 2,001 tokens against the FP16 prompt path (IQ3_XXS): 8K prompt KL 0.042 (0.1.38:
    0.054), top-1 93.5% (92.2%); 32K prompt KL 0.020 (0.019), top-1 95.3% (95.0%). It doesn't help every pack: the
    Coder at a 32K context reads a 30K prompt about 8% slower (with a 64K context it was about 6% faster).
    STRATA_RING_BYTES=0 restores 0.1.38's ring.
  • Very long contexts: the attention's top-k past the register kernel's reach (262K-524K cells) runs with a
    histogram per warp: a 243K-token prompt +26% on an RTX 3060 (#603).
  • More than one GPU (measured by their authors, not here: we have one GPU): setup now adds
    --remote-expert-opt (#578) to a config with two or more GPUs. It skips the host's work for tokens whose experts
    all run on the helper cards (dual RTX 4090: +63% mixed text, +132% code over the plain helper path), and setup --no-remote-expert-opt leaves it out. A layer split of 3+ GPUs overlaps its stages better (#598, 4x RTX 3060: +10
    to 28% on prompts). STRATA_PREFILL_HELP=1 lets a split's idle card stream a share of a one-chunk prompt's experts
    (#663). It is opt-in because it rounds differently. With one or two GPUs the pipeline runs as before.
  • Also: the Linux expert arena on transparent huge pages (#650; STRATA_NO_LARGEPAGES=1 keeps 4 KB pages), an
    opt-in AVX2 codebook gather for the IQ CPU kernels (STRATA_IQ256_GATHER=1, #622), and thread affinity on PCs with
    more than 64 CPUs (#626).

Several requests at once (#465, opt-in): "parallel": N in strata-<model>.json (or setup --parallel N)
decodes up to N conversations together in the engine. More requests wait for a free slot, and a long prompt gives
way to a short one at a chunk boundary. A request left alone goes back to the one-at-a-time path, and /metrics
shows each slot. Every slot's greedy answer is the same as when it runs alone. On a 12 GB card it cuts waiting time
and costs speed. RTX 5070, Q2_0, four requests at once: the last one starts after 1.8 s instead of 11.2 s, but
together they decode 11% slower (63.1 against 70.7 tok/s). A request alone loses 11% with 2 slots and 22% with 4,
because each slot's session takes 0.56 GiB from the expert cache. So setup recommends it only where the experts
mostly fit in VRAM. A 4-GPU layer split (4x 16 GB, IQ3_S, PR #559) served 8 requests at 360 tok/s in total against
120 for one. docs/BATCHING.md has the numbers and the options.

Older and other hardware, experimental: three opt-in paths that community members wrote and measured on their
own machines. We have none of this hardware. Each is compile-checked and unit-tested here, and the ready-made
engines and their output are unchanged (checked byte-identical).

  • Older NVIDIA cards (Pascal, Volta: P40, P100, GTX 10, V100, Titan V; #395 #600 #540 #655 #627): CUDA 13 cannot
    compile for them, so setup keeps a second engine built with CUDA 12.9 (strata-windows-x64-cuda12.zip on
    Windows, compiled on Linux). It is used only when you choose such a card: a PC with only Pascal / Volta cards, a
    card named with --gpu N / --gpus, or --cuda 12. --cuda 12 is also the way to run with an NVIDIA driver
    older than 580 (528+ on Windows, 525+ on Linux). The choice is kept per model. The Volta prompt attention and the
    BF16 path through FP16 / fp32 are compiled into this engine only. Untested here: every Pascal and Volta card, the
    old drivers, and an RTX 50 card in that engine (it gets a warning; keep it on its own model). docs/OLDER_GPUS.md
    has the reporters' numbers (V100: prompts 1,123-1,251 tok/s, UD-IQ4_XS). RTX 20 owners can try the same FP16
    tensor-core path for the prompt's BF16 products with STRATA_BF16_TC=1 (+15-18% on an RTX 2080 Ti, #655). Its
    sums are not bitwise cuBLAS's.
    Older AMD cards are built by hand: gfx906 (Instinct MI50 / MI60, Radeon VII; -DSTRATA_HIP_GFX906=ON, #638 #677:
    2x MI50, the Coder, 50 tok/s at 4K) and gfx1012 (RX 5500 XT, #442). The RX 6700 XT (gfx1031, #524) goes through
    setup.
  • Intel Arc (#423): maxfridbe's SYCL port of the engine, built from source on Linux with Intel oneAPI:
    ./setup.sh --backend sycl. Reported on the Arc Pro B70 / B50 and the B580 (Coder IQ1_M 70-78 tok/s on a B70) on
    earlier versions. Here it compiles (oneAPI 2026.1) and its kernel tests run on a CPU device. Untested: the 0.1.39
    port on an Arc, Windows (no build path yet; setup points at Linux), WSL2, the A-series and integrated Arc GPUs, and
    images. There is no ready-made Intel engine. docs/INTEL_ARC.md.
  • Older CPUs without AVX2 (#394 #595): AVX-only (Sandy / Ivy Bridge, Xeon E5 v1/v2, Bulldozer) and
    SSE4.2-only CPUs (Nehalem / Westmere). Run setup as usual: it warns and compiles the engine for that CPU
    (STRATA_ISA_FLOOR=avx or none, 10-20 minutes once) instead of stopping. Only the i-quant models run there, and
    the CPU's share is slow (forced on our Ryzen: 11-17 tok/s with the AVX build, 3.5-3.8 with SSE4.2, against 26).
    Untested here on a real old CPU; contributors ran earlier versions of it on Xeon E5-2680 / E5-2687W / X5690.
    docs/INSTALL.md#older-cpus-experimental.

Prompts with a RAM budget are fast again (#577): 0.1.38 decided too early whether to read the experts past the
file cache, and it compared against the size of every model file. On a 96 GB PC with UD-Q4_K_XL and a 72 GiB budget
it chose wrong, so every refill after a prompt read the drive (prompts 15-40% slower than 0.1.34). The choice now
counts only the expert bytes outside the RAM copy, and is made again once that copy is built. The tokens are the same
either way.

The OpenAI Responses API (#451): POST /v1/responses, so Codex CLI works with Strata (tested with Codex 0.160.0,
a tool loop included; later turns reused ~96% of the prompt from the cache). It is stateless, as Codex uses it, and
covers function tools, tool results, reasoning effort, JSON schemas and the streaming events. Not supported:
previous_response_id (nothing is stored), hosted tools and reasoning summaries. The Codex config.toml is in
docs/DETAILS.md.

Security: SECURITY.md says how to report a problem privately (GitHub's private vulnerability reporting) and what
the server exposes: 127.0.0.1 by default, the API key, the Host and Origin checks of 0.1.38, CORS, and the opt-in MCP
tools and request monitor.

Fixes from your reports:

  • A reply stuck on one token is ended (#606): a reply that repeats one token 256 times ends there
    ("repeat_stop_tokens" sets the length, 0 turns it off), and the q8_1 activations stay finite.
  • Starting on a tight card (#620): the native head and the logits are loaded before the expert arena, and a failed
    allocation names the free VRAM.
  • The RAM check before the arena (#633): a container whose memory limit is below the arena is refused with the
    numbers instead of being killed during the load. Less RAM available only warns.
  • A rotational disk (#605): --ple-io direct warns at start, setup keeps the n-gram table in RAM there, and the
    stall report counts disk waits.
  • layer_split in the config (#644) is checked before the start, takes a JSON list, and says the format.
  • Linux CUDA toolkits (#601): STRATA_NVCC picks the toolkit and CUDA_HOME is honoured.
  • Windows Pascal/Volta source builds (#585) link the shared CUDA runtime.
  • The hit rate (#588) names the PCIe share beside it.
  • The server: a malformed tools value is a 400 instead of a dropped connection (#592); a literal <think> in a
    message is read as text (#537); /load and /unload read the request body before replying (#630).
  • Setup: running it again keeps the run config's other keys (#629); --no-browser / "open_browser": false (#609
    #631); --vision-tokens N (#625); --draft-vocab fr, the English/code subset plus French (#597); a calibration on
    Linux HIP is kept for its card (#566); Unsloth's UD-IQ4_XS as an experimental choice (#621); on hybrid CPUs with
    more E-cores than P-cores, setup writes a recommended --pool-workers you can edit (#642).
  • The web page: a Model settings card (#564) and a Conversation cache card in the Monitor tab (#596).
  • Docs: the decode window profiler (#610) and the five reading errors #604 found.

New options (off by default; the default output is unchanged):

  • "effort_position": "end" (#458): a request that only changes the effort reuses the cached conversation.
  • "vram_elastic": true (#533): give VRAM back to other programs while the model runs, and take it again
    (POST /v1/vram).
  • STRATA_ARENA_MMAP=1 (Linux, PR #640): the expert arena as a read-only mapped file, for machines with little RAM
    whose GPUs hold most experts.
  • STRATA_STAGE_TRIM=1 (PR #639): with explicit --layer-split points, each GPU loads only its own layers' dense
    weights (2x MI50, the Coder: 8,819 -> 10,626 experts in VRAM). Opt-in for now on AMD and NVIDIA alike.
  • tools/make_profile.py --reorder (#589) ranks a routing trace ahead of the base profile.
  • Diagnosis switches for the gfx1201 prompt stalls and the gfx1030 verify timeouts (#579 #613 #541 #649).

Also: a dead engine is restarted for real, with the start retried and the last known context kept (#637).

Thanks to everyone who sent PRs, tests and reports, and especially to JeanP00l for the multi-GPU, gfx906 and restart
work, to maxfridbe for the Intel Arc engine, and to Stuart Chapin for the decode work in #646.

Checked before the release:

  • The same answers as 0.1.38 on all four quants (Q2_0, IQ3_XXS, IQ3_S, the Coder), 10/10 each with
    a fixed cache and no failures with the default settings. The prompt path's internal state is identical at 4K and
    20K tokens (Q2_0) and at 4K (IQ3_XXS, the Coder). Long prompts on a native pack differ by design (#583, above).
  • Speed, interleaved runs against 0.1.38, medians: decode +2.5% to +6% (10 pairs each, Q2_0 and IQ3_XXS), prompts
    at 4K and 32K the same to +2% with setup's default configs (5 pairs each), decode after them +5% to +7%.
  • Parallel requests (Q2_0): 4 slots of 128 tokens, each identical to the same request alone. These also matched
    their solo tokens: a prompt read beside two decoding slots, a prompt that gives way and goes on, a next turn from
    its slot, a slot request back on the solo path, and a turn checkpoint.
  • Kernel parity tests (11), real use at a 59K-token prompt (Q2_0, the Coder), Linux (WSL) Q2_0
    identical 10/10 to 0.1.38.
  • The CUDA 12 engine: strata-windows-x64-cuda12.zip builds (sm_60 to sm_89 + PTX) and gives the same answers as
    the CUDA 13 engine on the RTX 5070 (through its PTX), 10/10. We have no Pascal or Volta card, so this checks the
    build, not those cards.
  • AMD on Windows: the HIP zip builds; the AMD changes are untested on an AMD card here (we have none).
  • Intel Arc: the SYCL project configures in WSL (oneAPI 2026.1) and compiled on its branch.
  • Tests: tools/setup 376, server 268, each run twice. The README speed table was not re-measured for this
    release (its IQ2_XS files are not on the test PC).

Updating: run UPDATE.bat (Linux: ./update.sh). Setup installs engine 0.1.39.

The ready-made Strata engines for Windows, which START-HERE.bat fetches by itself:

  • strata-windows-x64.zip: NVIDIA (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0, needs an
    NVIDIA driver 580 or newer. Contents: strata.exe, strata-vision.exe (the optional image encoder), BUILD.json.
  • strata-windows-x64-cuda12.zip (experimental): older NVIDIA cards (Pascal and Volta: sm_60, sm_61, sm_70; it
    also has sm_75, sm_80, sm_86, sm_89 + PTX for a mixed PC), CUDA 12.9, needs an NVIDIA driver 528 or newer. Setup
    fetches it only for a model you put on such a card, or with --cuda 12. Contents: strata.exe, strata-vision.exe,
    BUILD.json.
  • strata-windows-x64-hip.zip: AMD (gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030), ROCm 10.2.0a20260930
    from AMD's TheRock builds, needs a current AMD driver. Contents: strata.exe, strata-device.exe, the HIP runtime
    next to them, BUILD.json, rocm\ (the ROCm libraries and their licenses).

Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going:

Buy Me A Coffee

Don't miss a new Strata release

NewReleases is sending notifications on new releases.