Hotfix for 0.1.40.2: Intel Arc install and A-series fixes, Windows AMD speed and runtime fixes, Docker config links, and a few smaller fixes. Default answers are unchanged. Update with UPDATE.bat (Linux: ./update.sh); setup replaces the engine with 0.1.40.3.
Intel Arc
- Install:
python3 sycl/setup_intel.py, as the 0.1.40.2 notes said, fails on Ubuntu's managed Python. The way in is:sudo apt install docker.io, add yourself to thedockergroup and log in again, then./setup.sh --backend sycl. docs/INTEL.md now walks through it, including the build (it was missing) and the A-series. - Arc A750 and other 8 GB A-series cards: setup now writes a config that starts (no
--stream-experts,--ple-io ram, a 300 MiB reserve, the English draft vocabulary, and the FP64 emulation settings these cards need for sampled answers). - The first answer after a start was garbage on the A750 ("!!!!!"): the GPU stopped waiting for the CPU's experts after a few tens of milliseconds, and a cold first request is slower than that. The wait is now long enough; the B70 build is unchanged. A750 IQ3_XXS: correct first and later answers, about 19 tok/s.
- Setup's Intel output is right now (no CUDA banner, no AMD pointer, no "every expert stays in VRAM", the real version number), an explicit
--vram-reserve-mibis kept, and the config'senvblock reaches the engine.
Windows AMD
- Slow answers on 16 GB+ cards (#1376): the automatic expert cache could leave only about 300 MiB of VRAM free; Windows then pages GPU memory and decode fell by half (RX 7900 XT, IQ3_S: 24 tok/s, 55 tok/s with room). An automatic cache on a 16 GB or larger card now keeps 2,560 MiB free. An explicit
--vram-reserve-mibis kept; NVIDIA and Linux are unchanged. Thanks to the reporter for the clean measurements. - The bundled HIP runtime is used again (#461): its dependencies (
rocm_kpack.dlland the Microsoft C++ runtime) were not next tostrata.exe, so Windows fell back to the driver's ownamdhip64_7.dll, which crashed prompts on an RX 7800 XT. They are now copied there, and the engine says which runtime it loaded and, if it is not the bundled one, why. We have no Windows AMD card to confirm this on; reports help. - Driver resets (VIDEO_ENGINE_TIMEOUT_DETECTED): a new troubleshooting entry lists what to send and what to try.
- The web dashboard says that Windows AMD GPU readings are not available yet, instead of showing empty tiles.
Linux AMD
- Setup warns when your user cannot open the GPU (not in the
rendergroup), with the command that fixes it; the engine no longer blames "another program" in that case. - A Radeon 780M (gfx1103) is no longer described as a Strix Halo.
Everyone
- Docker / Kubernetes (#1244): 0.1.40.2 replaced a config link made before the entrypoint, so a pod could silently start the plain setup config. A link into
/data/configis kept now,CONFIG=picks a config file explicitly, and the entrypoint prints which config it uses. - Web app (#1392): a turn with no answer text (only thinking, or an error) is kept in the history it sends.
- MTP: the draft layer's per-token router call has the same 512-expert guard as its siblings (#1357, thanks to rwkeyes).
- Tokenizer (#1385): hardened against an intermittent error seen on Python 3.14.
- Docs:
STRATA_STAGER_SLEEP(sleeping prompt-stager waits: much less CPU, 5-6% slower prompt reads on a Windows laptop).
Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going: