github Niko1221/Strata v0.1.40.3
Strata v0.1.40.3

3 hours ago

Hotfix for 0.1.40.2: Intel Arc install and A-series fixes, Windows AMD speed and runtime fixes, Docker config links, and a few smaller fixes. Default answers are unchanged. Update with UPDATE.bat (Linux: ./update.sh); setup replaces the engine with 0.1.40.3.

Intel Arc

  • Install: python3 sycl/setup_intel.py, as the 0.1.40.2 notes said, fails on Ubuntu's managed Python. The way in is: sudo apt install docker.io, add yourself to the docker group and log in again, then ./setup.sh --backend sycl. docs/INTEL.md now walks through it, including the build (it was missing) and the A-series.
  • Arc A750 and other 8 GB A-series cards: setup now writes a config that starts (no --stream-experts, --ple-io ram, a 300 MiB reserve, the English draft vocabulary, and the FP64 emulation settings these cards need for sampled answers).
  • The first answer after a start was garbage on the A750 ("!!!!!"): the GPU stopped waiting for the CPU's experts after a few tens of milliseconds, and a cold first request is slower than that. The wait is now long enough; the B70 build is unchanged. A750 IQ3_XXS: correct first and later answers, about 19 tok/s.
  • Setup's Intel output is right now (no CUDA banner, no AMD pointer, no "every expert stays in VRAM", the real version number), an explicit --vram-reserve-mib is kept, and the config's env block reaches the engine.

Windows AMD

  • Slow answers on 16 GB+ cards (#1376): the automatic expert cache could leave only about 300 MiB of VRAM free; Windows then pages GPU memory and decode fell by half (RX 7900 XT, IQ3_S: 24 tok/s, 55 tok/s with room). An automatic cache on a 16 GB or larger card now keeps 2,560 MiB free. An explicit --vram-reserve-mib is kept; NVIDIA and Linux are unchanged. Thanks to the reporter for the clean measurements.
  • The bundled HIP runtime is used again (#461): its dependencies (rocm_kpack.dll and the Microsoft C++ runtime) were not next to strata.exe, so Windows fell back to the driver's own amdhip64_7.dll, which crashed prompts on an RX 7800 XT. They are now copied there, and the engine says which runtime it loaded and, if it is not the bundled one, why. We have no Windows AMD card to confirm this on; reports help.
  • Driver resets (VIDEO_ENGINE_TIMEOUT_DETECTED): a new troubleshooting entry lists what to send and what to try.
  • The web dashboard says that Windows AMD GPU readings are not available yet, instead of showing empty tiles.

Linux AMD

  • Setup warns when your user cannot open the GPU (not in the render group), with the command that fixes it; the engine no longer blames "another program" in that case.
  • A Radeon 780M (gfx1103) is no longer described as a Strix Halo.

Everyone

  • Docker / Kubernetes (#1244): 0.1.40.2 replaced a config link made before the entrypoint, so a pod could silently start the plain setup config. A link into /data/config is kept now, CONFIG= picks a config file explicitly, and the entrypoint prints which config it uses.
  • Web app (#1392): a turn with no answer text (only thinking, or an error) is kept in the history it sends.
  • MTP: the draft layer's per-token router call has the same 512-expert guard as its siblings (#1357, thanks to rwkeyes).
  • Tokenizer (#1385): hardened against an intermittent error seen on Python 3.14.
  • Docs: STRATA_STAGER_SLEEP (sleeping prompt-stager waits: much less CPU, 5-6% slower prompt reads on a Windows laptop).

Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going:

Buy Me A Coffee

Don't miss a new Strata release

NewReleases is sending notifications on new releases.