AMD GPUs now set up on Windows too, Windows reads its experts from the SSD much faster, and Maya's logs and installer use its own name.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
- AMD on Windows, Strix Halo / Gorgon Halo included (experimental) (#69):
START-MAYA.bat --backend hipsets Maya up on Windows 10/11 as./maya.sh --backend hipdoes on Linux. That covers Ryzen AI Max 300 / 400 (Radeon 8050S / 8060S / 8065S) and the RX 7900 XT / XTX and RX 9070 / R9700.- Setup: finds the GPUs, uses AMD's HIP SDK (or offers AMD's ROCm SDK in
.venv), and builds the engine with Visual Studio 2022 or 2026. - The engine: handles a Windows APU's memory and Windows' way of submitting GPU work.
- Checked on a Ryzen AI Max+ PRO 495: setup builds and packs Maya-L, every expert fits on the GPU, and all 15 HIP checks pass.
- Setup: finds the GPUs, uses AMD's HIP SDK (or offers AMD's ROCm SDK in
- Windows reads experts from the SSD much faster (#73 by @klvnblst):
- The model files are unmapped after loading: while they stay mapped, Windows slows every direct read of them. On the RTX 5090 Laptop that reported it, decode is about 3.5x as fast.
- CPUs without hyper-threading leave a third of their cores free for the disk readers.
- The CPU lane times itself on experts read from RAM (#70 by @merbanan): it used to time experts sitting in the CPU's cache and overrate the CPU. +6% decode on a Ryzen 9 7950X with a Tesla V100, and the same plan at every start.
- Maya's own name (#74): the logs say
[maya]andmaya ..., and the installer no longer shows Strata's name. Settings, flags and the API are unchanged, and the README keeps the credit. - Windows setup in the README, step by step (#71 by @handmade0octopus): the prerequisites, copy-paste commands, and page-file advice that matches
--check. - Fix: a restart no longer mistakes the dashboard's saved settings for a model's config.
Checked
On 1x and 2x Tesla V100:
- the builds and parity tests, and identical greedy tokens with the CPU lane off;
- decode (writing the answer) and prefill (reading the prompt) the same within run-to-run variation;
- the server end to end, the setup screen in a real terminal, and the logs saying
[maya].
Also: the engine compiles on Windows, AMD on Windows was run on the author's Ryzen AI Max+ PRO 495, and the GitHub checks pass.