Maya-L runs at 16.4 tokens/s on an RTX 4090 with 192 GB of RAM under Windows 11 - a user's report, now in the README.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). Documentation only; nothing is downloaded again.
What's new
- A new speed report (#63 by @npc97): Maya-L on an RTX 4090 24 GB, a Ryzen 9 9950X3D and 192 GB of DDR5 under Windows 11.
- Every expert fits in VRAM or RAM, so nothing is read from the SSD.
- Decode (writing the answer): 16.4 tokens/s. Prefill (reading the prompt): 1371 tokens/s on an 8K-token prompt.
- It's in the README's user table, and the Windows section now says Maya runs on NVIDIA under Windows, with CUDA 13 working for RTX 20 and newer.
- CPU pinning is Linux-only (#64 by @needmorevram): the README now says
STRATA_GLM_CPU_PINdoes nothing on Windows or with one GPU.