Maya gets a screen of its own in the terminal, the crash after a long prompt is fixed, a stuck prompt now ends with an error instead of waiting forever, and Maya builds on Windows again.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
- A screen of its own in the terminal (#58 by @needmorevram):
- The setup's steps are tabs, the questions are arrow-key menus, and the compile and the download show their progress.
- Maya then runs on the same screen: a loading view while the experts warm up, the Monitor's numbers, and its log.
- Nothing is downloaded for it: its library ships with Maya, byte-identical to PyPI's.
--plain(or a pipe, or a service) keeps the plain text as before.
- No more "illegal memory access" after a long prompt, and no more 0xC0000005 on Windows / AMD (#39 by @boxwrench). It fixes three races in the expert tiers that NVIDIA and AMD share:
- the expert tables' update buffers were rewritten while their copy to the GPU was still in flight;
- a background expert move could land in a VRAM slot a prompt was borrowing;
- a route of NaN scores indexed the expert tables. Now that request fails with an error and the engine keeps running.
- Confirmed on an RTX 3090 + 3060 (#50: the 8K prompt after decoding no longer crashes) and on Windows with an RX 7900 XTX (#53: 5 crashes in 8 sessions before, 0 in 8 now).
- A stuck prompt ends (#40): when a request finishes no prompt layer and no token for 3 minutes, the engine stops with an error that says where it was stuck. On Windows it also writes
strata-stall-<pid>.dmp. The next request starts the engine again. Before, a prompt that stalled with the GPU idle waited forever.STRATA_WATCHDOG_Ssets the time;0turns it off. - Windows builds again (#54 by @jerem91150; #55 by @noahark had the same fix).
- Thinking off stays off (#54): the answer no longer lands inside an empty thinking block. This applies to every GPU.
- AMD:
- a Windows build script (#54);
- RDNA4 prefill (reading the prompt) +8-10% on R9700 / RX 9070 (#38 by @boxwrench).
- One GPU, prefill (reading the prompt) about +1% (#61 by @boxwrench): the last layer's unused outputs are skipped, and the output is the same.
- Robustness (#62 by @needmorevram):
- F16, BF16, Q2_K, Q4_1 and Q5_1 GGUFs read the right embedding rows. The published Maya models were not affected.
- Bad image files are refused instead of ending the engine.
- Switch chats while an answer is being written (#66 by @needmorevram): the other chat opens, and the answer carries on in its own chat.
- A setup for another context keeps your
--calibratetuning (#57). Before, it fell back to the engine's defaults.
Checked
On 1x and 2x Tesla V100, and on Windows:
- the builds and parity tests, including the engine's Windows build;
- identical greedy tokens with the CPU lane off on one and two GPUs, and with the graphs on;
- #50's sequence (every RAM-tier expert on the CPU, then long prompts) on one and two GPUs;
- the watchdog: a forced stall ends with an error and the next request is answered, and it never fires on normal work;
- decode (writing the answer) and prefill (reading the prompt) the same, prefill +1% on one GPU;
- the server end to end, #58's screen in a real terminal, and #66 in a browser against the running model;
- the GitHub checks.