A new model to download, Maya-M-Derisked. A split's GPUs now load at once, three GPUs or more no longer pin more RAM than the PC has, and the Ryzen AI Max APUs decode faster.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
- Maya-M-Derisked (#97): Maya-M with a directional weight modification by Blackfrost_AI that reduces blanket refusals.
- It is not a new quant: the same IQ2_S files, tensors, MTP draft block and size as Maya-M (116 GB), with some of its weights changed.
- It is experimental, and it lives in a repo of its own.
- Set it up with
./setup.sh --setup --model Maya-M-Derisked(Windows:START-MAYA.bat --setup --model Maya-M-Derisked), or pick it in the setup's model menu. - It reads pictures with Maya's vision files, like the other models.
- A split's GPUs load at once (#80 by @ksanislo): a thread per GPU, instead of one GPU after another.
- On 4x Tesla T4 with Maya-L, start to the first token went from 142 to 96 s.
- The GPUs plan their RAM tiers in turn, so each gets the same tier as before; then they pin and warm them at the same time.
STRATA_GLM_PARALLEL_LOAD=0keeps the old order.
- The CPU lane's timing, kept across starts (#81 by @ksanislo): with
STRATA_GLM_CPU_CAL=<file>, the timing at start (~1.7 s a GPU) is written once and read by later starts. A new build, another card or thread count times again, and a timing that decode finds off is dropped. - Three GPUs or more on Linux no longer freeze the desktop at start (#85, issue #77):
- the later GPUs' prompt buffers are set aside when the RAM tier is sized;
- each later GPU measures the free RAM again;
- a warning says when what is left is short.
- Faster decode on Ryzen AI Max APUs (#95, from the Gorgon Halo work): the decode's dense projections run in one RDNA3 kernel that keeps two blocks' loads in flight, with bit-for-bit the same results.
- Radeon 8065S with Maya-S: decode +3.1%.
- On by default on gfx115x.
STRATA_GLM_MV_RDNA=1turns it on for other gfx11 cards;0turns it off.
- The setup accepts a first shard of metadata only (#85, issue #83): a GGUF whose first shard holds only the tokenizer and settings, as unsloth's do, no longer stops the setup as incomplete.
- For cache studies (by @sociolog):
The README has rows for the new settings.
Checked
On 1x and 2x Tesla V100:
- the build and the GLM parity tests;
- identical greedy tokens to v1.0.31, the parallel and the one-by-one load alike;
- the same RAM tiers as v1.0.31, and decode the same within run-to-run variation;
- the server end to end: answers, pictures, context reloads, a clean stop;
- a stress run.
Maya-M-Derisked was downloaded and checked against its published sha256. Its answers in several languages, greedy and sampled, end on their own, with no loops and no stray characters.
On a Radeon 8065S (Gorgon Halo): the HIP build, and the new kernel bit-exact in all 333 parity cases and in Maya-S's greedy tokens.
On Windows: the engine compiles. Also: the GitHub checks pass.