github mw00/project-maya v1.0.32
Project Maya v1.0.32

6 hours ago

A new model to download, Maya-M-Derisked. A split's GPUs now load at once, three GPUs or more no longer pin more RAM than the PC has, and the Ryzen AI Max APUs decode faster.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • Maya-M-Derisked (#97): Maya-M with a directional weight modification by Blackfrost_AI that reduces blanket refusals.
    • It is not a new quant: the same IQ2_S files, tensors, MTP draft block and size as Maya-M (116 GB), with some of its weights changed.
    • It is experimental, and it lives in a repo of its own.
    • Set it up with ./setup.sh --setup --model Maya-M-Derisked (Windows: START-MAYA.bat --setup --model Maya-M-Derisked), or pick it in the setup's model menu.
    • It reads pictures with Maya's vision files, like the other models.
  • A split's GPUs load at once (#80 by @ksanislo): a thread per GPU, instead of one GPU after another.
    • On 4x Tesla T4 with Maya-L, start to the first token went from 142 to 96 s.
    • The GPUs plan their RAM tiers in turn, so each gets the same tier as before; then they pin and warm them at the same time.
    • STRATA_GLM_PARALLEL_LOAD=0 keeps the old order.
  • The CPU lane's timing, kept across starts (#81 by @ksanislo): with STRATA_GLM_CPU_CAL=<file>, the timing at start (~1.7 s a GPU) is written once and read by later starts. A new build, another card or thread count times again, and a timing that decode finds off is dropped.
  • Three GPUs or more on Linux no longer freeze the desktop at start (#85, issue #77):
    • the later GPUs' prompt buffers are set aside when the RAM tier is sized;
    • each later GPU measures the free RAM again;
    • a warning says when what is left is short.
  • Faster decode on Ryzen AI Max APUs (#95, from the Gorgon Halo work): the decode's dense projections run in one RDNA3 kernel that keeps two blocks' loads in flight, with bit-for-bit the same results.
    • Radeon 8065S with Maya-S: decode +3.1%.
    • On by default on gfx115x. STRATA_GLM_MV_RDNA=1 turns it on for other gfx11 cards; 0 turns it off.
  • The setup accepts a first shard of metadata only (#85, issue #83): a GGUF whose first shard holds only the tokenizer and settings, as unsloth's do, no longer stops the setup as incomplete.
  • For cache studies (by @sociolog):
    • STRATA_GLM_VRAM_EVICT=lru refills a layer's spares by evicting its least recently used expert (#78);
    • STRATA_GLM_ROUTE_LOG writes each route's 16 near misses (#90);
    • tools/glm_tier_replay.py replays a route log through a model of the VRAM tier's rules (#91).

The README has rows for the new settings.

Checked

On 1x and 2x Tesla V100:

  • the build and the GLM parity tests;
  • identical greedy tokens to v1.0.31, the parallel and the one-by-one load alike;
  • the same RAM tiers as v1.0.31, and decode the same within run-to-run variation;
  • the server end to end: answers, pictures, context reloads, a clean stop;
  • a stress run.

Maya-M-Derisked was downloaded and checked against its published sha256. Its answers in several languages, greedy and sampled, end on their own, with no loops and no stray characters.

On a Radeon 8065S (Gorgon Halo): the HIP build, and the new kernel bit-exact in all 333 parity cases and in Maya-S's greedy tokens.

On Windows: the engine compiles. Also: the GitHub checks pass.

Don't miss a new project-maya release

NewReleases is sending notifications on new releases.