github mw00/project-maya v1.0.25
Project Maya v1.0.25

latest releases: v1.0.27, v1.0.26
3 hours ago

Multi-GPU rigs of three cards or more now use all of the CPU (+19% decode on a 9-GPU split), --calibrate tunes on realistic text without touching your profile, and there are opt-in CUDA graphs for one GPU.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • One CPU pool for three GPUs or more (#60 by @needmorevram):
    • A token visits the cards one after another, so giving each card its own small share of the CPU left most of it idle. One pool now serves them all.
    • Maya-L on 9 GPUs: decode 25.7 -> 30.7 tokens/s, an 8K prompt 1265 -> 1466 tokens/s.
    • One and two GPUs are unchanged. STRATA_GLM_CPU_SHARED=1 shares the pool on two GPUs too, which was faster on a single-CPU 2x V100 box (30.0 vs 28.5 tokens/s).
  • --calibrate on realistic text (#60):
    • It now answers fresh prompts for every measurement, so it tunes for real chats rather than three repeated answers.
    • It works on a copy of your expert usage profile, so tuning no longer reorders the experts your next start loads first.
  • CUDA graphs for one GPU, opt-in (#59 by @handmade0octopus): STRATA_GLM_KDA_GRAPH=1.
    • The output is bit-identical.
    • RTX 4090D: +2.35% decode. Tesla V100: the same speed.

Checked

On 1x and 2x Tesla V100:

  • the builds and parity tests;
  • identical greedy tokens at the defaults, and with the graphs on;
  • #59's real-model test: 192 bit-for-bit comparisons;
  • decode and prefill (reading the prompt) unchanged at the defaults;
  • a full --calibrate run that left the usage file byte-for-byte unchanged;
  • the GitHub checks.

Don't miss a new project-maya release

NewReleases is sending notifications on new releases.