github mw00/project-maya v1.0.31
Project Maya v1.0.31

2 hours ago

The expert prefetch can now pay on a RAM-bound split: three new settings choose when its copy starts, how many GPU blocks it takes and which predictions it copies.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • The expert prefetch, tunable (#84 by @sociolog): STRATA_GLM_PREFETCH_N copies the next layer's predicted experts that are not in VRAM into its spare slots while a layer computes. As it was, it cost more than it saved on a split whose experts mostly come from RAM. Three settings for it:

    • STRATA_GLM_PREFETCH_AT=fetch: start the copy after the layer's own fetch, when the PCIe link is free (cpu: after its CPU-lane answer);
    • STRATA_GLM_PREFETCH_BLOCKS=<n>: how many GPU blocks the copy takes (half the SMs before);
    • STRATA_GLM_PREFETCH_RANK=<n>: copy only the prediction's first n guesses, the ones that are almost always right.

    On an RTX 3090 + 3060 with Maya-M at 21K context, decode (writing the answer) went from 19.64 to 20.22 tok/s with STRATA_GLM_PREFETCH_N=1 AT=fetch BLOCKS=4 RANK=2, and +4.7% over a 40-turn agent-like conversation.

    The prefetch stays off by default, and nothing else changes. The README has a row for it.

Checked

On 1x and 2x Tesla V100:

  • the build and the GLM parity tests;
  • identical greedy tokens with the CPU lane off;
  • decode the same within run-to-run variation (2 GPUs: 30.2 / 30.4 against 30.4 / 30.1 tok/s);
  • a run with the new settings on, which completes cleanly.

Also: the GitHub checks pass.

Don't miss a new project-maya release

NewReleases is sending notifications on new releases.