The expert prefetch can now pay on a RAM-bound split: three new settings choose when its copy starts, how many GPU blocks it takes and which predictions it copies.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
-
The expert prefetch, tunable (#84 by @sociolog):
STRATA_GLM_PREFETCH_Ncopies the next layer's predicted experts that are not in VRAM into its spare slots while a layer computes. As it was, it cost more than it saved on a split whose experts mostly come from RAM. Three settings for it:STRATA_GLM_PREFETCH_AT=fetch: start the copy after the layer's own fetch, when the PCIe link is free (cpu: after its CPU-lane answer);STRATA_GLM_PREFETCH_BLOCKS=<n>: how many GPU blocks the copy takes (half the SMs before);STRATA_GLM_PREFETCH_RANK=<n>: copy only the prediction's first n guesses, the ones that are almost always right.
On an RTX 3090 + 3060 with Maya-M at 21K context, decode (writing the answer) went from 19.64 to 20.22 tok/s with
STRATA_GLM_PREFETCH_N=1 AT=fetch BLOCKS=4 RANK=2, and +4.7% over a 40-turn agent-like conversation.The prefetch stays off by default, and nothing else changes. The README has a row for it.
Checked
On 1x and 2x Tesla V100:
- the build and the GLM parity tests;
- identical greedy tokens with the CPU lane off;
- decode the same within run-to-run variation (2 GPUs: 30.2 / 30.4 against 30.4 / 30.1 tok/s);
- a run with the new settings on, which completes cleanly.
Also: the GitHub checks pass.