Multi-GPU rigs of three cards or more now use all of the CPU (+19% decode on a 9-GPU split), --calibrate tunes on realistic text without touching your profile, and there are opt-in CUDA graphs for one GPU.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
- One CPU pool for three GPUs or more (#60 by @needmorevram):
- A token visits the cards one after another, so giving each card its own small share of the CPU left most of it idle. One pool now serves them all.
- Maya-L on 9 GPUs: decode 25.7 -> 30.7 tokens/s, an 8K prompt 1265 -> 1466 tokens/s.
- One and two GPUs are unchanged.
STRATA_GLM_CPU_SHARED=1shares the pool on two GPUs too, which was faster on a single-CPU 2x V100 box (30.0 vs 28.5 tokens/s).
--calibrateon realistic text (#60):- It now answers fresh prompts for every measurement, so it tunes for real chats rather than three repeated answers.
- It works on a copy of your expert usage profile, so tuning no longer reorders the experts your next start loads first.
- CUDA graphs for one GPU, opt-in (#59 by @handmade0octopus):
STRATA_GLM_KDA_GRAPH=1.- The output is bit-identical.
- RTX 4090D: +2.35% decode. Tesla V100: the same speed.
Checked
On 1x and 2x Tesla V100:
- the builds and parity tests;
- identical greedy tokens at the defaults, and with the graphs on;
- #59's real-model test: 192 bit-for-bit comparisons;
- decode and prefill (reading the prompt) unchanged at the defaults;
- a full
--calibraterun that left the usage file byte-for-byte unchanged; - the GitHub checks.