Better drafts for speculative decode on two GPUs or more: more drafts accepted, never slower.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
- Better drafts (#48 by @sociolog): when the model's draft block is missing an expert in VRAM, it now fetches the missing ones among its 2 most important experts instead of skipping them all.
- 2x Tesla V100: 82-83% of drafts accepted instead of 76-81%, and decode at 128K context 28.0-28.2 tokens/s instead of 27.0-28.1.
- RTX 3090 + 3060: 19.5 -> 20.6 tokens/s.
- It never measured slower, and the answers are the model's own.
STRATA_GLM_MTP_KEEP=<n>changes how many:0= the old behaviour.
Checked
- On 2x Tesla V100: the build, five draft settings over two rounds, and identical greedy tokens with the CPU lane off.
- The GitHub checks pass.