Several conversations at once on a layer split, a short request no longer waits for a long answer, and fixes from your reports: a start that failed on Windows, a speed that changed between restarts, ZFS, and --calibrate now tunes the disk reads too.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
- Several conversations at once on a layer split (#94 by @0xPreDa): with
STRATA_GLM_SEQS=<n>on two GPUs or more, n conversations decode together, each card of the split working on a different one instead of waiting for the others. On 4x RTX 4090, four conversations at once give 225 tok/s in total. It's off by default; one GPU gains nothing from it. - A short request no longer waits for a whole long answer (#93 by @0xPreDa): requests are served in arrival order, and with
STRATA_FAIR_SLICE_S=<s>a long answer that has been decoding for s seconds while another request waits lets it in, then continues from where it was. On 2x V100, a short question asked during a 700-token answer got its first word after 10 s instead of 60. --calibratetunes the disk reads too (#99, from #67): when part of the model is read from the SSD while it answers, the tuning also tries reading each expert in fewer, larger pieces, and keeps a size only when it is more than 3% faster. A Windows laptop on an Intel RST RAID decoded 15% faster with 4 pieces than with the default 8. A Linux NVMe stays fastest at 8, so nothing changes there.- Fixes from your reports (#99):
- A start that failed right after the RAM tier on Windows (
pinned disk staging did not allocate, #67): the small pinned buffers are now set up before the tier, so a cap on pinned memory makes the tier slightly smaller instead of stopping the start. - A speed that changed between restarts (#56): the CPU lane's timing at start is now taken over half a second, and keeps the fastest round, so a burst of other work at that moment no longer makes the whole session slower.
- ZFS: the memory ZFS's cache can give back counts as free RAM when the RAM tier is sized (#89).
- The prefetch says its settings in the engine log when it is on (#6).
- A start that failed right after the RAM tier on Windows (
Checked
On 1x and 2x Tesla V100:
- the build and the GLM parity tests;
- identical greedy and sampled tokens to v1.0.32, and the same answers across a conversation's turns;
- decode the same within run-to-run variation;
- several conversations at once writing the same tokens as one at a time;
- the server end to end and a stress run.
On Windows: the engine compiles. Also: the GitHub checks pass.