github mw00/project-maya v1.0.33
Project Maya v1.0.33

4 hours ago

Several conversations at once on a layer split, a short request no longer waits for a long answer, and fixes from your reports: a start that failed on Windows, a speed that changed between restarts, ZFS, and --calibrate now tunes the disk reads too.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • Several conversations at once on a layer split (#94 by @0xPreDa): with STRATA_GLM_SEQS=<n> on two GPUs or more, n conversations decode together, each card of the split working on a different one instead of waiting for the others. On 4x RTX 4090, four conversations at once give 225 tok/s in total. It's off by default; one GPU gains nothing from it.
  • A short request no longer waits for a whole long answer (#93 by @0xPreDa): requests are served in arrival order, and with STRATA_FAIR_SLICE_S=<s> a long answer that has been decoding for s seconds while another request waits lets it in, then continues from where it was. On 2x V100, a short question asked during a 700-token answer got its first word after 10 s instead of 60.
  • --calibrate tunes the disk reads too (#99, from #67): when part of the model is read from the SSD while it answers, the tuning also tries reading each expert in fewer, larger pieces, and keeps a size only when it is more than 3% faster. A Windows laptop on an Intel RST RAID decoded 15% faster with 4 pieces than with the default 8. A Linux NVMe stays fastest at 8, so nothing changes there.
  • Fixes from your reports (#99):
    • A start that failed right after the RAM tier on Windows (pinned disk staging did not allocate, #67): the small pinned buffers are now set up before the tier, so a cap on pinned memory makes the tier slightly smaller instead of stopping the start.
    • A speed that changed between restarts (#56): the CPU lane's timing at start is now taken over half a second, and keeps the fastest round, so a burst of other work at that moment no longer makes the whole session slower.
    • ZFS: the memory ZFS's cache can give back counts as free RAM when the RAM tier is sized (#89).
    • The prefetch says its settings in the engine log when it is on (#6).

Checked

On 1x and 2x Tesla V100:

  • the build and the GLM parity tests;
  • identical greedy and sampled tokens to v1.0.32, and the same answers across a conversation's turns;
  • decode the same within run-to-run variation;
  • several conversations at once writing the same tokens as one at a time;
  • the server end to end and a stress run.

On Windows: the engine compiles. Also: the GitHub checks pass.

Don't miss a new project-maya release

NewReleases is sending notifications on new releases.