github mw00/project-maya v1.0.21
Project Maya v1.0.21

latest releases: v1.0.23, v1.0.22
3 hours ago

Eight community pull requests: decode up to 8% faster at long context, smarter multi-GPU splits, the web app's pictures fixed, and every default fitted to your machine.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • A smaller KV cache, more room for experts (#41 by @merbanan).
    • At 128K context, 1.28 GB more VRAM goes to experts, with the same output.
    • Decode: one V100 with Maya-S24 18.2 -> 19.4 tokens/s; two V100s with Maya-S 26.3 -> 28.2.
    • Optional: STRATA_GLM_KV_INT8=1 stores the attention cache in 8-bit, for another 0.65 GB (19.6 tokens/s on the same V100). It measured the same quality against the full FP8 model.
  • Multi-GPU split and prompt chunks (#44 by @needmorevram).
    • --layer-split auto predicts every placement of the layers on 2-4 GPUs and takes the fastest.
    • --prefill auto sizes prompt chunks from what each card can lend: up to 32768 tokens on one GPU and 8192 on several, each the faster there.
    • On PCs with two CPU sockets (NUMA nodes) or more, each GPU's CPU lane gets CPUs of its own: decode 15.4 -> 19.8 tokens/s on a 2x Xeon. On one socket it stays off, which is faster there.
    • The second card's RAM tier no longer stalls for minutes on fragmented memory.
  • Pictures from the web app (#42 by @tanutanu56): the model was answering about a black image. Fixed.
  • Loop guard (#33 by @ksanislo): a reply stuck repeating a short pattern is closed - the thinking ends, a tool call is completed, or the turn ends - so agents carry on.
  • Thinking budget (#34 by @ksanislo): thinking always leaves room for the answer, so a small max_tokens no longer ends with no answer.
  • Speculative decode on three GPUs or more (#31 by @ksanislo): 4x Tesla T4 decode 15.0 -> 22.6 tokens/s.
  • CPUs without AVX2 (#32 by @ksanislo): older CPUs such as Sandy Bridge and Ivy Bridge can run Maya now.
  • Conversations that outlive a restart (#35 by @ksanislo): STRATA_GLM_SLOT_KEEP=1 keeps saved conversations across starts. A stop (Ctrl+C or SIGTERM) lets the running answer finish first.
  • Housekeeping: no third-party quant names in the repository; the installer's leftover Strata code is gone.

Checked

On 1x and 2x Tesla V100:

  • 239 Python tests and the CUDA parity tests;
  • identical greedy tokens with and without the new KV cache;
  • prefill (reading the prompt), decode (writing the answer), and quality against FP8;
  • the server end to end: an exact answer, a picture from the web app, thinking with a small max_tokens, the context reload 128K -> 16K -> 128K, and a graceful stop.

Don't miss a new project-maya release

NewReleases is sending notifications on new releases.