Maya-L is out: the closest Maya quant to the full FP8 model, with 99.2% of its zero-shot accuracy - plus GLM's own thinking levels, an exact context meter and an Update button in the dashboard.
Update: git pull, then ./setup.sh (Windows: START-MAYA.bat). From this version on, the dashboard's About > Update does it for you, and the model is never downloaded again.
What's new
- Maya-L (156.3 GB): Maya-M's recipe one step up.
- IQ3_S gate/up experts, IQ4_XS down projections, Q5_K in the most sensitive MoE layers, Q6_K attention and shared experts.
- It keeps 99.2% of the FP8 model's zero-shot accuracy, with the same score as FP8 on HellaSwag and PIQA.
- Its KL divergence is 35% below Maya-M's on the same engine (0.188 vs 0.291), and it picks the FP8 model's next token 90% of the time.
- No loops: two 14,000-token answers, at temperature 1.0 and greedy.
- The most demanding of the four: it is fastest when VRAM and RAM together hold most of its 156 GB.
- Set it up with
./setup.sh --setup --model Maya-L(Windows:START-MAYA.bat --setup --model Maya-L). Details: bench/results/MAYA-L.md.
- Thinking levels are GLM's own: Off, Low, High (the default) and Max in the dashboard;
none,low,highandmaxin the API.- OpenAI's
mediumis High andxhighis Max. Anthropic thinking budgets under 2K tokens are Low, under 8K High, else Max. - The old page's "High" asked GLM for Max, its longest thinking. Settings saved by the old page are translated once.
- OpenAI's
- The context meter counts exactly:
- While it answers: the request's prompt plus the tokens written so far.
- Otherwise: the chat and your draft through the model's own template and tokenizer (nothing is run). It is approximate only with pictures.
- Updates from the dashboard: About says when a new release is out, and its Update button:
- downloads the new code (git);
- compiles only the engine files that changed;
- loads the same model with the same settings. Nothing big is downloaded.
- If this folder can't update itself (not a git checkout, files changed by hand, or a server not started by
./setup.sh/START-MAYA.bat), it says why and gives the steps. - It asks GitHub for the latest release at most every six hours.
MAYA_UPDATE_CHECK=0turns that off.
- Monitor shows each GPU figure once: with two to four GPUs, each hardware card lists every GPU's own value under the total, and the repeated GPUs table is gone. From five GPUs, the table comes back.
- Fixed: saving the Chat settings for other apps failed when Min-p was above 0 (the Creative preset), so nothing was shared.
Checked
- Maya-L against the FP8 model on a Tesla V100: KL divergence, zero-shot accuracy, and the loop test.
- The thinking levels rendered with GLM's own chat template.
- The update against a throwaway git origin: a real fast-forward, then the restart into the new version. Every case that blocks it was checked too.
- The dashboard in the browser, at desktop and phone widths.
- The Python tests: 82 pass, 16 of them new.