A new dashboard: many chats, the context size from Settings (the model reloads with it), edit and regenerate, code previews, and an About page that finally shows the version.
Update: git pull, then ./setup.sh (Windows: START-MAYA.bat). Your chat from the old page moves into the new chat list by itself.
What's new
- Chats: every chat is kept in your browser (with its pictures and attached files) and listed in the sidebar, with search, rename and delete. "New chat" no longer wipes the last one. On a phone, the list slides in from the left.
- Edit a message and send it again (the answers after it are replaced), or regenerate the last answer.
- An answer that was streaming when the page reloaded is taken back up from the server.
- Context size in Settings: a slider from 4K to the model's trained 1M tokens.
- It shows what the size costs on your PC, from the engine's own measurement. On 2x V100: 3.3 GB at 128K, and about 249 more experts in VRAM at 64K.
- It warns past 256K (untested) and asks before it reloads the model (about a minute on 2x V100).
- Other apps get a 503 with Retry-After while it reloads. A size that doesn't start puts the old one back. The new size is saved for the next start.
- Instructions per chat (a system prompt, optionally for new chats too), sampling presets (Precise, Balanced, Creative), min-p, and a warning before Settings closes with unsaved changes.
- While it answers: a context meter for the chat, the prompt it read and reused under each answer, a "Latest" button, Esc to stop, Ctrl+Shift+O for a new chat.
- Code: syntax highlighting (bundled, nothing loaded from the web), wrap, download, and a sandboxed preview for HTML pages.
- Monitor:
- A reading from the last answer is now marked as such (it looked live). Idle charts say so instead of drawing a flat line.
- A table of GPUs on multi-GPU PCs.
- Each request's source (Chat / OpenAI / Anthropic) and prefill (reading the prompt) speed.
- Banners for a reload or a hot GPU, and tips from the expert tiers.
- Copy report for a GitHub issue: versions, PC, engine settings, and the telling lines of the engine log. No API key.
- About:
- The version, which was missing.
- The trained context, and the KV cache as it really is: 16-bit (it said "32-bit").
- Each GPU's PCIe link, and the model folder with its free disk space.
- Copy-ready snippets for Claude Code, OpenAI-style apps and curl.
- Phones and accessibility: no zoom into the message box on iPhones, larger buttons, safe areas, a status word in the header, and installable as an app. Focus stays inside Settings and dialogs, the arrow keys move the thinking level, and reduced motion is honoured.
Checked
- On 2x Tesla V100 with Maya-S, the context was changed 128K -> 96K -> 128K through the page:
- each reload took about a minute;
- chat requests got a 503 meanwhile;
- the run config's other settings stayed as they were;
- the model then answered exactly.
- In the browser, at desktop and phone widths: chats, edit, regenerate, the preview, the reload and the report.
- The server's tests: 102, plus 9 new ones for the context size and the report.