v1.2.0 — GLM-5.3-Flash, 32 GB streaming, zero-RAM prompt cache
- GLM-5.3-Flash: the new Sushi-2.4bpw pack runs on 128 GB Macs (KLD 0.074 against the BF16 model) with image and
video input, the full 1M context and DFlash2 drafting: 42–55 tok/s decode on an M5 Max. Up to four requests decode
together. - Streaming with MTP:
--ssd-budget-gbstreams any model's experts from the SSD, so Qwen3.8 Sushi-2bpw runs on a
32 GB Mac at about 20 tok/s. - Prompt cache with 0 GB of RAM: prompts are reused across turns from an SSD cache that sizes itself (up to 20 GB);
keeping them in RAM is now opt-in with--prefix-cache-mem. - Faster: MiMo decodes about 12% and prefills about 15% faster, Qwen3.8-Flash-Next prefills faster, and concurrent
requests decode together on every model. - Agents and API:
sushi launch grokis new and opencode 2.x works again; streamed and non-streamed answers match
byte for byte; penalties,repetition_penaltyandignore_eoswork; GLM history renders exactly as its template. - Flags:
--mtp-min-depth/--mtp-max-depthreplace--mtp-depth; new--no-mtp-lookupand--gpu-warm-secs;
--wired-margin-gibdefaults to 4 GiB. - Reliability: long GLM sessions, memory pressure and restarts are handled cleanly, and an idle streamed server no
longer uses CPU.
Thanks @cnsiva (request-budget defaults, repetition_penalty), @jasontitus (unique response IDs, the M1–M4 GLM fix,
portable GLM tests), @ShoichiTect (idle SSD-read workers) and @gomezvd (MTP on streamed Qwen).