github beamivalice/sushi v1.2.0
sushi v1.2.0

3 hours ago

v1.2.0 — GLM-5.3-Flash, 32 GB streaming, zero-RAM prompt cache

  • GLM-5.3-Flash: the new Sushi-2.4bpw pack runs on 128 GB Macs (KLD 0.074 against the BF16 model) with image and
    video input, the full 1M context and DFlash2 drafting: 42–55 tok/s decode on an M5 Max. Up to four requests decode
    together.
  • Streaming with MTP: --ssd-budget-gb streams any model's experts from the SSD, so Qwen3.8 Sushi-2bpw runs on a
    32 GB Mac at about 20 tok/s.
  • Prompt cache with 0 GB of RAM: prompts are reused across turns from an SSD cache that sizes itself (up to 20 GB);
    keeping them in RAM is now opt-in with --prefix-cache-mem.
  • Faster: MiMo decodes about 12% and prefills about 15% faster, Qwen3.8-Flash-Next prefills faster, and concurrent
    requests decode together on every model.
  • Agents and API: sushi launch grok is new and opencode 2.x works again; streamed and non-streamed answers match
    byte for byte; penalties, repetition_penalty and ignore_eos work; GLM history renders exactly as its template.
  • Flags: --mtp-min-depth/--mtp-max-depth replace --mtp-depth; new --no-mtp-lookup and --gpu-warm-secs;
    --wired-margin-gib defaults to 4 GiB.
  • Reliability: long GLM sessions, memory pressure and restarts are handled cleanly, and an idle streamed server no
    longer uses CPU.

Thanks @cnsiva (request-budget defaults, repetition_penalty), @jasontitus (unique response IDs, the M1–M4 GLM fix,
portable GLM tests), @ShoichiTect (idle SSD-read workers) and @gomezvd (MTP on streamed Qwen).

Don't miss a new sushi release

NewReleases is sending notifications on new releases.