github beamivalice/sushi v1.2.1
sushi v1.2.1

4 hours ago

v1.2.1 — Safer prompt cache and server hardening

  • Prompt cache: two sushi processes on the same model no longer share one SSD cache folder, which could return
    another session's answer; a restart on a nearly full disk keeps the cache, and loading a second model no longer
    deletes the first one's entries.
  • Crash fixes: an empty /v1/completions prompt, deeply nested JSON and oversized WebSocket messages are refused
    with a 400 instead of stopping the server.
  • GLM-5.3-Flash: after a client disconnects, the next request on a streamed GLM no longer fails; sushi pull
    now fetches the DFlash2 assistant, and deleting its BF16 source keeps DFlash2 on.
  • Tools and structured output: earlier tool calls with empty or non-object arguments render in the model's own
    format, anyOf/oneOf/$ref parameters get their real types, and JSON-schema output stops cleanly on numbers and
    bounded arrays; logit_bias: null is accepted again.
  • CLI: sushi launch quotes model names, sushi serve without --model honours the sampling flags, update
    no longer logs the API key, and sushi run accepts long pasted lines and filters terminal escapes from model output.
  • mlx-serve: the guest manifest now lists GLM-5.3-Flash.

Don't miss a new sushi release

NewReleases is sending notifications on new releases.