v1.2.1 — Safer prompt cache and server hardening
- Prompt cache: two sushi processes on the same model no longer share one SSD cache folder, which could return
another session's answer; a restart on a nearly full disk keeps the cache, and loading a second model no longer
deletes the first one's entries. - Crash fixes: an empty
/v1/completionsprompt, deeply nested JSON and oversized WebSocket messages are refused
with a 400 instead of stopping the server. - GLM-5.3-Flash: after a client disconnects, the next request on a streamed GLM no longer fails;
sushi pull
now fetches the DFlash2 assistant, and deleting its BF16 source keeps DFlash2 on. - Tools and structured output: earlier tool calls with empty or non-object arguments render in the model's own
format,anyOf/oneOf/$refparameters get their real types, and JSON-schema output stops cleanly on numbers and
bounded arrays;logit_bias: nullis accepted again. - CLI:
sushi launchquotes model names,sushi servewithout--modelhonours the sampling flags,update
no longer logs the API key, andsushi runaccepts long pasted lines and filters terminal escapes from model output. - mlx-serve: the guest manifest now lists GLM-5.3-Flash.