Splash 1.3.1 moves MLX targets and DFlash drafts onto the same block kernel library as GGUF targets, loads more checkpoints of the two supported model families, and reduces conversation-cache memory.
Breaking change: Splash packages
Splash packages (incoai/Qwen3.8-27B-Splash, incoai/Qwen3.6-35B-A3B-Splash and community packages) no longer load (#335). The MLX models below load the same weights as the two official packages:
splash serve --model mlx-community/Qwen3.8-27B-4bit
splash serve --model mlx-community/Qwen3.6-35B-A3B-4bitFor a fine-tune, use its own MLX or GGUF release. For an installed package, the error names its Hub cache folder, which can be deleted.
Models
These changes apply to Qwen3.8-27B and Qwen3.6-35B-A3B checkpoints and compatible fine-tunes:
- MLX affine formats at 2, 3, 4, 5, 6 and 8 bits, with groups of 32, 64 or 128, now load, including mixed precision. MLX mxfp4 checkpoints also load. For example:
mlx-community/Qwen3.8-27B-8bit(#336). - GGUF repositories with an F16 vision projector can serve images. The projector must preserve the BF16 weights accepted by Splash's vision loader (#330).
- MLX repositories that save the image processor in
processor_config.json, as newer Transformers releases do, now serve images (#353). OptiQ checkpoints install with--language-only(#349).
Speed and memory
All MLX targets and drafts now use the block kernel library used by GGUF targets (#345, #346). In the release check against 1.3.0, on a 40-core M5 Max:
- With 2–4 requests at once, the 27B MLX 4-bit target uses 4–26% less GPU time per decode step. Single-request GPU time is effectively unchanged.
- The tested 27B and 35B
UD-Q4_K_MGGUF targets use about 6–12% less GPU time per decode step at batch sizes 1–4. - The 27B and 35B MLX 4-bit targets use 3–16% less GPU time in the measured 4K-token cold and 544-row prefill runs.
On a 40-core M3 Max, the 35B MLX 4-bit target uses 8–18% less GPU time per decode step. Output tokens and draft acceptance matched 1.3.0 in all 45 backend-regression cases on each Mac.
A cached conversation state takes 153 MiB instead of 187 MiB on the 27B, and 64 MiB instead of 109 MiB on the 35B. The same cache budget can hold more conversation states (#329).
Other changes
/statusand/metricsreport the Mac's thermal state, and the log records each change, making heat-related slowdowns visible (#339)./v1/messagesand/v1/messages/count_tokensaccept Claude Code tool search'stool_referenceblocks inside tool results (#342, fixes #337).- The chat page renders Markdown and highlights code, pauses automatic scrolling while you read earlier output, and displays reply images as links. It also adds reply and code copying, individual conversation deletion, a collapsible sidebar and a bounded Thinking area (#343).
- After an idle release, restoring weights waits for memory instead of stopping the engine on an allocation refusal. A restore that remains blocked for 30 seconds returns a retryable 503; the engine stays available for the next request (#347).
Thanks to @komarov-dc, @linson007 and @shixianqin for their contributions.
Upgrade
brew update
brew upgrade incoai/tap/splashRestart the running server after upgrading. If it used a Splash package, change --model to an MLX or GGUF checkpoint as described above.
The persistent conversation cache written by 1.3.0 is not reused because cached states changed format. It refills as you use Splash; existing MLX and GGUF weight downloads can be reused.
Full changelog: 1.3.0...1.3.1