github incoai/splash 1.3.1
Splash 1.3.1

5 hours ago

Splash 1.3.1 moves MLX targets and DFlash drafts onto the same block kernel library as GGUF targets, loads more checkpoints of the two supported model families, and reduces conversation-cache memory.

Breaking change: Splash packages

Splash packages (incoai/Qwen3.8-27B-Splash, incoai/Qwen3.6-35B-A3B-Splash and community packages) no longer load (#335). The MLX models below load the same weights as the two official packages:

splash serve --model mlx-community/Qwen3.8-27B-4bit
splash serve --model mlx-community/Qwen3.6-35B-A3B-4bit

For a fine-tune, use its own MLX or GGUF release. For an installed package, the error names its Hub cache folder, which can be deleted.

Models

These changes apply to Qwen3.8-27B and Qwen3.6-35B-A3B checkpoints and compatible fine-tunes:

  • MLX affine formats at 2, 3, 4, 5, 6 and 8 bits, with groups of 32, 64 or 128, now load, including mixed precision. MLX mxfp4 checkpoints also load. For example: mlx-community/Qwen3.8-27B-8bit (#336).
  • GGUF repositories with an F16 vision projector can serve images. The projector must preserve the BF16 weights accepted by Splash's vision loader (#330).
  • MLX repositories that save the image processor in processor_config.json, as newer Transformers releases do, now serve images (#353). OptiQ checkpoints install with --language-only (#349).

Speed and memory

All MLX targets and drafts now use the block kernel library used by GGUF targets (#345, #346). In the release check against 1.3.0, on a 40-core M5 Max:

  • With 2–4 requests at once, the 27B MLX 4-bit target uses 4–26% less GPU time per decode step. Single-request GPU time is effectively unchanged.
  • The tested 27B and 35B UD-Q4_K_M GGUF targets use about 6–12% less GPU time per decode step at batch sizes 1–4.
  • The 27B and 35B MLX 4-bit targets use 3–16% less GPU time in the measured 4K-token cold and 544-row prefill runs.

On a 40-core M3 Max, the 35B MLX 4-bit target uses 8–18% less GPU time per decode step. Output tokens and draft acceptance matched 1.3.0 in all 45 backend-regression cases on each Mac.

A cached conversation state takes 153 MiB instead of 187 MiB on the 27B, and 64 MiB instead of 109 MiB on the 35B. The same cache budget can hold more conversation states (#329).

Other changes

  • /status and /metrics report the Mac's thermal state, and the log records each change, making heat-related slowdowns visible (#339).
  • /v1/messages and /v1/messages/count_tokens accept Claude Code tool search's tool_reference blocks inside tool results (#342, fixes #337).
  • The chat page renders Markdown and highlights code, pauses automatic scrolling while you read earlier output, and displays reply images as links. It also adds reply and code copying, individual conversation deletion, a collapsible sidebar and a bounded Thinking area (#343).
  • After an idle release, restoring weights waits for memory instead of stopping the engine on an allocation refusal. A restore that remains blocked for 30 seconds returns a retryable 503; the engine stays available for the next request (#347).

Thanks to @komarov-dc, @linson007 and @shixianqin for their contributions.

Upgrade

brew update
brew upgrade incoai/tap/splash

Restart the running server after upgrading. If it used a Splash package, change --model to an MLX or GGUF checkpoint as described above.

The persistent conversation cache written by 1.3.0 is not reused because cached states changed format. It refills as you use Splash; existing MLX and GGUF weight downloads can be reused.

Full changelog: 1.3.0...1.3.1

Don't miss a new splash release

NewReleases is sending notifications on new releases.