Splash 1.3.0 accelerates Qwen3.8-27B prefill by running its feed-forward layers across the GPU and Apple Neural Engine, reducing time to first token for long prompts. It also improves shared-prefix reuse, burst serving, model validation and agent connections.
Neural Engine prefill
- Enable GPU/Neural Engine FFN splitting by default for supported dense targets. The engine calibrates the split for each Mac and model, selects the prompt sizes that benefit, and remembers the result. Short prompts, Qwen3.6-35B-A3B MoE and decoding remain on the GPU.
- Use Hadamard-rotated W8A8 computation for the Neural Engine's part, staging weights from the existing model rather than keeping an int8 copy of the whole model.
- Stop the split and rerun the affected chunk on the GPU after an evaluation failure, timeout or non-finite result. An automatically selected split also stops if it loses to the GPU alone.
splash serve --disable-aneselects the GPU-only path. - Report the split's state, share, minimum prompt-chunk size, evaluations, timing and reruns under
ane_ffnin/status. Unload the Neural Engine program at idle release and reload it for the next request.
Measured cold-prompt TTFT for Qwen3.8-27B, one request:
| Mac | Weights | Input tokens | GPU only | GPU + Neural Engine | Speedup |
|---|---|---|---|---|---|
| M5 Max | MLX 4-bit | 14,096 | 14.09 s | 12.36 s | 1.14x |
| M5 Max | UD-Q4_K_M GGUF | 14,096 | 13.74 s | 12.24 s | 1.12x |
| M5 Pro | MLX 4-bit | 14,096 | 28.28 s | 21.57 s | 1.31x |
| M6, 24 GB | UD-IQ3_XXS GGUF | 14,096 | 45.77 s | 30.8–31.3 s | 1.46–1.49x |
| M3 Max, Low Power Mode | MLX 4-bit | 14,096 | 91.04 s | 62.43 s | 1.46x |
| M4 Max, 40-core GPU, 128 GB (community) | Q8_0 GGUF | 14,328 | 77.6–78.1 s | 40.7–45.9 s | 1.77x median |
The M4 Max community benchmark tested the final PR revision 80827e1, whose production engine sources match this release. It also measured 1.71x at 2,040 tokens, with TTFT falling from 8.5–9.2 s to 5.2–5.4 s. That run measured prefill, not decode or accuracy. The other rows are measurements recorded in the merged implementation's tests.
These are prefill measurements, not decode speedups. First startup can take longer while programs compile and the split calibrates. The split consumes memory that would otherwise hold KV cache: the measured 24 GB M6 automatic context was 61,433 tokens versus 73,721 with --disable-ane. An explicit context limit can reduce the split or select the GPU alone. The W8A8 path is numerically different from GPU-only execution; startup checks each program function against the GPU. The private ANE interface falls back to GPU-only execution when unavailable.
Shared-prefix caching
- Keep the last prefill checkpoint beside a completed prompt's replay state when space permits, so the next conversation sharing a long prefix can resume before its point of divergence instead of prefilling from zero (#301).
- In a measured 19K-token shared-prefix case, the second conversation's TTFT fell from 39.6 s to 6.3 s, reusing 16,384 tokens. This retains one additional reclaimable state per sufficiently long prompt (187 MiB on the 27B); with SSD caching, eviction can write that state to disk.
Serving and APIs
- Admit incoming connections without creating a thread for each connection; use a fixed worker pool and deliver retryable 503 responses under overload instead of resetting burst connections.
- Return errors in the requested API's format, including early Anthropic Messages errors. Return server faults as server errors rather than misclassifying them as invalid JSON; include
Retry-Afterfor retryable systemone errors. - Bound JSON depth to 128 levels and tool/response-format schema depth to 64. Improve schema caches, cached PDF access, disconnect handling and answer-slot preparation for large systemone prompts.
- Keep
/statusreads from changing memory admission hysteresis, and report held requests using admission's own decisions.
Models, agents and compatibility
- Check model configuration and quantization using the engine's own rules before downloading weight files. Improve errors for unsupported models and malformed GGUF/configuration data, and validate the exact buffer and weight extents GPU kernels read.
- Add
--portto all five agent launchers, fixing connections to servers on non-default ports (#311). Put an agent's own--portafter--. - Keep API key values out of generated Hermes profiles, serialize Pi configuration updates, and let another model's download proceed without holding the environment setup lock.
- Publish Hub snapshot pins by atomic rename, fixing failures on filesystems such as ExFAT that do not support hard links.
- Tune selected single-lane paired N256 projection plans; improve persistent-cache restart coverage, installed-client checks and native build coverage in CI.
Upgrade
The native/server protocol changes from 7 to 8. Restart existing Splash servers after upgrading so both components run the new version. Existing installed GGUF models refresh their derived metadata at the next start without downloading their weights again.
Full changelog: 1.2.1...1.3.0