Splash 1.2.0 adds persistent conversation caching, OpenAI text completions, more sampling controls, and broader GGUF compatibility. It also unifies Metal memory residency, replaces sparse KV mapping, improves concurrent serving and memory handling on 24 GB Macs, and removes the extra prepared-weight copy on disk.
Persistent conversation cache
- Add
--persistent-cachewith--max-cache-diskto reuse durable conversation prefixes after a clean restart, a compatible upgrade, or a crash.--cache-dirselects the location; the default is~/Library/Caches/Splash/prefix-cache. - Persist recent conversation restore points in the background, including points still in RAM. A clean shutdown flushes pending restore points for up to six seconds. Crash recovery restores points that reached disk, rather than guaranteeing every recent token was saved.
- Isolate caches by model manifests and cache layouts, lock each namespace, bound storage by the configured quota, and validate recovered data. Invalid or incompatible entries fall back to recomputation.
- Persistence is off by default.
--max-cache-diskalone keeps its temporary, session-local behavior. Persistent mode can write more to SSD when RAM is plentiful.
Metal memory residency and KV cache
- Extend the residency management previously used for weights to every backend buffer: weights, KV extents, recurrent state, draft rings and scratch. Buffers join one residency set attached to the command queue and stay wired between requests during active serving, avoiding repeated residency declarations per dispatch. This does not keep weights resident indefinitely: residency lapses after ten minutes without a command, and idle weight memory is explicitly released as described below.
- Replace placement-sparse KV buffers and 64 KiB heap map/unmap operations with ordinary shared Metal buffer extents, reached through GPU-address page tables. Remove the sparse mapping queue, event waits, mapping timeout and startup sparse-support probe, eliminating that uninterruptible mapping path during serving and teardown.
- Move SSD KV eviction and restoration directly between shared extents and the disk tier on its I/O worker. Remove the GPU staging ring, KV-copy kernel and copy-only GPU commands; SSD-offload configurations regain the staging ring's memory, accounting for the state staging buffer.
- Compact live KV pages into other extents so empty extents can actually be released to macOS. Memory-pressure recovery now returns allocated memory, rather than only evicting cached pages inside still-allocated extents.
- Stop submitting GPU work during backend teardown. Residency expiry uses a precompiled minimal compute kernel instead of a lazily compiled blit, improving shutdown and recovery behavior.
APIs, sampling and structured output
- Add
POST /v1/completionsfor a raw text prompt or token IDs, with streaming and usage reporting. This initial endpoint supports one completion; prompt batches, suffix, echo, logprobs, best-of and multiple choices are unsupported. Its default output limit is 16 tokens. - Add
ignore_eosto Chat and text completions. It suppresses model stop tokens while respecting explicit stop strings and the output budget; it cannot be combined with tools or structured-output grammars. - Add
presence_penalty,frequency_penalty,repetition_penaltyandmin_pto Chat, text completions and Responses. Replace the previous top-32 candidate sampling path with exact top-k/top-p over the whole vocabulary, including under DFlash speculative decoding. Arbitrary top-k values are accepted;0and-1disable the top-k filter. Omitted penalties remain neutral. Non-emptylogit_biasremains unsupported; null and an empty object are accepted. - Enforce required and named
tool_choiceand strict tool schemas across Chat, Responses and Messages. Fix schema-pattern propagation, bound constrained whitespace and referenced patterns, and reject unsupported schemas before prefill. - Accept
chat_template_kwargs, includingenable_thinking, with checks for reserved fields and conflicting reasoning settings. The built-in chat page now defaults to the server's configured reasoning effort. - Preserve text emitted after tool calls and the chronological block order in Messages and Responses. Fix incremental tool-argument streaming, whitespace around tool-call framing, partial calls at the output limit, and streamed/non-streamed output consistency. Remove leading answer newlines immediately after
</think>and avoid replacement characters when a token limit cuts a UTF-8 character.
Serving and client behavior
- Add
--decode-shareto balance active decoding against contended prefill. The default is0.5;0retains one-command alternation. Improve short-prefill completion, priority-based lane preemption and resumption, and scheduling while other lanes wait for disk I/O or grammar masks. - Add repeatable
--allowed-originfor explicitly permitted browser/app origins, including custom schemes used by desktop apps. - Add
--announce-served-nameso API responses report the first--served-model-namealias./v1/modelslists it first;/statuscontinues to identify the loaded model. - Add
splash serve --offlineto start an already installed model without the Hub probe. It does not download missing models. - Forward
--request-timeoutand--queue-sizethroughsplash serve. Requests no longer have a default 1,800-second timeout; an explicit timeout remains available. - Chat and Responses requests that omit an output limit can use the remaining context instead of an implicit 32,768-token cap. Explicit limits receive clearer context-capacity errors. Messages reports
model_context_window_exceededwhen context is exhausted; text completions retains its 16-token default. - Add Fish completion for commands and installed models, installed by Homebrew alongside Bash and Zsh completions.
- Run Hermes in a dedicated profile under the user's Hermes root and fix argument forwarding. Configure clients from
/v1/modelswithout waiting for/status. Pi receives a 32K output allowance, lowered by Pi to the remaining context; OpenCode and Hermes retain a quarter-window allowance for contexts below 128K. - Give the built-in chat page a favicon and keep it from scrolling or changing its waiting state on empty keepalive chunks.
Model loading, memory and performance
- Load MLX, GGUF, draft and vision weights directly into Metal buffers from the original model files at each start. Remove the extra prepared-weight disk copy and its preparation cache. After ten idle minutes, release weight memory and reload from the original files on the next request, while continuing to answer status and cancellation messages.
- Support GGUF token/embedding tables stored as
IQ4_XS,IQ4_NLorIQ3_S, and quantized GDN alpha/beta gate pairs supported by the loader. Accept the legacy RoPE type key and whole-number floating values in model configurations. - Probe Apple GPU family 11 explicitly. Improve affine Q4 split-K execution on Apple10/Apple11 (M5/M6), reduce decode CPU/driver overhead, prepare sampling pipelines at startup, and improve shared output-head and draft execution.
- Fix partial-tile GGUF prefill synchronization and Q4 prefill shader validation on macOS 26.5. Fix vision FP32 residual parity.
- Improve serving on memory-squeezed 24 GB Macs: preserve a minimal serving footprint during memory warnings, return empty KV allocations to macOS, allow admitted requests to grow within the configured budget, and avoid repeatedly restarting queue wait limits.
- Keep the reusable conversation state before the chat template's generation prompt, and prioritize unfinished requests' replay points for the next turn. Improve checkpoint/branch-point placement and prefix leases, and let shared-prefix requests wait for a producer restoring from SSD. Close admission behind requests waiting for memory so later small requests cannot indefinitely bypass them.
- Allocate vision scratch dynamically for the images being encoded instead of reserving its maximum working set at startup. Share image embeddings across requests and retained states, skip images already covered by a restored prefix, and reclaim unused vision arenas before useful embeddings. The advertised text context no longer pays the former fixed vision-scratch reserve.
- Fix SSD slot reclamation and quota accounting, including shrinking files and punching freed slots. Preserve a state's only copy when a disk restore cannot be promoted into the RAM cache. Disk-full or unwritable-tier failures disable affected writes while serving continues; filesystem allocation can briefly exceed the logical quota until freed slots are punched.
Reliability and compatibility
- Recover engines that fail or stop responding even when idle. Bound repeated crash/restart loops, report retryable failures while recovering, and return
engine_failedif recovery stops./readynow follows engine readiness without treating an ordinary busy tick as a failure. - Isolate request-level sampling failures, keep native input and cancellation responsive, bound grammar work, cancel image/document preparation on deadlines, and discard status from replaced engines. A grammar mask left unanswered for five seconds fails that request with retryable
mask_timeoutinstead of holding its batch./readyremains healthy under memory warning and returns 503 under critical pressure. - Return JSON HTTP errors for malformed requests, accept matching Bearer and
x-api-keycredentials together, and improve responses to overloads, stalled uploads and connection bursts. Full queues return 503 withRetry-After. - Messages requests ending in an assistant prefill now return 400 rather than silently starting a new assistant turn.
- Seeded sampled output can differ from earlier versions because draft proposals and sampling execution changed; a seed does not promise identical token streams across versions or different concurrent/cache schedules.
- Direct users of the old server CLI should remove
--max-new-tokensand set output budgets in requests.server.pyand nativeserve-nativenow take one model root directory instead of separate target/draft directories. - Monitoring consumers should update for status schema 6: cache/KV and lane fields are normalized and duplicate counters removed.
splash_ttft_secondsbecomessplash_http_ttft_seconds, and/status.latency.ttftbecomes/status.latency.http_ttft. Addmetrics.decode_cycle_msandsplash_decode_cycle_milliseconds_totalto include host work between decode completions. Native wire protocol is now 7; use the server and engine from the same package.
Developer tools
- Simplify fixed kernel execution plans and remove unused tuning/graph-confirmation paths. Add
weight-digestsand baseline weight-image comparison tooling. Consolidate architecture checks, protocol golden vectors, memory-policy tests and real-client test harnesses. These changes support development and release verification.
Upgrade
Requires Apple M3 or newer and macOS 26.4+. After the Homebrew tap update is merged:
brew update
brew upgrade incoai/tap/splash
splash --versionStop running Splash servers before upgrading, then restart them. Existing model downloads and user data remain reusable. The old ~/Library/Caches/Splash/weights prepared cache is no longer used and can be deleted manually; SPLASH_WEIGHT_CACHE is removed. Keep the original model files available: startup and restoration after idle now read them directly, including from external disks.
Hermes users whose older Splash launcher installed managed tools under the old Splash runtime directory should run hermes pm install in a normal shell to restore launchers to the user's Hermes tools before removing that old directory. Existing sessions there can be retained for migration.
Optional persistent cache:
splash serve --model mlx-community/Qwen3.6-35B-A3B-4bit --max-cache-disk 16G --persistent-cache