github incoai/splash 1.2.0
Splash 1.2.0

one hour ago

Splash 1.2.0 adds persistent conversation caching, OpenAI text completions, more sampling controls, and broader GGUF compatibility. It also unifies Metal memory residency, replaces sparse KV mapping, improves concurrent serving and memory handling on 24 GB Macs, and removes the extra prepared-weight copy on disk.

Persistent conversation cache

  • Add --persistent-cache with --max-cache-disk to reuse durable conversation prefixes after a clean restart, a compatible upgrade, or a crash. --cache-dir selects the location; the default is ~/Library/Caches/Splash/prefix-cache.
  • Persist recent conversation restore points in the background, including points still in RAM. A clean shutdown flushes pending restore points for up to six seconds. Crash recovery restores points that reached disk, rather than guaranteeing every recent token was saved.
  • Isolate caches by model manifests and cache layouts, lock each namespace, bound storage by the configured quota, and validate recovered data. Invalid or incompatible entries fall back to recomputation.
  • Persistence is off by default. --max-cache-disk alone keeps its temporary, session-local behavior. Persistent mode can write more to SSD when RAM is plentiful.

Metal memory residency and KV cache

  • Extend the residency management previously used for weights to every backend buffer: weights, KV extents, recurrent state, draft rings and scratch. Buffers join one residency set attached to the command queue and stay wired between requests during active serving, avoiding repeated residency declarations per dispatch. This does not keep weights resident indefinitely: residency lapses after ten minutes without a command, and idle weight memory is explicitly released as described below.
  • Replace placement-sparse KV buffers and 64 KiB heap map/unmap operations with ordinary shared Metal buffer extents, reached through GPU-address page tables. Remove the sparse mapping queue, event waits, mapping timeout and startup sparse-support probe, eliminating that uninterruptible mapping path during serving and teardown.
  • Move SSD KV eviction and restoration directly between shared extents and the disk tier on its I/O worker. Remove the GPU staging ring, KV-copy kernel and copy-only GPU commands; SSD-offload configurations regain the staging ring's memory, accounting for the state staging buffer.
  • Compact live KV pages into other extents so empty extents can actually be released to macOS. Memory-pressure recovery now returns allocated memory, rather than only evicting cached pages inside still-allocated extents.
  • Stop submitting GPU work during backend teardown. Residency expiry uses a precompiled minimal compute kernel instead of a lazily compiled blit, improving shutdown and recovery behavior.

APIs, sampling and structured output

  • Add POST /v1/completions for a raw text prompt or token IDs, with streaming and usage reporting. This initial endpoint supports one completion; prompt batches, suffix, echo, logprobs, best-of and multiple choices are unsupported. Its default output limit is 16 tokens.
  • Add ignore_eos to Chat and text completions. It suppresses model stop tokens while respecting explicit stop strings and the output budget; it cannot be combined with tools or structured-output grammars.
  • Add presence_penalty, frequency_penalty, repetition_penalty and min_p to Chat, text completions and Responses. Replace the previous top-32 candidate sampling path with exact top-k/top-p over the whole vocabulary, including under DFlash speculative decoding. Arbitrary top-k values are accepted; 0 and -1 disable the top-k filter. Omitted penalties remain neutral. Non-empty logit_bias remains unsupported; null and an empty object are accepted.
  • Enforce required and named tool_choice and strict tool schemas across Chat, Responses and Messages. Fix schema-pattern propagation, bound constrained whitespace and referenced patterns, and reject unsupported schemas before prefill.
  • Accept chat_template_kwargs, including enable_thinking, with checks for reserved fields and conflicting reasoning settings. The built-in chat page now defaults to the server's configured reasoning effort.
  • Preserve text emitted after tool calls and the chronological block order in Messages and Responses. Fix incremental tool-argument streaming, whitespace around tool-call framing, partial calls at the output limit, and streamed/non-streamed output consistency. Remove leading answer newlines immediately after </think> and avoid replacement characters when a token limit cuts a UTF-8 character.

Serving and client behavior

  • Add --decode-share to balance active decoding against contended prefill. The default is 0.5; 0 retains one-command alternation. Improve short-prefill completion, priority-based lane preemption and resumption, and scheduling while other lanes wait for disk I/O or grammar masks.
  • Add repeatable --allowed-origin for explicitly permitted browser/app origins, including custom schemes used by desktop apps.
  • Add --announce-served-name so API responses report the first --served-model-name alias. /v1/models lists it first; /status continues to identify the loaded model.
  • Add splash serve --offline to start an already installed model without the Hub probe. It does not download missing models.
  • Forward --request-timeout and --queue-size through splash serve. Requests no longer have a default 1,800-second timeout; an explicit timeout remains available.
  • Chat and Responses requests that omit an output limit can use the remaining context instead of an implicit 32,768-token cap. Explicit limits receive clearer context-capacity errors. Messages reports model_context_window_exceeded when context is exhausted; text completions retains its 16-token default.
  • Add Fish completion for commands and installed models, installed by Homebrew alongside Bash and Zsh completions.
  • Run Hermes in a dedicated profile under the user's Hermes root and fix argument forwarding. Configure clients from /v1/models without waiting for /status. Pi receives a 32K output allowance, lowered by Pi to the remaining context; OpenCode and Hermes retain a quarter-window allowance for contexts below 128K.
  • Give the built-in chat page a favicon and keep it from scrolling or changing its waiting state on empty keepalive chunks.

Model loading, memory and performance

  • Load MLX, GGUF, draft and vision weights directly into Metal buffers from the original model files at each start. Remove the extra prepared-weight disk copy and its preparation cache. After ten idle minutes, release weight memory and reload from the original files on the next request, while continuing to answer status and cancellation messages.
  • Support GGUF token/embedding tables stored as IQ4_XS, IQ4_NL or IQ3_S, and quantized GDN alpha/beta gate pairs supported by the loader. Accept the legacy RoPE type key and whole-number floating values in model configurations.
  • Probe Apple GPU family 11 explicitly. Improve affine Q4 split-K execution on Apple10/Apple11 (M5/M6), reduce decode CPU/driver overhead, prepare sampling pipelines at startup, and improve shared output-head and draft execution.
  • Fix partial-tile GGUF prefill synchronization and Q4 prefill shader validation on macOS 26.5. Fix vision FP32 residual parity.
  • Improve serving on memory-squeezed 24 GB Macs: preserve a minimal serving footprint during memory warnings, return empty KV allocations to macOS, allow admitted requests to grow within the configured budget, and avoid repeatedly restarting queue wait limits.
  • Keep the reusable conversation state before the chat template's generation prompt, and prioritize unfinished requests' replay points for the next turn. Improve checkpoint/branch-point placement and prefix leases, and let shared-prefix requests wait for a producer restoring from SSD. Close admission behind requests waiting for memory so later small requests cannot indefinitely bypass them.
  • Allocate vision scratch dynamically for the images being encoded instead of reserving its maximum working set at startup. Share image embeddings across requests and retained states, skip images already covered by a restored prefix, and reclaim unused vision arenas before useful embeddings. The advertised text context no longer pays the former fixed vision-scratch reserve.
  • Fix SSD slot reclamation and quota accounting, including shrinking files and punching freed slots. Preserve a state's only copy when a disk restore cannot be promoted into the RAM cache. Disk-full or unwritable-tier failures disable affected writes while serving continues; filesystem allocation can briefly exceed the logical quota until freed slots are punched.

Reliability and compatibility

  • Recover engines that fail or stop responding even when idle. Bound repeated crash/restart loops, report retryable failures while recovering, and return engine_failed if recovery stops. /ready now follows engine readiness without treating an ordinary busy tick as a failure.
  • Isolate request-level sampling failures, keep native input and cancellation responsive, bound grammar work, cancel image/document preparation on deadlines, and discard status from replaced engines. A grammar mask left unanswered for five seconds fails that request with retryable mask_timeout instead of holding its batch. /ready remains healthy under memory warning and returns 503 under critical pressure.
  • Return JSON HTTP errors for malformed requests, accept matching Bearer and x-api-key credentials together, and improve responses to overloads, stalled uploads and connection bursts. Full queues return 503 with Retry-After.
  • Messages requests ending in an assistant prefill now return 400 rather than silently starting a new assistant turn.
  • Seeded sampled output can differ from earlier versions because draft proposals and sampling execution changed; a seed does not promise identical token streams across versions or different concurrent/cache schedules.
  • Direct users of the old server CLI should remove --max-new-tokens and set output budgets in requests. server.py and native serve-native now take one model root directory instead of separate target/draft directories.
  • Monitoring consumers should update for status schema 6: cache/KV and lane fields are normalized and duplicate counters removed. splash_ttft_seconds becomes splash_http_ttft_seconds, and /status.latency.ttft becomes /status.latency.http_ttft. Add metrics.decode_cycle_ms and splash_decode_cycle_milliseconds_total to include host work between decode completions. Native wire protocol is now 7; use the server and engine from the same package.

Developer tools

  • Simplify fixed kernel execution plans and remove unused tuning/graph-confirmation paths. Add weight-digests and baseline weight-image comparison tooling. Consolidate architecture checks, protocol golden vectors, memory-policy tests and real-client test harnesses. These changes support development and release verification.

Upgrade

Requires Apple M3 or newer and macOS 26.4+. After the Homebrew tap update is merged:

brew update
brew upgrade incoai/tap/splash
splash --version

Stop running Splash servers before upgrading, then restart them. Existing model downloads and user data remain reusable. The old ~/Library/Caches/Splash/weights prepared cache is no longer used and can be deleted manually; SPLASH_WEIGHT_CACHE is removed. Keep the original model files available: startup and restoration after idle now read them directly, including from external disks.

Hermes users whose older Splash launcher installed managed tools under the old Splash runtime directory should run hermes pm install in a normal shell to restore launchers to the user's Hermes tools before removing that old directory. Existing sessions there can be retained for migration.

Optional persistent cache:

splash serve --model mlx-community/Qwen3.6-35B-A3B-4bit --max-cache-disk 16G --persistent-cache

Full changelog

Don't miss a new splash release

NewReleases is sending notifications on new releases.