github incoai/splash 1.2.1
Splash 1.2.1

4 hours ago

Splash 1.2.1 fixes requests failing across Mac sleep, preserves tool-call arguments, and makes idle memory release configurable.

Tool calls

  • Fix tool fields being renamed or dropped when the model writes parameters in a different order, and string/number unions losing values such as 12:00 or 0800 (#293, #294).
  • With automatic tool choice and no strict tools or response format, let the model write calls without an output grammar. Required/named tool choice, disabled parallel calls, strict tools and structured response formats retain their applicable constraints.
  • Read arguments according to the chat template and the tool's declared JSON types. Preserve parameter contents containing tool-format tags, multiple calls and text after calls; a tool-call opening also ends unclosed reasoning.
  • Return generated arguments to the client without post-generation schema validation; the client handles tool argument errors. Request schemas are still checked before generation, and strict tools retain schema constraints. Repeated parameters keep their first value so streaming and complete responses agree.
  • Allow ignore_eos with unconstrained tools; it remains incompatible with output grammars.

Mac sleep

  • Measure engine timeouts, keep-alives and execution durations using awake time. A sleep no longer consumes the Metal command watchdog, resource waits or idle memory timers (#275).
  • Prevent automatic system sleep while requests run. The display may still sleep; --allow-idle-sleep opts out. Explicit sleep and lid closure are still allowed.
  • A real-model request resumed after a four-minute sleep without an engine restart. Deep standby lasting hours and driver-reported GPU failures remain outside this fix. An explicitly timed request still in transit to the engine when sleep begins can expire on arrival.

Memory and timeout options

  • Add splash serve --idle-release DURATION|off (#292). The default remains 10m; examples include 30m, 2h and 90s. off keeps weights allocated and backend buffers wired between requests. KV memory reclamation under pressure remains active.
  • Expose weights.idle_release_seconds, weights.released and weights.restores in /status; the interval is null when idle release is disabled.
  • Accept seconds or an s, m or h suffix for --request-timeout, using the same duration format as idle release. Existing numeric values remain seconds; omitting the option still means no request timeout.

Thinking and chat templates

  • Pass reasoning_effort: "none" as well as enable_thinking: false when reasoning is disabled. This fixes non-thinking Chat and Anthropic Messages requests on Nex-N2.5-mini, an existing Qwen3.6-35B-A3B-compatible fine-tune; judgments and systemone use the same template options.
  • Read the request's chat_template_kwargs.enable_thinking once: a boolean overrides the effort, null follows the request/server effort, and other values return 400. Explicit true with effort none enables thinking at the template's default effort.

Full changelog: 1.2.0...1.2.1

Don't miss a new splash release

NewReleases is sending notifications on new releases.