Splash 1.2.1 fixes requests failing across Mac sleep, preserves tool-call arguments, and makes idle memory release configurable.
Tool calls
- Fix tool fields being renamed or dropped when the model writes parameters in a different order, and string/number unions losing values such as
12:00or0800(#293, #294). - With automatic tool choice and no strict tools or response format, let the model write calls without an output grammar. Required/named tool choice, disabled parallel calls, strict tools and structured response formats retain their applicable constraints.
- Read arguments according to the chat template and the tool's declared JSON types. Preserve parameter contents containing tool-format tags, multiple calls and text after calls; a tool-call opening also ends unclosed reasoning.
- Return generated arguments to the client without post-generation schema validation; the client handles tool argument errors. Request schemas are still checked before generation, and strict tools retain schema constraints. Repeated parameters keep their first value so streaming and complete responses agree.
- Allow
ignore_eoswith unconstrained tools; it remains incompatible with output grammars.
Mac sleep
- Measure engine timeouts, keep-alives and execution durations using awake time. A sleep no longer consumes the Metal command watchdog, resource waits or idle memory timers (#275).
- Prevent automatic system sleep while requests run. The display may still sleep;
--allow-idle-sleepopts out. Explicit sleep and lid closure are still allowed. - A real-model request resumed after a four-minute sleep without an engine restart. Deep standby lasting hours and driver-reported GPU failures remain outside this fix. An explicitly timed request still in transit to the engine when sleep begins can expire on arrival.
Memory and timeout options
- Add
splash serve --idle-release DURATION|off(#292). The default remains10m; examples include30m,2hand90s.offkeeps weights allocated and backend buffers wired between requests. KV memory reclamation under pressure remains active. - Expose
weights.idle_release_seconds,weights.releasedandweights.restoresin/status; the interval isnullwhen idle release is disabled. - Accept seconds or an
s,morhsuffix for--request-timeout, using the same duration format as idle release. Existing numeric values remain seconds; omitting the option still means no request timeout.
Thinking and chat templates
- Pass
reasoning_effort: "none"as well asenable_thinking: falsewhen reasoning is disabled. This fixes non-thinking Chat and Anthropic Messages requests on Nex-N2.5-mini, an existing Qwen3.6-35B-A3B-compatible fine-tune; judgments and systemone use the same template options. - Read the request's
chat_template_kwargs.enable_thinkingonce: a boolean overrides the effort,nullfollows the request/server effort, and other values return 400. Explicittruewith effortnoneenables thinking at the template's default effort.
Full changelog: 1.2.0...1.2.1