I've been busy, busy, busy and have not had as much time for llama-swap as I
would like. This release includes reliability improvements, improved
compatibility and some small quality of life improvements. Thank you for the
amazing contributors for this release! Everyone who files issues or submits a
PR is greatly appreciated.
- PR #1213 Add confirmation UI and progress tracking for WoL wake: wol-proxy gets a
-require-confirmflag that shows a "Start server" page instead of waking on every request, plus a retro loading page with a progress bar estimated from the last wake duration by @mostlygeek - PR #1215 internal/router: strip browser headers before proxying to peers: fix UI chats failing with a 401 against an api.anthropic.com peer, which treated the forwarded browser Origin header as a direct browser call; Origin, Referer, Sec-Fetch-* and Access-Control-Request-* are now removed (#1214) by @mostlygeek
- PR #1212 internal/process: report loading progress from the health check: parse
progress/messagefrom not-ready health check JSON bodies, show it on the Models and model detail pages, and restart thehealthCheckTimeoutcountdown when progress changes so long loads don't time out (#1079, #1208) by @mostlygeek - PR #1196 internal/perf: restart loop to trigger GPU collector when the monitor still running: fix GPU stats stopping for good after
nvidia-smiexited; the exit error is now logged and the collector restarts with a 5 to 30 second backoff (#1155) by @hpdkhoa - PR #1211 internal/process: add a per-model disableKeepAlives option (#1205): fix intermittent 502
proxy error: EOFwhen a POST reused a pooled connection that llama-server closed after a streamed response;disableKeepAlives: trueopens a new connection per request by @johnnieCR - PR #1199 fix(metrics): show oMLX token rates in Activity: read oMLX's
usage.prompt_tokens_per_secondandusage.generation_tokens_per_secondas a fallback so rates no longer show as unknown by @junmo-kim - PR #1198 internal/capcompat: add Gufo inference engine prober: add a default capability prober that uses only
/v1/modelsdata as the fallback for servers without a custom prober (vLLM, Gufo, SGLang, etc.), replacing the vLLM-specific one by @berney