github mostlygeek/llama-swap v261

4 hours ago

This release has a bunch of small compatibility improvements requested by the community.
Also the changelog summaries have been changed to only AI for the list below. These notes at
the top will be written by humans (or my dog) for humans. The AI's writing more correct but
fuck it we have too much AI slop in our lives already.

  • PR #1189 Increase ComfyUI concurrency limit from 50 to 999: fix ComfyUI stuck on the splash screen via /comfyui because one tab makes more parallel requests than the old minimum of 50 allowed (#1188) by @mostlygeek
  • PR #1187 docker/unified: compile sm_70 in the CUDA 12 image for V100: fix ik-llama-server aborting on V100 GPUs, where only sm_61 PTX without a WMMA flash-attention kernel was available (#1185) by @mostlygeek
  • PR #1158 Hardware Detection: detect Intel GPUs on Linux via xpu-smi, including dedicated Arc memory, and read the Metal version on macOS across system_profiler key names by @anantshri
  • PR #1183 docs/kb: document that tailcat.models accepts selectors: note that aliases, selectors, profile pins and peer model names are accepted, with a selector example by @mostlygeek
  • PR #1182 Build audio.cpp with static eSpeak-ng and install phoneme data: link eSpeak-ng statically and ship its phoneme data so Kokoro, SanoTTS and Inflect phonemization works without a writable cache (#1181) by @mostlygeek

Don't miss a new llama-swap release

NewReleases is sending notifications on new releases.