🚀 New features
- Audio input for omni and MLX models — record and play audio right in the chat
- The app now remembers window size and position across restarts
- Download a model during onboarding and you'll jump straight into chat, with progress still visible
- Smarter sampling: recommended defaults for Gemma 4 QAT models, while keeping any values you've customized
- MLX models now have a proper Download section on the Hub (name, MLX badge, file size, Download / New Chat)
🔧 Improvements & fixes
- Fixed CUDA and Vulkan auto-detection on Windows — the right GPU backend now gets picked and downloaded automatically
- More reliable backend downloads overall (shared / NAT / VPN networks, fewer GitHub rate-limit failures, offline fallback preserved)
- Better Metal GPU recovery: clearer out-of-memory messages and auto-reload after compute errors
- Fixed system prompt being dropped in regular local-model chats
- Clickable links in the model-policy error banner
- Quieter logs (no more misleading CORS warnings for non-browser clients)
🙏 Contributors
Thanks to @Vect0rM, @Albert-Atomic for their contributions to this release!