github unslothai/unsloth v0.1.806-beta
2x Faster Qwen3.8-Flash + GLM-5.3-Flash MTP

3 hours ago

Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it.
Also our new release includes 170+ training, chat, hardware, and performance improvements.

Highlights

  • Smoother model loading (less errors) across local servers and connected providers.
  • Safer chat edits that preserve tool cards, reply details, and conversation branches.
  • New local media APIs for video, audio, and MLX-served models.
  • New audio support with new models, progress tracking including: MiniMax-Music3, Higgs, MOSS and more!
  • Improved multi-GPU planning, memory fitting, and split-model training.
  • Strengthened AMD/ROCm detection, installation, and GPU compatibility.
  • Upgraded MCP, Deep Research, OAuth, and agent tool reliability.

Qwen3.8-Flash + GLM-5.3-Flash

  • Qwen and GLM now generate faster with MTP enabled by default.
  • Use GLM tools across longer, multi-turn chats.
  • Qwen automatically applies the recommended settings for thinking and non-thinking modes.

Download Qwen3.8-Flash-Next and GLM-5.3-Flash. See the Qwen guide and GLM guide for recommended settings and available GGUFs.

Faster MLX inference

  • Fine-tune both large MoE models with text or images on Apple Silicon using MLX.
  • Long Qwen chats now run much faster on Mac, with follow-up turns up to 30x faster.
  • MLX models now use their full context size and support much longer batched generation.
  • MLX releases GPU memory more cleanly between generation bursts and model switches.
  • Serve MLX models through Unsloth's OpenAI-compatible API.

Audio

  • Added support for MiniMax-Music3, Higgs, MOSS audio models.
  • Added live progress updates while audio is being generated.
  • Audio clips can now be archived and managed.
  • Improved reliability with custom TTS playback fixes, Whisper pairing checks, and stronger audio testing.

Chat + tools

  • Run several tool calls at once without mixing up their arguments.
  • Keep tools available when chatting with images.
  • Each chat keeps its MCP connection for faster tool calls.
  • Local models can edit code using Codex’s apply_patch tool.
  • Continue long chats with images and other media using Auto Compaction.
  • Review and approve Deep Research plans before research starts.

Training + hardware

  • Train larger models across multiple GPUs with automatic placement.
  • AMD installs choose the best build across Windows and Linux, with BF16 on more GPUs.
  • Export GLM-5.3 MLX fine-tunes to GGUF.
  • Choose custom GGUF shard sizes and save locations.

API + Desktop

  • Generate videos through the new OpenAI-compatible Videos API.
  • Updates download in the background and install when you restart.
  • Choose a custom port for LAN access.
  • Generate audio with Higgs, MOSS and MiniMax models.
  • Track audio generation progress and archive finished clips.
  • Model downloads show clearer progress and can switch from Xet to HTTP automatically.

Download Unsloth Desktop

Unsloth Desktop is free and open source. Download it for:

  • Windows
  • macOS
  • Linux

🦥 Download Unsloth Desktop

What's Changed

New Contributors

Full Changelog: v0.1.804-beta...v0.1.806-beta

Don't miss a new unsloth release

NewReleases is sending notifications on new releases.