github unslothai/unsloth v0.1.804-beta
Qwen3.8-Flash-Next + GLM-5.3-Flash

5 hours ago

Qwen3.8-Flash-Next and GLM-5.3-Flash can now run locally in Unsloth!

  • Run Qwen3.8-Flash on 75GB RAM, GLM-5.3-Flash on 102GB RAM+VRAM
  • 5x Faster inference for RAM offloading
  • "Infinite" repeated compaction now works
  • 100+ chat, reliability and performance improvements

Qwen Guide: https://unsloth.ai/docs/models/qwen3.8-next
Qwen GGUFs: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF
GLM Guide: https://unsloth.ai/docs/models/glm-5.3-flash
GLM GGUFs: https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF

Highlights

  • Qwen3.8-Flash-Next on 75GB RAM
  • GLM-5.3-Flash on 102GB total memory
  • Smarter GPU + RAM offloading - run larger models with less setup
  • Chats recover after disconnects instead of losing the reply
  • See what fits before loading with clearer memory estimates

Qwen3.8-Flash-Next

Qwen3.8-Flash-Next is a new 125B multimodal reasoning model and an early preview of Qwen4's architecture.

  • The 1-bit Unsloth Dynamic GGUF runs on 75GB RAM or unified memory.
  • It's 79% smaller than BF16 while retaining 80% top-1 accuracy.
  • Chat with text and images using up to 262K context.
  • Switch between None, Low, Medium and Extra High reasoning.
  • Preserved Thinking keeps reasoning consistent across longer chats.

GLM-5.3-Flash

GLM-5.3-Flash is Z.ai's new 320B multimodal model, with only 18B parameters active at a time.

  • Run the 1-bit model on 102GB of combined RAM + VRAM.
  • Chat with text, images and long documents using up to 1M context.
  • Switch between Low, High and Max reasoning.
  • Stronger coding, agent and vision performance than GLM-5.2.
  • Recommended settings are applied automatically in Unsloth.

Chat + tools

  • Local chats resume after a disconnect instead of losing the reply.
  • Deep Research keeps going when a provider asks it to slow down.
  • Vision chats now handle multiple images properly.
  • Images returned by MCP tools appear directly in chat.
  • Export chats as JSONL for backups or use in other tools.
  • Adjust Auto Compaction for longer chats, or turn it off.
  • Collapse tool activity by default for cleaner agent chats.

Models + performance

  • Large GGUFs automatically split across GPU and system RAM.
  • See estimated memory usage before loading a model.
  • View VRAM usage directly from your downloaded models.
  • Model settings stay saved when switching chats.
  • Search and download embedding models directly from Hugging Face.
  • Text-to-speech models only load when you actually use them.

Desktop + reliability

  • Linux voice recording fixed.
  • NVIDIA + Wayland interface freezes fixed.
  • AMD model loading crashes fixed.
  • llama.cpp models now load from Windows profiles with non-English characters.
  • Non-English web links now work properly as chat sources.
  • Desktop download links always point to the latest stable release.

What's Changed

New Contributors

Full Changelog: v0.1.803-beta...v0.1.804-beta

Don't miss a new unsloth release

NewReleases is sending notifications on new releases.