github unslothai/unsloth v0.1.902-beta
Command Palette + Desktop UI/UX

3 hours ago

This release brings faster navigation, shareable run settings, and clearer errors to Unsloth Desktop. It also 4x speeds up Laya decisions, expands hosted Decision API support, and keeps NVFP4, INT4, and MXFP4 checkpoints in 4-bit during LoRA training.

Highlights

  • Command palette on Cmd/Ctrl+P. Also, you can now share your GGUF model's run settings
  • Laya decisions up to 4.1x faster, plus connections to TypeSafe and other hosted providers
  • Clearer errors that explain failed loads and generations, with a View logs button
  • NVFP4 LoRA with 45% lower peak memory on Qwen3.8-27B-NVFP4
  • INT4 + MXFP4 checkpoints stay packed during LoRA, including Kimi-K2.7-Code and gpt-oss

Desktop + chat

  • Cmd/Ctrl+P opens a command palette, and GGUF run settings can be shared as a link or saved without loading.
  • Continue response extends a finished reply, and Fix with the model puts HTML canvas errors in the message box.
  • A warning flags models that only partly fit on the GPU, and the Images gallery loads thumbnails first.
  • The desktop app adds Cmd/Ctrl + and - zoom, Windows 11 style buttons on Windows and Linux, and a check before Update stops training.

Lower-memory checkpoint training

Pre-quantized 4-bit checkpoints now train on their exact published weights. Training reads them packed, so a loaded model needs about as much memory as its 4-bit checkpoint.

  • Qwen3.8-27B-NVFP4 peaks at 40.2 GB instead of 72.9 GB on one RTX PRO 6000, with 11% shorter steps.
  • Dense compressed-tensors INT4 checkpoints train as published on NVIDIA GPUs. Qwen3-8B w4a16 peaks at 8.5 GB, not 18.8 GB, on a B200.
  • Kimi-K2.7-Code trains on 6 B200s with its INT4 experts packed, where a 16-bit copy would need about 2 TB.
  • gpt-oss loaded with load_in_4bit = False keeps its MXFP4 experts, so gpt-oss-120b trains on one B200.

How it works

  • Triton kernels dequantize each layer on the fly, bit for bit matching compressed-tensors, and backward keeps only the packed weights.
  • INT4 stays exact instead of being re-quantized to NF4, which rounds every weight twice and raised Qwen3-8B perplexity by 1.6%.
  • For inference-only gpt-oss, UNSLOTH_MXFP4_KEEP_PACKED=0 restores the native load, which prefills faster.

Decision API + Laya

Laya decision models now answer short requests up to 4.1x faster, and hosted decision models work too.

  • Add TypeSafe or another hosted decision provider under Connections, then pick its models in Settings > API.
  • With the Decision API on, enable Unsloth Decisions in the chat's MCP menu for a decide tool.
  • On a B200, a 3-question guardrail check takes 2.8 ms instead of 11.6 ms. Long inputs run at about the same speed.
  • Laya loads in about 9 s instead of 17 s and needs about half the peak host RAM.

Image + video

  • Image models keep INT8 or FP8 when offloading. On a 10 GiB-capped B200, Z-Image takes 2.3 s, not 38.4 s.
  • Qwen-Image-2.1 renders 1536x1536 images up to 3.6x faster on a Radeon 8060S.

Training + models

  • Fine-tune T5, T5Gemma, BART, Marian and other text encoder-decoders with FastModel.
  • The new unsloth eval command runs lm-evaluation-harness tasks on a checkpoint or LoRA adapter.
  • Gemma 4 31B training and batched generation now work on models split across GPUs.

AMD + Apple Silicon

  • On Windows, RX 5000 through RX 9000 cards get PyTorch from AMD's multi-arch index.
  • MLX moves to mlx 0.32.3 and mlx-vlm 0.7.4, and quantized-KV vision loads can batch replies.
  • On a Mac, export LoRA adapters as PEFT or GGUF for GPU machines.

Download Unsloth Desktop

Unsloth Desktop is free and open source. Download it for:

Platform Link
Windows Download
macOS Download
Linux x64 / Ubuntu (deb) Download
Linux ARM64 / Ubuntu 24.04+ (deb) Download
Linux x64 (AppImage) Download
Windows ARM64 Download

What's Changed

  • Bump install.sh / install.ps1 pin to unsloth>=2026.9.12, unsloth-zoo>=2026.9.8 by @danielhanchen in #12209
  • Give the Kimi processor test's stand-in tokenizer convert_tokens_to_ids by @danielhanchen in #12206
  • Follow chat export calls in the cancelled-save contract instead of a fixed file list by @danielhanchen in #12203
  • Add Idefics3 support (Granite Docling VLM) by @danielhanchen in #4241
  • Repair the checks red on main after today's Mistral-format, GKD, laya, llm-compressor and HF token merges by @danielhanchen in #12218
  • Load compressed-tensors packed INT4 checkpoints for QLoRA without a 16-bit copy (exact packed INT4 by default, NF4 fallback), and train remote DeepSeek-V3 MoE code (Kimi-K2.7-Code) by @danielhanchen in #11537
  • fix(studio): prevent terminal tool hangs in credential scan (#12048) by @Souravrajvi0 in #12087
  • Default loader.py's Mistral-format names in the precision-conflict harness by @danielhanchen in #12229
  • Studio: say why a generation or a model load failed, and offer the log by @danielhanchen in #8804
  • Train Kimi-K3 with its MXFP4 experts kept packed and dequantized on the fly by @danielhanchen in #11750
  • fix: align GKD Liger loss with the native objective by @taking-lying-flat in #12007
  • Studio: resolve repo-relative image paths in the simple image+text VLM converter by @breken-ai in #11897
  • fix(studio): reuse hardware monitor drag and resize handling for API monitor by @Imagineer99 in #12204
  • Unsloth Studio (AMD/Windows): real torchao INT8 / FP8 export on Windows ROCm by @danielhanchen in #7102
  • Studio: clamp unsupported reasoning effort to offered levels by @oobabooga in #12208
  • Studio: load Laya without randomly initialising its vocabulary embedding by @danielhanchen in #12210
  • Scope the runtime encoding scan past vendored laya while its loader supplies UTF-8 by @danielhanchen in #12230
  • Studio: load Images strip tiles as thumbnails instead of full PNGs by @oobabooga in #12228
  • Images: smaller prompt box radius, inset scrollbar, taller default by @shimmyshimmer in #12223
  • Library delete dialog: no dark mode border, cancel on outside click by @shimmyshimmer in #12222
  • Studio: laya download confirm by @NilayYadav in #12196
  • feat(studio): default the model source to ModelScope where Hugging Face is restricted by @Lyxot in #12197
  • Studio: fix prompt timestamp layout, accessibility, and imported times by @shimmyshimmer in #12173
  • Studio: explain failed and recurring llama.cpp runtime repairs by @umran666 in #12106
  • Studio: faster Laya inference (marker-only head, CUDA graphs) by @danielhanchen in #12224
  • Library: equal size chat cards and no loading placeholders by @shimmyshimmer in #12221
  • Run the user message timestamp Playwright check in Frontend CI by @danielhanchen in #12255
  • Studio: say when a model only partly fits on the GPU by @danielhanchen in #8812
  • Studio: push only the exported merged model to the Hub on a Mac by @NilayYadav in #12239
  • Fix sidebar More menu hover and collapsed border hover by @Sneakr in #12216
  • Studio: keep earlier messages when you send audio by @NilayYadav in #12244
  • Studio: show the model a phone photo the right way up by @NilayYadav in #12241
  • Studio: name the llama.cpp graph-scheduler abort and stop reloading into it by @danielhanchen in #6741
  • Fix NaN gradients from compiled flex attention on short static training batches by @danielhanchen in #12207
  • Studio: stop chats with images from blocking other chats on GGUF models by @NilayYadav in #12236
  • Studio: stop tool calls from getting stuck in Running forever by @NilayYadav in #12234
  • Studio: read a remote image URL on safetensors and MLX models instead of ignoring it by @NilayYadav in #12243
  • Composer attachments: first page previews, Library icons, lighter cards by @shimmyshimmer in #12261
  • Keep import unsloth working when vLLM needs transformers 5 by @danielhanchen in #12268
  • Fall back to PEFT's LoRA forward for float32 base weights by @danielhanchen in #7020
  • Reuse the dY @ B.t() product in the fused LoRA backward by @danielhanchen in #12212
  • Studio: let the chat model use the Decision API through MCP by @NilayYadav in #12232
  • Read the llama.cpp changelog after it settles, by content rather than innerText by @danielhanchen in #12273
  • Average GRPO gradients across DDP ranks, and keep MiniCPM3's pre-head scaling by @danielhanchen in #12193
  • Add Qwen2.5 Coder model support to registry and chat templates by @danielhanchen in #4221
  • Studio: do not fail a cached dataset load on a bookkeeping write by @BardiaKoopah in #8057
  • Wait for the model-config Load to finish before the test moves on by @danielhanchen in #12270
  • Scale the attachment preview text with the UI font size by @danielhanchen in #12276
  • Unsloth Studio: explain why Model Memory skips the RAM lock when a model is fully on the GPU by @danielhanchen in #9571
  • Studio: title a chat with the connection that answered it, not the local model by @Lyxot in #9106
  • Fix packed and grouped Q2 GGUF variant detection by @Imagineer99 in #7033
  • Improve empty text dataset validation by @Imagineer99 in #6646
  • Studio: fit the context again when MTP is forced on a reload by @NilayYadav in #12231
  • Studio: show an error when Gemini fails a reply partway through by @NilayYadav in #12242
  • ci: enforce npm ci on Studio install paths (follow-up to #5604) by @danielhanchen in #5616
  • Studio: let the dataset check accept chat columns that training already supports by @NilayYadav in #11849
  • Unsloth Studio / Desktop: allow text-encoder nvfp4 on pre-Blackwell GPUs and explain TE precision refusals by @mahiatlinux in #11413
  • refactor(trainer): add warning for ignored eval_steps by @danielhanchen in #4225
  • fix: handle zero-strided tensors in fast_rope_embedding (#3781) by @danielhanchen in #4233
  • Studio: hide the stray scrollbar under the composer at fractional zoom by @shimmyshimmer in #12279
  • Unsloth Studio / Desktop (AMD): "Automatic" llama.cpp follows the installed torch on a mixed NVIDIA+AMD host by @LeoBorcherding in #12250
  • Unsloth Studio / Desktop (AMD): explain why an NVIDIA+AMD host installed with CUDA_VISIBLE_DEVICES="" shows no GPU by @LeoBorcherding in #12249
  • Studio desktop: let the sidebar shrink to 224px by @shimmyshimmer in #12283
  • Studio: reopen the artifact panel after it has been dragged shut by @danielhanchen in #12278
  • Studio: say when an attached PDF has no readable text by @NilayYadav in #12237
  • Studio: read attached HTML files in the encoding they declare by @NilayYadav in #12235
  • Decline the fused LoRA kernels under FSDP, where the weights are shard views by @danielhanchen in #11132
  • Studio: allow long inputs for embedding models by @NilayYadav in #11721
  • feat(mlx): support DDP in CLI training by @Lyxot in #7166
  • Studio: save the chat template a base model was trained with by @NilayYadav in #12240
  • Unsloth Studio (AMD RDNA 1, Windows): install ROCm torch for the RX 5700 XT from AMD's multi-arch index by @LeoBorcherding in #11755
  • Studio: keep pop-up messages off the Run settings panel by @NilayYadav in #11486
  • Fix DDP "marked ready twice" for VLMs with CPU offload + TiledMLP by @danielhanchen in #4240
  • Studio: show download progress during desktop updates by @NilayYadav in #11291
  • Unsloth Studio / Desktop: add Canvas's console errors, and offer them to the model by @LeoBorcherding in #11744
  • fix(prompt storage): follow up, restore prompt list bookmarking, the kind pill and drag autoscroll by @LeoBorcherding in #9788
  • Raise the minimum typer version so the unsloth command starts by @NilayYadav in #11148
  • Unsloth Studio: run fp16 Laya checkpoints in fp16 on MLX by @Lyxot in #12256
  • Qwen3.5: skip compiled regions on eager decode steps, make CUDA graph decode opt-in by @danielhanchen in #12213
  • Studio: stop the stall watchdog from killing Xet downloads that are still receiving data by @NilayYadav in #11594
  • Studio: keep what you typed with a document when a long chat is compacted by @NilayYadav in #12238
  • Train LoRA on NVFP4 compressed-tensors checkpoints with the weights kept packed (W4A16) by @danielhanchen in #12185
  • Studio: keep working when GitHub rate limits or is down by @NilayYadav in #10670
  • Studio tests: forget the managed provider URL setting cache around each test by @danielhanchen in #12288
  • Unsloth Studio / Desktop: show the GPU llama.cpp runs on when it is CUDA or ROCm and differs from training by @LeoBorcherding in #12251
  • Unsloth Studio / Desktop: support custom vision projectors by @mahiatlinux in #11352
  • Studio: Respect the thinking budget on /v1/messages by @NilayYadav in #11724
  • Keep a model's no-placement parameters on CPU (Qwen3.8-Flash-Next n-gram table) by @danielhanchen in #12141
  • Studio: keep unsloth start alive while a LoRA adapter's base model downloads by @NilayYadav in #10550
  • Baseline the thirteen unsloth-zoo 2026.9.8 findings after review by @danielhanchen in #12292
  • Studio: show attachments when editing a message by @NilayYadav in #11054
  • Unsloth Studio installer (AMD/Windows): install PyTorch for RX 5000 to 9000 cards from AMD's multi-arch index by @LeoBorcherding in #11846
  • Fix: Support past_key_values in model.generate for multi-turn conversations by @danielhanchen in #4232
  • Studio: command palette (Cmd/Ctrl+P) by @NilayYadav in #6846
  • Unsloth Studio / Desktop: reload llama.cpp models on reconnect by @mahiatlinux in #11341
  • Harden timing-dependent Studio backend tests (nvidia-smi cache, cold /api/health) by @danielhanchen in #12275
  • Let FastModel and FastLanguageModel load the same family in one process, in either order by @danielhanchen in #12219
  • Studio: save a model's run settings without loading it by @NilayYadav in #10216
  • Keep LoRA on dense Linears sharing a name with fused MoE experts in PEFT's v5 conversion by @danielhanchen in #12281
  • Train gpt-oss MXFP4 LoRA through unsloth-zoo's packed experts when load_in_16bit is not set by @danielhanchen in #11929
  • feat(studio): reuse MLX VLM prompt snapshots in the resident vision batch by @Lyxot in #12263
  • Let Gemma 4 31B train when it is split across GPUs by @NilayYadav in #12233
  • Preserve repo id case when resolving a model name by @BardiaKoopah in #8058
  • feat(studio): Mac adapter-format option for LoRA export, with GGUF adapters on macOS by @Lyxot in #7539
  • Studio: list large MCP tool schemas compactly and load the full schema on demand by @NilayYadav in #11046
  • Support for Seq2Seq Models (T5, T5Gemma, etc.) by @danielhanchen in #4226
  • Studio: run large Laya requests in token-budgeted chunks by @danielhanchen in #12271
  • Studio: stabilize generalized compare mode (clean replacement) by @Imagineer99 in #7407
  • Add unsloth eval CLI command by @NilayYadav in #6824
  • Reload LoRA adapters trained with added tokens by @danielhanchen in #4219
  • Unsloth Studio / Desktop: expose prefill progress in API monitor by @mahiatlinux in #11161
  • Studio: one top row and layout fixes for the Windows desktop app by @mahiatlinux in #12070
  • Studio desktop: Cmd/Ctrl + and - zoom with a zoom popup by @shimmyshimmer in #12280
  • Studio: use the Hugeicons internet icon for every globe by @shimmyshimmer in #12321
  • Studio: use FileEmpty02Icon for generic file chips by @shimmyshimmer in #12324
  • Studio: mark adjusted backup timestamps as estimated by @shimmyshimmer in #12322
  • Studio: stop polling focus while message menus are open by @shimmyshimmer in #12323
  • Studio desktop tests: keep temp paths distinct when the clock repeats by @danielhanchen in #12296
  • Studio tests: check the compare composer's paste path, not one spelling of it by @danielhanchen in #12325
  • Tests: let the TiledMLP DDP workers import unsloth_zoo on a CPU runner by @danielhanchen in #12329
  • Studio: use Refresh01Icon and FileEmpty02Icon everywhere by @shimmyshimmer in #12331
  • Studio: Library grid cards keep one icon position and use the full width by @shimmyshimmer in #12330
  • Studio: tidy the Skills dialog rows and fields by @shimmyshimmer in #12332
  • Studio: Doc01 icon for Word and Google Docs files by @shimmyshimmer in #12340
  • Studio: lift Library grid icons to the card's middle by @shimmyshimmer in #12341
  • Studio: more edge padding on the model picker dropdowns by @shimmyshimmer in #12338
  • Studio: open the sidebar More menu on hover again by @shimmyshimmer in #12339
  • Studio: use the standard chevrons in place of Hugeicons' curved ones by @shimmyshimmer in #12337
  • Studio: stop normal chat from opening the API monitor by @Imagineer99 in #12305
  • Studio: fix Java initialization and home in the Linux tool sandbox by @oobabooga in #12294
  • Studio: don't re-prompt a finished code answer on safetensors and MLX models by @NilayYadav in #12309
  • Studio: keep answer columns out of the system prompt by @NilayYadav in #12306
  • unsloth start pi and dsh: stop cutting every reply at 8,192 tokens by @NilayYadav in #12315
  • Studio: flag responses that appear to stop while quoting a token by @oobabooga in #12259
  • Studio: keep web search source links for unsloth start claude by @NilayYadav in #12307
  • Move the attention mask to each layer's device in the fast decode loops by @danielhanchen in #12290
  • fix(studio): mark CSV exports as UTF-8 so Excel shows non-Latin text by @Mathews-Tom in #11971
  • Studio: keep code indentation when an HTML file is attached to a chat by @NilayYadav in #12312
  • Audio: keep PyAV decoding working on PyAV 19, and stop the tests needing an AMR encoder by @danielhanchen in #12295
  • Studio: accept multiple audio files per message by @shimmyshimmer in #12267
  • Parallel-isolation guard: exempt the nvidia-smi fake's poll deadline by @danielhanchen in #12353
  • Studio: ask before Update stops a training run in the desktop app by @NilayYadav in #12313
  • Validate unsloth train flags the way the config file is validated by @Abhishek-B-R in #11780
  • Fix three main CI failures from the Mac adapter export and seq2seq PRs by @danielhanchen in #12354
  • Studio frontend test: find the More flyout by its props in any order by @danielhanchen in #12356
  • Studio: chunk large Qwen-Image-2.1 attention queries on ROCm by @oobabooga in #12335
  • Studio desktop: survive AppKit exceptions during event dispatch on macOS by @shimmyshimmer in #12277
  • Studio: fix the faint seam beside the chat composer by @Imagineer99 in #12304
  • feat(studio): batch KV-quantized and TurboQuant MLX loads with prompt snapshot reuse by @Lyxot in #12343
  • Studio: load plain FP8 encoder safetensors without torchao by @oobabooga in #12333
  • Studio: let the Mac window be moved during an app update by @NilayYadav in #12316
  • Studio: keep menus and popovers below the desktop titlebar by @shimmyshimmer in #12348
  • fix(studio): record the bit widths of prequantized MLX models by @Mathews-Tom in #11940
  • Fix the ShareGPT mapping for Llama 3.1, Qwen and Gemma chat templates by @NilayYadav in #12314
  • Studio: report CUDA graphs as on once the deferred speed profile arms them by @danielhanchen in #12301
  • Keep packed INT4 layers off the fused inference kernel in train mode by @danielhanchen in #12349
  • Studio: isolate the Windows Terminal on MXC's default tier by @Etherll in #12328
  • Studio: prompt caching for OpenRouter by @Lyxot in #12264
  • Force non-reentrant gradient checkpointing for DeepSeek-V4.1 by @danielhanchen in #12318
  • Studio: count tokens for templates that refuse an empty chat by @shimmyshimmer in #12336
  • Studio: catch the tensor split abort behind a gdb backtrace by @oobabooga in #12293
  • Keep frozen BatchNorm running stats fixed during LoRA training by @danielhanchen in #12319
  • Widen the pinned child GPU mask when extra args name a companion device (#11810) by @deepspace28 in #11823
  • Studio: turn train on completions back on when leaving CPT by @NilayYadav in #12308
  • Studio: fix Qwen-Image-2.1 recompiles and Ideogram 4 attention and device errors on torch 2.11 to 2.14 by @danielhanchen in #12350
  • Studio: fix the desktop island corner and border seam by @Imagineer99 in #12285
  • Move the Unsloth Studio MLX pins to mlx 0.32.3 and mlx-vlm 0.7.4 by @Lyxot in #12334
  • Honour FP8Linear.block_size in the patched FP8 forward (32x32 block checkpoints) by @danielhanchen in #12317
  • Write the Ollama Modelfile from the trained chat template by @NilayYadav in #12311
  • Fix MiniMax-H3 compiled render on torch 2.12, 2.13 and 2.14 by @danielhanchen in #12326
  • Studio: load LTX-2 and LTX-2.3 on the pinned transformers 5.5 by @danielhanchen in #12299
  • scan_packages: refuse VCS, URL and local-path specs before pip download by @danielhanchen in #12357
  • Studio: let GGUF vision models see dark text in transparent images by @NilayYadav in #12310
  • Keep masked head_dim 256 SDPA training off cuDNN attention on SM100 (torch 2.14 NaN grads) by @danielhanchen in #12344
  • Unsloth Studio (AMD): run native image generation on an NVIDIA card next to ROCm torch by @LeoBorcherding in #12252
  • fix(lora): preserve all active adapters with PEFT fallbacks by @taking-lying-flat in #12214
  • Studio: load downloaded tokenizers offline, size unsloth mirrors from the family table by @danielhanchen in #12300
  • Build the marlin_gemm call from the op schema so packed INT4 inference works on vLLM 0.29 by @danielhanchen in #12320
  • Studio video: keep an explicit fast LTX-2.3 single-file load resident when it fits by @danielhanchen in #12345
  • Studio: keep MiniMax-H3's int8 denoiser and compile when it has to offload by @danielhanchen in #12286
  • Model-config UI test: stop waiting on a request whose finish event never arrives by @danielhanchen in #12363
  • Studio: keep INT8 and compile on conventional video models when they have to offload by @danielhanchen in #12289
  • Studio frontend test: drain the previous runtime's writes before each reasoning-effort scenario by @danielhanchen in #12367
  • Baseline the moved transformers 5.18.0 testing_utils polling-loop site after review by @danielhanchen in #12368
  • Studio: keep int8 / fp8 image transformers quantised when the load has to offload by @danielhanchen in #12287
  • Studio: style canvas notices like toasts by @shimmyshimmer in #12358
  • Studio: keep the model name ahead of its format and quant when the header is tight by @shimmyshimmer in #12396
  • Studio: compact context usage ring when the chat header is squeezed by @shimmyshimmer in #12355
  • fix(device_type): flush and fence through the shared device helpers by @li-lizhe in #12284
  • Studio: open a skill when its row chevron is clicked by @NilayYadav in #12375
  • Studio: keep the answers typed into a PDF form when it is attached to a chat by @NilayYadav in #12377
  • Studio: stop offering Claude sampling settings that are silently ignored by @NilayYadav in #12383
  • Studio: count companion assets in the Hub download size of image and video GGUFs by @oobabooga in #12370
  • Feat share model run settings through links by @Sneakr in #11710
  • studio: make media family overrides structural by @Imagineer99 in #10150
  • Studio: keep and Vec text in chat replies by @NilayYadav in #12376
  • Harden installer source selection, ROCm helper staging and npm scanner cleanup by @danielhanchen in #12402
  • Remove xFormers built for another torch after Linux repair by @tweekli in #11783
  • truststore: keep TLS verification on when handshakes overlap on one context by @danielhanchen in #12403
  • Read a sign-flipped SVD basis as the same subspace in the Q-GaLore schedule by @danielhanchen in #12399
  • Create the SentencePiece scratch directory without a check-then-create race by @danielhanchen in #12400
  • Only turn on HF_HUB_ENABLE_HF_TRANSFER in synthetic.py when hf_transfer is installed by @danielhanchen in #12406
  • Studio: train an uploaded CSV's NA, None and 00501 cells as written by @NilayYadav in #12379
  • Studio: fix Base vs LoRA compare for voice messages on fine-tuned Whisper by @NilayYadav in #12381
  • Compare Kaggle reference repo ids without case and prefetch the Qwen3 4bit repo under its new spelling by @danielhanchen in #12411
  • Seed preview: check path-backed image cells against the account's workspace by @danielhanchen in #12404
  • Studio: say Pin in model menus, and drop the border on right-click submenus by @shimmyshimmer in #12414
  • Stop the Chat UI persisted-monitor reset from running script in the stale page by @danielhanchen in #12412
  • Studio: keep a local model's built-in system prompt when the date setting is on by @NilayYadav in #12382
  • Studio: let the Decision API use TypeSafe, Liquid AI, OpenRouter and other System One servers by @NilayYadav in #12373
  • Keep the MLX adamw_8bit optimizer instead of collapsing it to adamw by @vineethsaivs in #9950
  • Treat {{ and }} in a merged prompt as literal braces by @vineethsaivs in #10742
  • fix(studio): keep scoped download progress tied to current files by @Imagineer99 in #12390
  • Studio: continue finished replies and resume GGUF reasoning by @oobabooga in #12371
  • Studio: keep Qwen-Image-2.1 image conditioning finite on ROCm by @oobabooga in #12360
  • Restore full-rank gradients after Q-GaLore updates by @vineethsaivs in #10878
  • Studio: fix PDF previews after a PDF attachment is extracted by @shimmyshimmer in #12346
  • Studio: stop the backend crashing on startup when memory is tight by @NilayYadav in #12374
  • Studio: warn that a Web share link from a remote Studio exposes its address by @danielhanchen in #12413
  • Studio: smaller scroll to bottom button, visible in dark mode, with a setting to hide it by @shimmyshimmer in #12417
  • Stop conversation_extension crashing when the caller keeps their columns by @vineethsaivs in #8373
  • Fix desktop icon clarity with supplied artwork and Tauri resource icons by @wasimysaid in #12352
  • Studio: always show composer attachments as cards, rename sent layouts by @shimmyshimmer in #12418
  • Studio: remove the Projects section setting from Chat settings by @shimmyshimmer in #12421
  • Studio: keep the arrow cursor on a sent prompt's time by @shimmyshimmer in #12420
  • Keep an explicit HF_HUB_ENABLE_HF_TRANSFER and install hf_transfer in Core CI by @danielhanchen in #12394
  • Studio: remove the white seam above the chat panel on Windows dark mode by @danielhanchen in #12425
  • Studio: avoid slow cold MIOpen searches on gfx1151 by @oobabooga in #12359
  • Studio: faster MiniMax-H3 GGUF renders (resident under memory auto, sd.cpp pin upgrade, speed_mode=max kernels) by @danielhanchen in #12405
  • Studio: keep Word footnotes and endnotes in chats and knowledge bases by @NilayYadav in #12378
  • Studio: make Compare in Chat load the full fine-tune that just finished by @NilayYadav in #12380
  • Stop the model selector's format and quant suffix clipping descenders by @danielhanchen in #12427
  • fix: allow configuring Pi output token limit by @Imagineer99 in #12393
  • Studio: keep the sidebar's bottom fade in step with the list by @shimmyshimmer in #12423
  • Studio: keep the 1024 canvas when auto precision picked the Qwen-Image-2.1 quant by @danielhanchen in #12428
  • Studio: move Blender MCP setup out of Manage MCP servers into the composer by @NilayYadav in #12430
  • Studio: explain the all-columns-dropped recipe error in UI terms by @danielhanchen in #12416
  • Studio: Qwen-Image-2.1 placement from measured sizes, int8 under offload on torchao 0.17, balanced fit check by @danielhanchen in #12408
  • Read the desktop New chat button contract token by token by @danielhanchen in #12426
  • feat(studio): configurable RAG upload extensions via RAG_UPLOAD_EXTS by @Souravrajvi0 in #11499
  • Unsloth Studio: return the spoken text from /audio/generate instead of a truncated label by @LeoBorcherding in #12386
  • Studio: keep PyTorch mirror leaves out of query tokens by @alkinun in #10546
  • perf(studio): reuse embeddings for identical files in linked folders by @Mathews-Tom in #11972
  • Studio: render MCP Apps widgets in the chat thread by @NilayYadav in #9301
  • Bump install.sh / install.ps1 pins to unsloth>=2026.9.13, unsloth-zoo>=2026.9.9 by @danielhanchen in #12432
  • fix(studio): use resolved public id in embeddings/completions monitor by @Souravrajvi0 in #9346
  • Studio: keep rounded boxes rounded when they scroll by @shimmyshimmer in #12431
  • Studio: browse temporary Linux mounts under /media and /mnt by @Souravrajvi0 in #10048
  • Make the registry's default quant_type usable by @vineethsaivs in #9887
  • Size the packed attention mask at the padded length, not the token count by @vineethsaivs in #8278
  • Keep the GRPO eval batch a multiple of num_generations by @vineethsaivs in #12362
  • Prefetch Tauri NSIS and WebView2 tools before the Windows desktop build by @danielhanchen in #12434
  • fix(studio): retry custom gateways with max_completion_tokens after max_tokens 400 (#10787) by @Souravrajvi0 in #10837
  • Collapse the llama.cpp install branch that never branched by @vineethsaivs in #10268
  • Resolve a sentence-transformers modules.json module class through the same trust gate as upstream by @danielhanchen in #12436

New Contributors

Full Changelog: v0.1.900-beta...v0.1.902-beta

Don't miss a new unsloth release

NewReleases is sending notifications on new releases.