0.2.7
Patch Changes
-
3fe2c30Thanks @thrgreenwald! - - Magnitude is available for Arch Linux, Omarchy, and other Arch-based distributions (x64) as a pacman package. Install it withsudo pacman -U ./magnitude-desktop.pkg.tar.zstor the install script, and in-app updates install through pacman. As with the Debian and RPM packages, pacman refuses to upgrade or remove Magnitude while it is running. -
c74b8b0Thanks @thrgreenwald! - - Fix OpenClaw requests hanging: itsautomationstool kept the server busy without answering and could push the loaded model out of memory. These requests now answer in seconds.- Fix Claude Code and Oh My Pi requests failing with "Too many items" on MiniCPM5 and LFM2.5, and Oh My Pi failing on Qwen3.5 with reasoning off.
- A response cut off by its output limit in the middle of a tool call now ends as an ordinary length stop (
length,max_tokens, orincomplete) instead of failing with an error. - Fix connecting Codex from the Magnitude app on Windows, which always failed.
-
b2f79f0Thanks @anerli! - - Faster local inference across the board: decoding is up to about 3× faster with several requests at once, and prompt processing up to 4.5× faster (Gemma 4 E2B on Apple Silicon), on Apple Silicon and NVIDIA GPUs.- The one-time optimization after a download now runs in its own process within a fixed time budget, and loading a model never tunes again afterwards. A model also starts optimizing much sooner on first load (Qwen 3.8 27B: 1.9 s instead of 12.6 s).
- Gemma 4 models use their assistant model as a draft head for faster generation.
- Qwen3.5 4B and 9B now use their DFlash drafters for speculative decoding, so they generate considerably faster.
- Faster prefill and decode on Apple Silicon Macs (M1 through M4): weights use a new tiled layout, prompt processing packs tokens, and attention runs on the matrix and scalar units together. The one-time optimization now plans its time across every kernel instead of searching them in order.
- Faster time to the first token for models with a built-in drafting head: each prompt chunk is no longer computed twice.
- Follow-up turns in a conversation reuse the previous turn's prompt even when the template renders past turns differently, so later turns start sooner.
-
44e9fd5Thanks @thrgreenwald! - - Fix a request without aseedrepeating the same sample on every retry, which made some requests (such as Gemma 4 E2B over the Responses API with reasoning off) return an empty answer every time. Each request now samples with a fresh seed unless it names one, and the Responses and Anthropic APIs acceptseedlike Chat Completions. -
7d62661Thanks @thrgreenwald! - - Fix long agent sessions on Macs running the GPU out of memory, which unloaded the model mid-session. Memory used for earlier turns is now released before the GPU fills.- Fix NVIDIA GPUs crashing with an illegal memory access when several requests run together after a long prompt.
- Fix large images and long prompts resetting AMD GPUs on Linux and unloading the model: GPU work is now split so no single piece runs past the driver's time limit.
- On computers without a supported GPU, replies now stream as they are written instead of arriving all at once, stopping or abandoning a request stops it promptly instead of holding up the next one, and generation is faster on Intel and other x86 processors.
- Fix the one-time optimization never running on Windows, which left every model on slower untuned kernels.
- Fix Qwen3.5 4B and 9B never finishing their one-time optimization on Macs, which re-ran it on every load.
- Fix the optimization failing on GPUs with 8 GB of memory, and on computers without a supported GPU when optimization runs again.
- Restarting Magnitude or the computer no longer repeats the optimization on the next load.
- The first use of a model on AMD GPUs on Linux is much faster: the optimization now prepares every kernel instead of leaving them to the first load.
- Fix image requests failing with "insufficient memory" after a long session on AMD GPUs on Linux.
- Fix Magnitude closing at launch when it is opened while a previous instance is still quitting, or while it takes over from
magnitude serve. magnitude models loadno longer reports a transport error when a model takes longer than five minutes to load.- When the inference engine stops unexpectedly, its exit code and log now appear in Magnitude's logs.
-
a54897bThanks @thrgreenwald! - - Fix large images failing in vision models. An image that needed more rows than the engine's per-launch limit failed; images now encode up to 4,096 cells, and larger images are resized to fit. -
d066f74Thanks @thrgreenwald! - - Fix agent checkpoints and undo failing on Linux systems whose default file permissions are group-writable, such as Ubuntu.- Purging the Debian package removes Magnitude's installation state, and the login item no longer points at a removed application.
- If an unfinished package change keeps Magnitude from starting on Linux, opening it from the application menu now shows a notice with repair steps instead of doing nothing.
pacman -Qkkreports no permission mismatch for the Arch package.nohup magnitude serve &keeps running after you log out of an SSH session.
-
7f65eb3Thanks @anerli! - - macOS loads a model that fits beside wired and compressed memory instead of refusing it, including right after a download.- Qwen 3.8 27B loads on NVIDIA GPUs, and Gemma 4 E2B and E4B no longer hit an illegal memory access on NVIDIA GPUs.
- The one-time optimization is more reliable: it no longer settles on slow defaults for large contexts (Gemma 4 12B at a 65,536-token context decoded at 4.8 tok/s), tunes sliding-window layers at the history they keep, and finishes in bounded time on first load.
- Muse Glimmer can turn reasoning off.
-
9452ff4Thanks @thrgreenwald! - - A download that stops receiving data now reports that it can't reach the model source and offers Retry, which resumes from where it stopped, instead of freezing.- Clicking Download right after cancelling a download now starts it again instead of doing nothing.
- An expired or revoked
HF_TOKENno longer blocks downloads of public models. - Removing a model on Windows now deletes its files and frees the disk space.
- Fix retrying a model download with a draft model after a network outage, which could fail again.
-
4dba3cdThanks @thrgreenwald! - - Fix a stray horizontal scrollbar on Windows and Linux in windows narrower than about 1,380 pixels. Page content now fits beside the vertical scrollbar instead of extending under it. -
541ccf1Thanks @thrgreenwald! - - Magnitude writes logs again:service.log,inference.loganddesktop.login~/.magnitude/logsrecord the service, the inference engine and the app, including crashes and restarts. Each file is size-capped. Attach them when reporting a problem.