github magnitudedev/magnitude @magnitudedev/cli@0.2.7

3 hours ago

0.2.7

Patch Changes

  • 3fe2c30 Thanks @thrgreenwald! - - Magnitude is available for Arch Linux, Omarchy, and other Arch-based distributions (x64) as a pacman package. Install it with sudo pacman -U ./magnitude-desktop.pkg.tar.zst or the install script, and in-app updates install through pacman. As with the Debian and RPM packages, pacman refuses to upgrade or remove Magnitude while it is running.

  • c74b8b0 Thanks @thrgreenwald! - - Fix OpenClaw requests hanging: its automations tool kept the server busy without answering and could push the loaded model out of memory. These requests now answer in seconds.

    • Fix Claude Code and Oh My Pi requests failing with "Too many items" on MiniCPM5 and LFM2.5, and Oh My Pi failing on Qwen3.5 with reasoning off.
    • A response cut off by its output limit in the middle of a tool call now ends as an ordinary length stop (length, max_tokens, or incomplete) instead of failing with an error.
    • Fix connecting Codex from the Magnitude app on Windows, which always failed.
  • b2f79f0 Thanks @anerli! - - Faster local inference across the board: decoding is up to about 3× faster with several requests at once, and prompt processing up to 4.5× faster (Gemma 4 E2B on Apple Silicon), on Apple Silicon and NVIDIA GPUs.

    • The one-time optimization after a download now runs in its own process within a fixed time budget, and loading a model never tunes again afterwards. A model also starts optimizing much sooner on first load (Qwen 3.8 27B: 1.9 s instead of 12.6 s).
    • Gemma 4 models use their assistant model as a draft head for faster generation.
    • Qwen3.5 4B and 9B now use their DFlash drafters for speculative decoding, so they generate considerably faster.
    • Faster prefill and decode on Apple Silicon Macs (M1 through M4): weights use a new tiled layout, prompt processing packs tokens, and attention runs on the matrix and scalar units together. The one-time optimization now plans its time across every kernel instead of searching them in order.
    • Faster time to the first token for models with a built-in drafting head: each prompt chunk is no longer computed twice.
    • Follow-up turns in a conversation reuse the previous turn's prompt even when the template renders past turns differently, so later turns start sooner.
  • 44e9fd5 Thanks @thrgreenwald! - - Fix a request without a seed repeating the same sample on every retry, which made some requests (such as Gemma 4 E2B over the Responses API with reasoning off) return an empty answer every time. Each request now samples with a fresh seed unless it names one, and the Responses and Anthropic APIs accept seed like Chat Completions.

  • 7d62661 Thanks @thrgreenwald! - - Fix long agent sessions on Macs running the GPU out of memory, which unloaded the model mid-session. Memory used for earlier turns is now released before the GPU fills.

    • Fix NVIDIA GPUs crashing with an illegal memory access when several requests run together after a long prompt.
    • Fix large images and long prompts resetting AMD GPUs on Linux and unloading the model: GPU work is now split so no single piece runs past the driver's time limit.
    • On computers without a supported GPU, replies now stream as they are written instead of arriving all at once, stopping or abandoning a request stops it promptly instead of holding up the next one, and generation is faster on Intel and other x86 processors.
    • Fix the one-time optimization never running on Windows, which left every model on slower untuned kernels.
    • Fix Qwen3.5 4B and 9B never finishing their one-time optimization on Macs, which re-ran it on every load.
    • Fix the optimization failing on GPUs with 8 GB of memory, and on computers without a supported GPU when optimization runs again.
    • Restarting Magnitude or the computer no longer repeats the optimization on the next load.
    • The first use of a model on AMD GPUs on Linux is much faster: the optimization now prepares every kernel instead of leaving them to the first load.
    • Fix image requests failing with "insufficient memory" after a long session on AMD GPUs on Linux.
    • Fix Magnitude closing at launch when it is opened while a previous instance is still quitting, or while it takes over from magnitude serve.
    • magnitude models load no longer reports a transport error when a model takes longer than five minutes to load.
    • When the inference engine stops unexpectedly, its exit code and log now appear in Magnitude's logs.
  • a54897b Thanks @thrgreenwald! - - Fix large images failing in vision models. An image that needed more rows than the engine's per-launch limit failed; images now encode up to 4,096 cells, and larger images are resized to fit.

  • d066f74 Thanks @thrgreenwald! - - Fix agent checkpoints and undo failing on Linux systems whose default file permissions are group-writable, such as Ubuntu.

    • Purging the Debian package removes Magnitude's installation state, and the login item no longer points at a removed application.
    • If an unfinished package change keeps Magnitude from starting on Linux, opening it from the application menu now shows a notice with repair steps instead of doing nothing.
    • pacman -Qkk reports no permission mismatch for the Arch package.
    • nohup magnitude serve & keeps running after you log out of an SSH session.
  • 7f65eb3 Thanks @anerli! - - macOS loads a model that fits beside wired and compressed memory instead of refusing it, including right after a download.

    • Qwen 3.8 27B loads on NVIDIA GPUs, and Gemma 4 E2B and E4B no longer hit an illegal memory access on NVIDIA GPUs.
    • The one-time optimization is more reliable: it no longer settles on slow defaults for large contexts (Gemma 4 12B at a 65,536-token context decoded at 4.8 tok/s), tunes sliding-window layers at the history they keep, and finishes in bounded time on first load.
    • Muse Glimmer can turn reasoning off.
  • 9452ff4 Thanks @thrgreenwald! - - A download that stops receiving data now reports that it can't reach the model source and offers Retry, which resumes from where it stopped, instead of freezing.

    • Clicking Download right after cancelling a download now starts it again instead of doing nothing.
    • An expired or revoked HF_TOKEN no longer blocks downloads of public models.
    • Removing a model on Windows now deletes its files and frees the disk space.
    • Fix retrying a model download with a draft model after a network outage, which could fail again.
  • 4dba3cd Thanks @thrgreenwald! - - Fix a stray horizontal scrollbar on Windows and Linux in windows narrower than about 1,380 pixels. Page content now fits beside the vertical scrollbar instead of extending under it.

  • 541ccf1 Thanks @thrgreenwald! - - Magnitude writes logs again: service.log, inference.log and desktop.log in ~/.magnitude/logs record the service, the inference engine and the app, including crashes and restarts. Each file is size-capped. Attach them when reporting a problem.

Don't miss a new magnitude release

NewReleases is sending notifications on new releases.