github jundot/omlx v0.7.0
0.7.0

2 hours ago

oMLX 0.7.0

After a long wait, oMLX 0.7.0 is out. oMLX 0.7.0 brings faster Qwen3.8, GLM-5.3-Flash and MiMo V2 inference, a completely rebuilt memory guard, and fixes for the regressions reported against the release candidate. The notes below cover changes since 0.7.0rc1.

Download: macOS 26 / 27 | macOS 15 Sequoia

Screenshot 2026-10-01 at 00 21 47

Faster Qwen3.8-Flash-Next

  • Prefill. Wider prefill steps, exact hyper-connection fusions, QSA attention on the M5 tensor units and wider native QSA tiles. By @jonathan308 in #3980, #3981, #3982, #4020; @jerryfane in #3934.
  • MoE prefill. Fused gate/up gathers with a SwiGLU epilogue and segmented NAX tile scheduling, with identical greedy output. Oversized sorted MoE calls on M5 run in one dispatch instead of 32K-row slices. By @jonathan308 in #3995, #4022, #4029; @jerryfane in #3954.
  • Decode. Fused MoE, DeltaNet and attention kernels for one-token decode and MTP verify windows, and routed experts run in two launches instead of five. By @jerryfane in #4024, #4038, #4039, #4041; @wyanzhao in #3912.
  • Exact Lightning MTP. For a single request, verify rows are now bit-identical to serial decode, so turning Lightning MTP on does not change greedy output. By @jerryfane in #4023.
  • Qwen3.5/3.6 GDN prefill. A software-pipelined GDN recurrence speeds up Qwen3.5/3.6 VLM prefill. By @jonathan308 in #4006.

Faster Qwen3.8 Lightning MTP and DFlash2

DFlash2 and Lightning MTP now do less work per verify cycle on Qwen3.8-27B and Qwen3.8-Flash-Next, with the largest gains at long contexts. The fp16 checkpoints used on GPUs without native bf16 now use the same verify kernels. Flash-Next no longer falls back to plain decode at 64K on M3 Ultra, and after a performance park MTP resumes with its full head history instead of an empty one. #3958.

Faster GLM-5.3-Flash

  • Prefill. Pipelined per-layer evaluation, fused hyper-connection kernels, a blocked KDA recurrence, tensor-unit DSA indexer and sparse MLA attention, and 4096-token chunks on M5 GPUs. By @jonathan308 in #3971, #3983, #3984, #3985, #3986, #3988, #3996, #4025. This builds on the earlier Qwen4 prefill port by @williamxie1989 in #3944.
  • Decode. Exact fused decode and MTP verify kernels on M5 GPUs, plus cheaper decode steps on every Mac, with unchanged output. By @jonathan308 in #3989, #4019, #4026.
  • DFlash2. GLM-5.3-Flash now works as a DFlash2 target with the incoai/GLM-5.3-Flash-DFlash2 draft. By @jonathan308 in #3979.

Faster MiMo V2

Blocked sliding-window attention, JIT NAX attention kernels for 192/128 head dims, long-context attention in key-range passes, a fused decode and verify layer path, and wider prefill chunks (4096 tokens, 8192 on M5 GPUs). MiMo experts also combine through the fused weighted-sum kernel, and MiMo and GLM-5.3 run expert gate/up as one gather. By @jonathan308 in #3972, #3973, #3990, #3991, #3992, #3994, #4032, #4033, #3977, #3993.

A Completely Rebuilt Memory Guard

The previous memory guard could refuse a prompt even when there was still memory to spare, or try one when there wasn't enough. The memory guard in 0.7.0 was rebuilt from scratch. It helps oMLX use as much of your memory as it safely can, and with the aggressive tier that means close to all of it. Tiers now set how much memory stays free for other apps: safe keeps about 20% of RAM (6-16 GB), balanced about 8% (3-8 GB) and aggressive about 2% (1.5-4 GB). If you turned the memory guard off before, try turning it back on. #3933.

New Features and Improvements

  • Faster image follow-ups on Qwen3.8-Flash-Next. Vision features are cached per image and only new images are encoded, within a 1 GiB memory budget. In a four-turn screenshot conversation, a text-only follow-up started in 0.33s instead of 3.32s. By @williamxie1989 in #3955.
  • Lightning MTP with expert offload for Qwen3.8-Flash-Next. The MTP head stays resident while backbone experts stream from the checkpoint. By @samad-aghaei in #3935.
  • Live concurrency changes. max_concurrent_requests applies without unloading models. By @zviratko in #3765.
  • Download queues. Each row shows a live transfer rate, queues survive a restart, progress follows xet's fetch phase and cancel stops the transfer right away. By @LXD-8 in #3875. Also addresses #505.
  • DeepSeek Harness. Launch DeepSeek Harness (dsh) with oMLX as its provider. By @LXD-8 in #3950.
  • Less GPU wake-up delay. oMLX keeps the GPU out of its idle power state for up to 5 minutes after the last request, which removed about 1 to 1.5 seconds from the next time to first token in MiMo tests. The interval is server.gpu_keep_warm_interval in settings.json. By @jonathan308 in #3974.
  • Less per-request tokenizer work. The streaming detokenizer is built once per tokenizer instead of on every request, which saves about 45 ms per request on Qwen3.8. By @jonathan308 in #3975.
  • Faster expert offload on slow storage. Offloaded expert reads overlap with resident expert compute during decode. By @srcterm in #4037.
  • Exact prefix reuse for split-GDN SpecPrefill. Stable system and tool prefixes are no longer prefilled on every request. By @jerryfane in #3908.
  • Embedding precision. bf16 Qwen3-Embedding-0.6B and -8B checkpoints are promoted to fp16, which keeps their output within the conformance tolerance. By @ernestas-poskus in #3640.
  • Accuracy runs. Local runs record answers cut off by the token limit. By @beatakouchnir in #3945.
  • Cluster SSH accounts. Set a persistent SSH account for each paired Mac in the wizard, and pairing shares SSH account names. By @lybertybox443 in #4075; 0905fad.

Bug Fixes

  • ANE prefill. A TypeError in 0.7.0rc1 disabled ANE prefill and left no GDN procedures eligible. Addresses #3957.
  • oQ A8 prefill on M5. Dense Qwen 4-bit projections silently fell back to W4A16 in 0.7.0rc1. Addresses #3952.
  • Clustering. Helper Python processes no longer inherit the app's PYTHONHOME, rank failures are reported instead of masked, and ranks follow the mlx-lm 0.32 server and pipeline APIs. Addresses #3937, #3965, #3966 and #3967.
  • Muse Glimmer with DFlash. The dflash-mlx pin now works with mlx-lm 0.32 for Muse Glimmer, and Muse Glimmer DFlash can reuse prefix snapshots. Addresses #4009.
  • Accuracy benchmarks. Client disconnects no longer stall a run, saving settings during a run no longer cancels it, forced LM loads skip model families that only mlx-vlm implements, and Metal out-of-memory load failures are no longer cached as broken files. Addresses #3956, #3960 and #3961, and part of #3959.
  • Lightning MTP for MLX-converted MiMo checkpoints. The MTP sidecar loads with the correct tensor-parallel split, raising acceptance to 83-89% and decode from 43-45 to 53-57 tok/s on M3 Ultra. By @jonathan308 in #3976.
  • Lightning MTP batching. A late join hands off instead of replaying a row's whole history, and greedy verify picks tokens like the serial sampler. By @yoyo930021 in #4031; @jerryfane in #4050.
  • Decode attention. Declared threadgroup limits fix incorrect native decode attention on M1 Max, and the decode fallback keeps causal masking. By @samfenwick in #4063, #4064.
  • Qwen3.5/3.6 VLM numerics. GDN q/k normalization now matches mlx-lm and the reference model. SSD prefix blocks computed with the old numerics are rebuilt once. 5f70e46.
  • SSD cache growth on hybrid models. GDN, QSA, GLM-5 linear and DeepSeek V4.1 layouts kept one full-state cache tail per turn. Tails from two turns back are now deleted. a28e5a8.
  • Sampling. Small top_p values no longer mask every token in bf16 and fp16, and XTC no longer takes its cutoff across batch rows. 4255940.
  • Qwen3.8-Flash-Next loading. The n-gram table shards are released during load, so larger quants like oQ8e load reliably. By @agg23 in #4087.
  • DeepSeek V4.1 images. A literal image placeholder in text no longer causes repeated 400 errors, and /v1/responses drops input images for text-only models. By @Mast0089 in #4011. Addresses #4010 and #4069.
  • Gemma 4 audio. Audio weights with Hugging Face names are recognized, so gemma-4-E2B-it transcribes audio instead of dropping it. Audio models key the prefix cache on the audio content, so two clips with the same prompt no longer share a cached transcript, and oQ conversion loads 0-dim tensors, so gemma-4-E2B-it can be quantized. By @vichzp in #4117, #4118, #4119. Addresses #4114, #4115 and #4116.
  • Gemma 4 multi-audio and chat template. Several audio clips in one request no longer fail with a 500, and later clips keep their own placeholders. Prompts render with the tokenizer chat template, so each audio marker stays in its own turn and list system content is no longer printed as a Python list. 077f6b4, 8cc02d6.
  • Qwen fused GDN norm. The fused GDN norm uses the same SiLU arithmetic as the served path, which removes small mismatches in fp16 and on large negative gates. By @samfenwick in #4122.
  • Tool calling. Hy3 reads its tool parser template from the tokenizer, and repositories that label non-JSON tool grammars as json_tools use the parser their template needs. By @lybertybox443 in #4076; 0724417.
  • Prefill admission after eviction. A request no longer gets an HTTP 400 right after the pool frees enough memory for it, because the scheduler refreshes its memory sample before the final admission check. By @samfenwick in #4124. Addresses #4059.
  • Structured output. JSON schema decoding bounds whitespace, so a string closed early no longer spins until max_tokens. By @saichowdary007 in #3321. Addresses #2990.
  • GLM-5.3 indexer parity. Fused decode keeps the same top-k order as the general indexer when native top-k is unavailable. By @lybertybox443 in #4098.
  • Cluster reliability. SSD prefixes are restored after a memory cache rejection, admitted prefill respects SSD snapshot boundaries, staging uses the published worker shim, SFTP transfers keep literal paths, SSH probing stops after a transport failure, and pairing can retry after an initial connection failure. By @lybertybox443 in #4078, #4079, #4092, #4093, #4094, #4095, #4096.
  • Unexpected model reloads. TurboQuant bit depth is stored the same way in settings and profiles, so model aliases no longer trigger reloads. By @zviratko in #3616.
  • Qwen3-VL embeddings. Position ids are recomputed per request. By @zviratko in #3874. Addresses #3731.
  • Gemma 4 oQ conversions. The MoE router stays at source precision. By @andyoneal in #4107.
  • Usage statistics. Non-streaming responses report token rates, and rates for uncached prefill are correct. By @junmo-kim in #4065.
  • Throughput bench. Runs that stop within the first 16 tokens no longer report a decode rate made of a few burst tokens. ffeeb73.
  • Log noise. Forced-offload decisions are logged on each model load instead of on every model list refresh. By @cropduster in #4100. Addresses #4099.
  • macOS app. IPv6 hosts are bracketed in API endpoint URLs, and the Claude Code context-scaling controls the server no longer supports are removed. By @agrover in #4043; @LXD-8 in #3914.

Dependency Updates

  • mlx-lm 94cdcae (main after 0.32.0) and dflash-mlx 0.1.10+omlx.9.

New Contributors

@jerryfane in #3908, @wyanzhao in #3912, @samad-aghaei in #3935, @Mast0089 in #4011, @yoyo930021 in #4031, @srcterm in #4037, @agrover in #4043, @junmo-kim in #4065, @lybertybox443 in #4075, @agg23 in #4087, @vichzp in #4117.

Release notes from the release candidate and the development releases cover the rest of the 0.7.0 cycle.

Changes since 0.7.0rc1 | Full changelog from 0.6.4

Don't miss a new omlx release

NewReleases is sending notifications on new releases.