0.21.0
Added
- Qwen-Image-2.1 now supports instruction-based editing with up to ten reference images, RGBA output, and prefix KV caching through mflux-generate-qwen-2.1-edit. The new QwenImage21Edit Python API can quantize and export complete vision-capable checkpoints; the existing generation command and text-only exports remain available. (#741)
- Logging is now configured at CLI boot; use
--verboseto make it more verbose. (#747) - Qwen-Image-2.1 edit: new
--mask-image,--auto-mask,--strength,--enhance-prompt,--verifyand--use-step-cacheoptions formflux-generate-qwen-2.1-edit. They give masked edits, automatic object masks, subtler edits, prompt rewrite, a self-check with retries, and faster runs. (#749) - Added Qwen Image 2.1 LoRA support for PEFT
.defaultadapters via--lora. (#756) - Qwen-Image-2.1: new
--scheduler viggle_turboformflux-generate-qwen-2.1andmflux-generate-qwen-2.1-edit. With Viggle's distilled turbo LoRA and--steps 6, it generates and edits in 6 steps instead of 40. (#764) - New model: inclusionAI Ming-Image-0.1-Design (
mflux-generate-ming), a design-focused text-to-image model with RGBA output; its 16B MoE text encoder can be quantized separately with--text-encoder-quantize. (#765) - Qwen Image 2.1 now accepts common transformer. and diffusion_model. LoRA exports, including fused gate_up weights and modulation/timestep embedding adapters. (#768)
- Z-Image now loads ComfyUI LoRAs with fused QKV projections and normalization/bias deltas. Adapters containing direct deltas require the default baked inference mode. (#772)
- Support independently installed mflux namespace extensions, enabling community mflux.web UI packages to coexist across editable checkouts and separate import directories. (#776)
- Qwen-Image-2.1 reference editing now supports LoRA adapters through the same mappings as text-to-image, including baking controls and metadata replay. Both commands share the Transformer core and component-loading mechanism while retaining compatibility with existing model exports. (#777)
mflux-generate-z-image-turbonow exposesZImageTurboCommand.loadandZImageTurboCommand.generate, so Python code can load the model once and generate many images with the command's own flag handling. (#780)mflux-generate-z-imageandmflux-generate-z-image-controlnetnow exposeloadandgeneratefor Python callers, and all three Z-Image commands addvalidate, which checks a request without loading weights. (#785)mflux-savecan now save the SeedVR2 3B and 7B upscalers.mflux-upscale-seedvr2 --model <dir>reads the variant from the saved weights. (#800)mflux-generate-ernie-image-turboandmflux-generate-ernie-imagenow exposevalidate,loadandgeneratefor Python callers. (#815)
Improved
- Qwen-Image-2.1 generates up to ~1.8x faster: denoise steps run ~12% faster (cached text prefix, bounded to the two most recent embeddings; fused Q/K norm+rope Metal kernel, disable with MFLUX_QWEN21_DISABLE_FUSED_PROLOGUE=1), and
generate_image(..., teacache_ratio=0.25)reuses the previous noise on the least-changing steps for a further 1.4x (1.77x at 0.4). (#778) - Qwen Image 2.1's step reuse is now
--step-cache-ratio(--teacache-ratiostill works) and the ratio is recorded in image metadata and replayed by--config-from-conf. The logic is shared (mflux.models.common.step_cache.StepCache) so other models can adopt it. (#779) - MFlux now imports OpenCV and PyTorch only in the code paths that use them. Commands that do not use these libraries start without them. (#787)
- LoRA and LoKr adapters now run in the precision the model was loaded in, not float32. Training and LoRA inference on bfloat16 models use less memory and run faster. (#801)
Fixed
- bump anyio due to CVE-2026-63374 & CVE-2026-64847. MFlux doesn't import it directly - it is used transitively by httpx through huggingface-hub (#744)
- mflux now refuses a checkpoint whose weight names do not match the model, with an error that names the part and the missing weights. Before, such a checkpoint loaded with random weights and produced noise. This mostly affects checkpoints converted for other programs. (#758)
- DoRA adapters saved with
lora_A/lora_Bmatrices (the PEFT layout) now stop with an error on every model, instead of loading without their magnitude. LoKr adapters withdora_scalestill load. (#774) - Krea 2 checkpoints saved by earlier releases (mflux shards under
transformer/) load again instead of coming up untrained. (#792) - Reference images with transparency are composited over white before the model sees them; masks keep their transparent background as "preserve". (#793)
- Qwen-Image-2.1 edit: a multi-seed run now finds the auto-mask and rewrites the prompt one time, not one time for each seed. Edit verification now reads plain-text replies with any spacing. (#796)
--low-ramnow frees the transformer before VAE decode in Z-Image, Krea-2, FLUX.2 Klein, Ming-Image, ERNIE-Image and Ideogram-4. Before, the transformer weights stayed in memory during decode, which increased the peak memory. (#802)- Image metadata now keeps LoRA scales, the ControlNet strength and the Redux strengths at full precision. Before, these values were rounded to 2 decimals, so
--config-from-confdid not replay the same run. (#805)
Changed
- Three CLI flags have new names:
--make-conf(was--metadata),--no-exif(was--no-metadata) and--config-from-conf(was--config-from-metadata). The old names continue to work as aliases. (#797)
Internal
- Add Qwen-Image-2.1 to the CI Manifest script (
scripts/ci_extract_models.py) .. and improve the Agent directions & Skills to remind AI Agents to do this task. (#742)