github mflux-community/mflux v.0.21.0
Release 0.21.0

4 hours ago

0.21.0

Added

  • Qwen-Image-2.1 now supports instruction-based editing with up to ten reference images, RGBA output, and prefix KV caching through mflux-generate-qwen-2.1-edit. The new QwenImage21Edit Python API can quantize and export complete vision-capable checkpoints; the existing generation command and text-only exports remain available. (#741)
  • Logging is now configured at CLI boot; use --verbose to make it more verbose. (#747)
  • Qwen-Image-2.1 edit: new --mask-image, --auto-mask, --strength, --enhance-prompt, --verify and --use-step-cache options for mflux-generate-qwen-2.1-edit. They give masked edits, automatic object masks, subtler edits, prompt rewrite, a self-check with retries, and faster runs. (#749)
  • Added Qwen Image 2.1 LoRA support for PEFT .default adapters via --lora. (#756)
  • Qwen-Image-2.1: new --scheduler viggle_turbo for mflux-generate-qwen-2.1 and mflux-generate-qwen-2.1-edit. With Viggle's distilled turbo LoRA and --steps 6, it generates and edits in 6 steps instead of 40. (#764)
  • New model: inclusionAI Ming-Image-0.1-Design (mflux-generate-ming), a design-focused text-to-image model with RGBA output; its 16B MoE text encoder can be quantized separately with --text-encoder-quantize. (#765)
  • Qwen Image 2.1 now accepts common transformer. and diffusion_model. LoRA exports, including fused gate_up weights and modulation/timestep embedding adapters. (#768)
  • Z-Image now loads ComfyUI LoRAs with fused QKV projections and normalization/bias deltas. Adapters containing direct deltas require the default baked inference mode. (#772)
  • Support independently installed mflux namespace extensions, enabling community mflux.web UI packages to coexist across editable checkouts and separate import directories. (#776)
  • Qwen-Image-2.1 reference editing now supports LoRA adapters through the same mappings as text-to-image, including baking controls and metadata replay. Both commands share the Transformer core and component-loading mechanism while retaining compatibility with existing model exports. (#777)
  • mflux-generate-z-image-turbo now exposes ZImageTurboCommand.load and ZImageTurboCommand.generate, so Python code can load the model once and generate many images with the command's own flag handling. (#780)
  • mflux-generate-z-image and mflux-generate-z-image-controlnet now expose load and generate for Python callers, and all three Z-Image commands add validate, which checks a request without loading weights. (#785)
  • mflux-save can now save the SeedVR2 3B and 7B upscalers. mflux-upscale-seedvr2 --model <dir> reads the variant from the saved weights. (#800)
  • mflux-generate-ernie-image-turbo and mflux-generate-ernie-image now expose validate, load and generate for Python callers. (#815)

Improved

  • Qwen-Image-2.1 generates up to ~1.8x faster: denoise steps run ~12% faster (cached text prefix, bounded to the two most recent embeddings; fused Q/K norm+rope Metal kernel, disable with MFLUX_QWEN21_DISABLE_FUSED_PROLOGUE=1), and generate_image(..., teacache_ratio=0.25) reuses the previous noise on the least-changing steps for a further 1.4x (1.77x at 0.4). (#778)
  • Qwen Image 2.1's step reuse is now --step-cache-ratio (--teacache-ratio still works) and the ratio is recorded in image metadata and replayed by --config-from-conf. The logic is shared (mflux.models.common.step_cache.StepCache) so other models can adopt it. (#779)
  • MFlux now imports OpenCV and PyTorch only in the code paths that use them. Commands that do not use these libraries start without them. (#787)
  • LoRA and LoKr adapters now run in the precision the model was loaded in, not float32. Training and LoRA inference on bfloat16 models use less memory and run faster. (#801)

Fixed

  • bump anyio due to CVE-2026-63374 & CVE-2026-64847. MFlux doesn't import it directly - it is used transitively by httpx through huggingface-hub (#744)
  • mflux now refuses a checkpoint whose weight names do not match the model, with an error that names the part and the missing weights. Before, such a checkpoint loaded with random weights and produced noise. This mostly affects checkpoints converted for other programs. (#758)
  • DoRA adapters saved with lora_A / lora_B matrices (the PEFT layout) now stop with an error on every model, instead of loading without their magnitude. LoKr adapters with dora_scale still load. (#774)
  • Krea 2 checkpoints saved by earlier releases (mflux shards under transformer/) load again instead of coming up untrained. (#792)
  • Reference images with transparency are composited over white before the model sees them; masks keep their transparent background as "preserve". (#793)
  • Qwen-Image-2.1 edit: a multi-seed run now finds the auto-mask and rewrites the prompt one time, not one time for each seed. Edit verification now reads plain-text replies with any spacing. (#796)
  • --low-ram now frees the transformer before VAE decode in Z-Image, Krea-2, FLUX.2 Klein, Ming-Image, ERNIE-Image and Ideogram-4. Before, the transformer weights stayed in memory during decode, which increased the peak memory. (#802)
  • Image metadata now keeps LoRA scales, the ControlNet strength and the Redux strengths at full precision. Before, these values were rounded to 2 decimals, so --config-from-conf did not replay the same run. (#805)

Changed

  • Three CLI flags have new names: --make-conf (was --metadata), --no-exif (was --no-metadata) and --config-from-conf (was --config-from-metadata). The old names continue to work as aliases. (#797)

Internal

  • Add Qwen-Image-2.1 to the CI Manifest script (scripts/ci_extract_models.py) .. and improve the Agent directions & Skills to remind AI Agents to do this task. (#742)

Don't miss a new mflux release

NewReleases is sending notifications on new releases.