0.22.0
Improved
mflux-generate-krea2,mflux-generate-lens,mflux-generate-booguandmflux-generate-ideogram4now exposevalidate,loadandgeneratefor Python callers. (#821)mflux-generate-qwenandmflux-generate-qwen-editnow exposevalidate,loadandgeneratefor Python callers, and stop with an error when--modelnames another model instead of ignoring it. (#836)
Fixed
mflux-generate-z-imageandmflux-generate-ernie-imagenow keep a scheduler given as--scheduler=NAME. Before, that spelling was replaced by the command's default scheduler. (#849)
Changed
- Z-Image runs its transformer in bfloat16 instead of float32. Denoising is about 30% faster on an M5, and images for the same seed change slightly. Add
--float32to reproduce an image from an older release. (#803) - Z-Image now writes a warning when a prompt fills the 512-token limit. Before, the model ignored the extra text with no message. (#810)
- LoRA training: a new adapter now starts with
lora_Bat zero, so step 0 is the base model. Before, each new adapter added a random change to every patched layer, which caused a grid pattern in the step-0 preview. (#811) - FLUX.2 klein, Z-Image and Qwen-Image 2.1 take a new opt-in
--compute-precision float16(compute_precision=mx.float16in Python) that runs only the attention and feed-forward layers in float16. On an M1 Max it made each step 8–35% faster and a whole generation 7–29% faster, at the same peak memory; the image for a given seed changes slightly. It is off by default. (#820) - Update urllib3 to 2.8.0 or later to fix HTTPS proxy TLS and streamed-response denial-of-service vulnerabilities. (#822)
- A LoRA baked into a model quantized at load (
-q) on FLUX, FLUX.2, Qwen-Image, Krea 2 and Z-Image now has the same effect as the live adapter. It used to lose part of it, about 20% for a weak LoRA at scale 0.5 on q8. Layers below 8 bits keep their bits instead of being re-quantized at q8. (#824) - Loading a SeedVR2 model saved with
mflux-saveno longer warns that the folder "contains no files matching" the original file names. (#828) - Add Schedulers documentation (#829)
mflux-generate-qwen-2.1andmflux-generate-qwen-2.1-editno longer read every weight into memory at load. For an 8-bit model at 512x512 the peak drops from about 24 GB to about 13 GB, and--low-rambrings it to about 8 GB again in most runs. (#833)mflux-generate-qwen-2.1-edit --schedulernow reaches the edit (it was checked, then ignored), and--enhance-promptrewrites with every reference image instead of only the first.QwenImage21Editgainsrewrite_prompt()and aschedulerargument, and takes PIL images inimage_paths. Both Qwen-Image-2.1 commands record a scheduler other thanlinear, so--config-from-confreplays it. (#835)- Require fsspec 2026.6.0 or later, which fixes a code execution vulnerability in its ReferenceFileSystem (GHSA-27vj-qcqg-25rc). mflux doesn't use fsspec directly. (#838)
- FLUX.1, FIBO and Qwen-Image load the bias of their last normalization layer again. mflux 0.16.0 to 0.21.0 dropped it, so their images change: slightly for FLUX.1, visibly for Qwen-Image. A model saved with those versions generates as before; save it again from the original weights to include the bias. (#839)
- New command
mflux-generate-qwen-2.1-controlnet: Qwen-Image-2.1 with Alibaba's Fun ControlNet Union. One checkpoint takes Canny, depth, pose and five more kinds of control image, and inpaints through the same branch with--image-pathand--mask-image. (#842) - The Qwen-Image-2.1 edit's step cache no longer skips steps in runs under 10 steps, such as the 6-step viggle_turbo schedule, where skipping changed the image noticeably. Baking a LoRA now warns when the baked weights hold less than 90% of its update, and suggests --no-bake-lora. (#843)
- Qwen-Image-Edit-2511 runs as itself:
mflux-generate-qwen-edit --model qwen-image-edit-2511(orqwen-edit-2511) loads the 2511 weights and reads the reference images at timestep 0, as the checkpoint expects. Until now those names loaded 2509. Without--modelthe command still runs 2509. (#844)