Highlights
vLLM-Omni v0.31.0rc1 includes 110 merged changes since v0.30.0. This prerelease aligns with vLLM 0.31.0, adds Ray-based multi-node inference, expands music generation and realtime serving, and improves speech and multimodal inference performance.
Core Architecture & Runtime
- Align vLLM-Omni with vLLM 0.31.0, including CUDA/ROCm base images and native auxiliary-output integration. Update routed-expert tests to cover the supported V2 runner path. (#8459, #8500)
- Add an optional Ray executor for multi-node inference, with cluster-wide stage GPU allocation and environment forwarding. (#7938)
- Extend shared execution for mixed V1/Model Runner V2 pipelines, add byte backpressure to chunk transfers, and bound duplex output delivery and cancellation. (#8184, #7916, #7644)
- Add an independent
token_idsoutput type and enforce canonical output-type names. (#7843)
Model Support & Realtime Serving
- Add YuE2-3B text-to-music through
/v1/audio/speech, with lyrics, style instructions and optional ABC-score conditioning. Add async single-stage synthesis and CUDA graph acceleration. (#7886, #8407) - Add SheetSage2 preprocessing and an optional YuE2 score-conditioned request export. This is an official-backend preprocessing utility, rather than a native vLLM model or a new serving endpoint. (#8371)
- Add SenseNova-U1-A3B-MoT, restore PersonaPlex serving on the unified full-duplex framework, and extend OpenAI-compatible realtime serving with Qwen3-Omni and incremental video prefill. (#8172, #7695, #7285, #5894)
- Add AuK online serving and DiT CUDA graphs, alongside encoder, codec and attention optimizations. (#7469, #7881, #8300, #8328, #8305)
Speech & Audio Performance
- Add opt-in single-stage Qwen3-TTS on one CUDA GPU, with streaming codec decoding inside the Talker; the existing two-stage pipeline remains available. Improve CUDA throughput and first-audio latency. (#8259, #8163)
- Improve MiniCPM-o 4.5 with block-aligned duplex KV sliding windows, whole-Euler Code2Wav CUDA graphs, turn-mode MRv2 and batched audio inference. (#7821, #8007, #8222)
- Improve CosyVoice3 with a GPU-resident F0 predictor, pure-PyTorch vocoder, opt-in packed Flow/batched HiFT execution, bounded streaming graphs and device-dependent defaults. Restore Fish Speech/CosyVoice3 audio quality at high concurrency. (#7518, #8224, #8420, #8433, #8422)
- Improve MOSS-TTS streaming codec ramps, seeded batch generation and MRv2 serving. Add Higgs Audio v3 MRv2 streaming and Ascend VoxCPM2 LocDiT NPUGraph acceleration. (#8124, #7922, #8213, #8226, #7275)
Diffusion, Image & Video
- Add MiniMax-H3 latent super-resolution, high-resolution refinement and
res_multistepsampling; avoid unnecessarily enlarging reference inputs. (#8322, #8378, #8253) - Add Boogu-Image request-level batching, SenseNova online FP8, and bounded long-video restoration with SeedVR2. (#7227, #7955, #8102)
- Improve MammothModa2 with TeaCache, continuous batching, fused QK normalization/RoPE and text-to-image serving fixes. (#5357, #7954, #7969, #7199, #7293, #7482)
- Add HSDP pre-sharded loading, prototype FA4 diffusion attention execution contracts, and fix layerwise-offload shutdown, Cosmos3 FP32 sampling state and washed-out LingBot Video output. (#7948, #7379, #8453, #7592, #7037)
Compatibility & Packaging
- Output-type names now require canonical spellings:
text,image,audio,latentandtoken_ids. Legacy aliases and malformed values are rejected; existing+/,combinations andNonedefaults remain supported. (#7843) - Unsupported async-chunk configurations are disabled with a warning, and async chunk defaults depend on pipeline support. (#5099)
- Explicitly package runtime deployment files, chat templates, kernel sources and model assets, with source-archive/wheel checks. Fix remaining Python 3.10 compatibility issues. (#7956, #7985)
Validation & Known Issues
- The release is pinned to
6dd0d1f9310f7598b773c796b434c8000c2816ec, the merge commit of #8500, containing the vLLM 0.31.0 alignment and routed-expert test correction. - The completed CUDA nightly build #16714 was reused for this release. Its recorded job states are 54 passed, 3 failed and 4 broken; this is not an entirely passing nightly run.
- The Omni single-GPU job reported 20 passed, 1 failed and 3 skipped. The failing case is MiniCPM-o
test_text_to_audio_long_form_001[async_chunk], which failed its audio/text content assertion. This release does not claim that issue is resolved. - Two Wan2.2 jobs (T2V function and device-postprocess equivalence) failed with exit status
-1, without an assigned agent, start time or test log. Their test execution remains unverified. The broken steps include ready/merge/weekly uploads gated off for this nightly configuration and the email-distribution step.
Release Artifacts
- Install from PyPI: vllm-omni 0.31.0rc1. The published
vllm_omni-0.31.0rc1-py3-none-any.whlmatches the validated CI artifact; version metadata, installation,twine check, and all 136 runtime package-data files were checked. - CUDA images:
vllm/vllm-omni:v0.31.0rc1andvllm/vllm-omni:latest, with amd64 and arm64 manifests, based on vLLM 0.31.0. See DockerHub tags. - Release artifacts were built and published by omni-release #3462, pinned to the release commit above.
What's Changed
- [Bugfix] Apply function module marker per-model instead of hardcoding diffusion by @khairulkabir1661 in #5875
- Add Intel Arc BMG recipes for validated P0 diffusion models. by @Joshna-Medisetty in #7537
- [CI][Diffusion] Add tiny model builder for Krea2Pipeline by @khairulkabir1661 in #6421
- [CI][Diffusion] Restore Krea2 model type marker by @andyluo7 in #8164
- [Refactor][Diffusion] Read the resolved offload policy outside the config layer by @specture724 in #7327
- [Core][Diffusion] Add byte backpressure to chunk transfers by @specture724 in #7916
- [Model] Optimize Qwen3-TTS streaming throughput and first audio on CUDA by @Sy0307 in #8163
- feat(magi2): add fused SwiGLU7 activation by @yeahdongcn in #7252
- [Diffusion] Enable request-level batching for Boogu-Image (TI2I) by @ShengleiFu in #7227
- [Misc] Remove stale offline word_timestamps example resurrected by #5146 by @twu3202 in #7233
- feat: configure Torch Dynamo recompile limit from environment by @yeahdongcn in #7243
- [Doc] Update README and docs for v0.30.0 by @hsliuustc0106 in #8175
- [Core][Model] Fit PersonaPlex into the Unified Full-duplex Framework (RFC #7181 PR 2) by @chickeyton in #7695
- ci: extend AMD Z-Image cold-start timeout by @andyluo7 in #8195
- [Model] Support SenseNova-U1-A3B-MoT (MoE) by @nussejzz in #8172
- [Hardware][Ascend] Add FastH3 four-step deployment to MiniMax-H3 NPU … by @kinnara1989 in #7149
- [Performance][MiniCPM-o] Zero-copy block-aligned KV sliding window with in-place Re-RoPE for duplex Stage-0 by @BeatSeat in #7821
- test: keep Hunyuan LoRA CPU path on native linear by @andyluo7 in #8166
- [Perf][PersonaPlex] Serve 16 concurrent realtime duplex sessions on one GPU by @linyueqian in #8192
- [Misc] Sort OMNI_PIPELINES registry alphabetically by @HelloWorldU in #8208
- [Bugfix][MiniCPM-o] Forward a speech unit closed by LISTEN after turn_eos to the Talker by @chickeyton in #8227
- [TTS] Improve voice delete error message by @clodaghwalsh17 in #7795
- feat(cosyvoice3): GPU-resident F0 predictor & pure-PyTorch vocoder (Task C2) by @BeatSeat in #7518
- [Perf][Diffusion] support Sensenova online fp8 by @Dmaner in #7955
- feat(magi2): add fused BF16 routed MoE path by @yeahdongcn in #7206
- [BugFix] Add field constraints to /v1/omni/sleep and /v1/omni/wakeup by @Shaun-Walsh in #4740
- [Kernel][MammothModa2] Use shared RMSNorm in the DiT by @Jerry2423 in #7489
- [Bugfix] Raise when a diffusion stage fails sleep or wake_up by @cr-gao in #7811
- [Perf][TTS] Run talker_mtp eagerly when a batch carries explicit seeds by @jingchengtian in #8009
- [Tests] Use get_file_store_init_method helper from vllm by @NickCao in #7528
- [Bugfix] Avoid enlarging MiniMax H3 reference images and videos by @lishunyang12 in #8253
- [Frontend] OpenAI-compliant fullduplex /v1/realtime with Qwen3-Omni support by @vraiti in #7285
- [Kernel] Improve 2 backends importing FlashInfer: FLASHINFER_ATTN and TRTLLM_ATTN by @xrq-phys in #7431
- [Performance][MiniCPM-o] Whole-Euler CUDA graphs for Code2Wav with a shared attention arena by @BeatSeat in #8007
- [Feature][TTS] Port static chunk ramp to Moss-TTS streaming codec decode by @Wallbreazzz in #8124
- [Bugfix][BAGEL] Keep replicated prefill and cache-update phases off the SP strategy by @nussejzz in #8177
- [Frontend] Tune reference cache limits and registered MOSS voice requests by @gcanlin in #7883
- SeedVR2: bounded long-video restoration by @0z5a in #8102
- [Bugfix] Fix Qwen3-Omni realtime routing with typed stage configs by @xuhuan51 in #8279
- [Core] Extend the shared multi-stage execution framework for mixed V1/MRv2 pipelines by @Sy0307 in #8184
- test: define Wan fused BF16 parity budget by @andyluo7 in #8188
- [Bugfix] Break duplex OpenAI import cycle by @NickCao in #8287
- [Model] Add opt-in packed Flow and batched HiFT inference for CosyVoice3 by @Sy0307 in #8224
- [Config] Disable Async Chunk for Unsupported Pipelines by @alex-jw-brooks in #5099
- [Perf][TTS] Add per-row generator support to MOSS-TTS talker to eliminate serial fallback for seeded batched requests by @jingchengtian in #7922
- [Model] Add MiniCPM-o 4.5 turn-mode MRv2 and batched audio inference by @Sy0307 in #8222
- [Tests] Fix Qwen3-Omni E2E LiveKit Test by @vraiti in #8294
- [Perf][Auk] VAE cuda graph and tiling by @BruceLoveDecimal in #7881
- [Feat][Perf][AuK]Support Online Serving and DiT CUDAGraph by @sphinxkkkbc in #7469
- [Hardware][Ascend] Accelerate VoxCPM2 LocDiT with NPUGraph by @Big2Wheel in #7275
- [Perf][AuK] Compile the DiT step, use cuDNN attention and warm up at startup by @linyueqian in #8300
- [Bugfix] Do not block diffusion client shutdown on a dead subprocess by @cr-gao in #7812
- [Model] Make VoxCPM2 generation settings configurable by @lolyhop in #7010
- [Experimental][Perf] Eliminate redundant host copies from generated-video transport by @mglyn in #7355
- [Benchmark] Add MammothModa2 startup/loading breakdown by @BuckMulligan99 in #7312
- [Bugfix] Fix the remaining Python 3.10 breakages by @armaanamatya in #7985
- [Doc] Update vLLM-Omni WeChat QR code by @david6666666 in #8301
- [Bugfix][Helios] Guard cross-attn KV cache against address reuse by @yancaocn in #8064
- [CI/Build] Split long weekly TTS jobs and route slow full-model cases to nightly by @yenuo26 in #8239
- [CI/Build] Refresh A3 and MiniCPM-o H100 performance baselines by @congw729 in #6538
- [Bugfix] Fix Qwen3-Omni audio-input mean_e2el_ms regression by @psv666 in #8306
- [Model] Enable Higgs Audio v3 MRV2 streaming and fix mixed-prefill capture bounds by @Sy0307 in #8226
- [Bugfix] Defer VoxCPM2 NPU patch to avoid platform import cycle by @amy-why-3459 in #8330
- [Model][Cosmos3] Keep diffusion sampling state in FP32 by @yzhautouskay in #7592
- [Model] Add MiniMax-H3 latent super-resolution and hi-res refinement by @princepride in #8322
- [Model] Add YuE2-3B text-to-music (model + /v1/audio/speech adapter) by @Darcy-Lee in #7886
- [CI/Build] Package runtime data explicitly by @dougbtv in #7956
- fix(ming): build ISTFT spectrogram in float32 by @Joshna-Medisetty in #7538
- fix: repair ROCm test routing, CPU timeouts, and Breeze autotuning by @andyluo7 in #8342
- [Bugfix] Fix washed-out LingBot Video outputs by @wtz2333 in #7037
- [Model] Optimize Qwen3-Omni MRv2 handoff, code prediction and audio decoding by @Sy0307 in #8223
- test: measure offload memory inside diffusion workers by @andyluo7 in #8189
- [CI/Build] Restore streaming executor fixture contract by @princepride in #8418
- [feature] Add HSDP pre-sharded option by @MaciejBalaNV in #7948
- [Core] Prototype diffusion attention execution contracts with FA4 by @rahul-steiger-nv in #7379
- [Model] Add single-stage Qwen3-TTS with in-Talker streaming codec decode by @Sy0307 in #8259
- [AR-Diffusion] Page in kernel-legal blocks so the default resolution can run by @linzhenpl07 in #6481
- [Model] YuE2: async single-stage synthesis and CUDA graph acceleration by @Sy0307 in #8407
- [Bugfix] Restore Fish Speech and CosyVoice3 audio quality at high concurrency by @Sy0307 in #8422
- [FEAT][Fronted] Support incremental prefill for realtime video streaming by @ZhengWG in #5894
- [Bugfix] Initialize eager state in Higgs test fixture by @NickCao in #8426
- [Perf][AuK] Precompute adaLN modulations per request and run the conv position embedding as one GEMM by @linyueqian in #8328
- [Perf][AuK] Graph the encoder prefill, cache reference latents and fuse the codec activation by @linyueqian in #8305
- [Model] Optimize CosyVoice3 packed streaming with bounded HiFT graphs by @Sy0307 in #8420
- [Bugfix] Skip async-chunk SHM put for stage-0-final requests by @ZhengWG in #7245
- [Model] Optimize MOSS Local 1.5 MRV2 serving and progressive chunks by @Sy0307 in #8213
- [CI/Build] Fix MOSS CPU tests on ROCm workers by @andyluo7 in #8447
- [Bugfix] Route MOSS reference attention through ROCm FlashAttention by @andyluo7 in #8463
- fix: preserve Torch seeded sampling on ROCm by @andyluo7 in #8341
- fix: preserve CosyVoice full payload and live conditioning by @andyluo7 in #8343
- [Model] Select CosyVoice3 streaming defaults by device capability by @Sy0307 in #8433
- [Model] Batch Qwen3-TTS Talker prompt projection and settled-row preprocessing by @Sy0307 in #8477
- test: bound CPU vocoder reference threads by @andyluo7 in #8347
- [Bugfix][CI/Build] Fix layerwise offload shutdown and enable it for the LTX-2 distilled T2V accuracy test by @nikiwang-125 in #8453
- [Feature] Add MiniMax H3 res_multistep sampler by @nodeeeeee in #8378
- [Bugfix] Add TRTLLM attention custom op and execution contract by @rahul-steiger-nv in #7214
- ci: route AR paged attention tests to GPU lanes by @andyluo7 in #8346
- [Perf] Support TeaCache for MammothModa2 by @nagisa-kunhah in #5357
- [Perf] Support continuous batching for MammothModa2 by @nagisa-kunhah in #7954
- [Bugfix][Examples][MammothModa2] Serve T2I via LLM-typed DiT stage (#7199) + benchmark harness and profiler wiring by @quan-yi-ai in #7293
- [Bugfix] Validate MammothModa2 text-to-image dimensions by @TonybotNi in #7482
- [Model] Fuse MammothModa2 QK norm and RoPE by @Levius-Fubuki in #7969
- [Core][Frontend] Bound duplex output delivery and reject cancelled audio by @NolenLiang in #7644
- [Bugfix] Wire output buffers into duplex model test harnesses by @linyueqian in #8489
- [feature] Enable multi-node inference through Ray backend by @MaciejBalaNV in #7938
- [Model] Add SheetSage2 preprocessing and YuE2 score handoff by @princepride in #8371
- [Model][Core] Avoid unused CosyVoice3 payloads and redundant AR output work by @Sorbett in #8473
- [Rebase] Align vLLM-Omni with vLLM v0.31.0 by @tzhouam in #8459
- [Core] Add token_ids output type and enforce canonical names by @Jungle430 in #7843
- [Bugfix] Update routed-expert tests for native auxiliary output by @linyueqian in #8500
New Contributors
- @kinnara1989 made their first contribution in #7149
- @HelloWorldU made their first contribution in #8208
- @xuhuan51 made their first contribution in #8279
- @BuckMulligan99 made their first contribution in #7312
- @Darcy-Lee made their first contribution in #7886
- @nikiwang-125 made their first contribution in #8453
- @quan-yi-ai made their first contribution in #7293
- @TonybotNi made their first contribution in #7482
- @Sorbett made their first contribution in #8473
- @Jungle430 made their first contribution in #7843
Full Changelog: v0.30.0...v0.31.0rc1