github vllm-project/vllm-omni v0.31.0rc1

pre-release4 hours ago

Highlights

vLLM-Omni v0.31.0rc1 includes 110 merged changes since v0.30.0. This prerelease aligns with vLLM 0.31.0, adds Ray-based multi-node inference, expands music generation and realtime serving, and improves speech and multimodal inference performance.

Core Architecture & Runtime

  • Align vLLM-Omni with vLLM 0.31.0, including CUDA/ROCm base images and native auxiliary-output integration. Update routed-expert tests to cover the supported V2 runner path. (#8459, #8500)
  • Add an optional Ray executor for multi-node inference, with cluster-wide stage GPU allocation and environment forwarding. (#7938)
  • Extend shared execution for mixed V1/Model Runner V2 pipelines, add byte backpressure to chunk transfers, and bound duplex output delivery and cancellation. (#8184, #7916, #7644)
  • Add an independent token_ids output type and enforce canonical output-type names. (#7843)

Model Support & Realtime Serving

  • Add YuE2-3B text-to-music through /v1/audio/speech, with lyrics, style instructions and optional ABC-score conditioning. Add async single-stage synthesis and CUDA graph acceleration. (#7886, #8407)
  • Add SheetSage2 preprocessing and an optional YuE2 score-conditioned request export. This is an official-backend preprocessing utility, rather than a native vLLM model or a new serving endpoint. (#8371)
  • Add SenseNova-U1-A3B-MoT, restore PersonaPlex serving on the unified full-duplex framework, and extend OpenAI-compatible realtime serving with Qwen3-Omni and incremental video prefill. (#8172, #7695, #7285, #5894)
  • Add AuK online serving and DiT CUDA graphs, alongside encoder, codec and attention optimizations. (#7469, #7881, #8300, #8328, #8305)

Speech & Audio Performance

  • Add opt-in single-stage Qwen3-TTS on one CUDA GPU, with streaming codec decoding inside the Talker; the existing two-stage pipeline remains available. Improve CUDA throughput and first-audio latency. (#8259, #8163)
  • Improve MiniCPM-o 4.5 with block-aligned duplex KV sliding windows, whole-Euler Code2Wav CUDA graphs, turn-mode MRv2 and batched audio inference. (#7821, #8007, #8222)
  • Improve CosyVoice3 with a GPU-resident F0 predictor, pure-PyTorch vocoder, opt-in packed Flow/batched HiFT execution, bounded streaming graphs and device-dependent defaults. Restore Fish Speech/CosyVoice3 audio quality at high concurrency. (#7518, #8224, #8420, #8433, #8422)
  • Improve MOSS-TTS streaming codec ramps, seeded batch generation and MRv2 serving. Add Higgs Audio v3 MRv2 streaming and Ascend VoxCPM2 LocDiT NPUGraph acceleration. (#8124, #7922, #8213, #8226, #7275)

Diffusion, Image & Video

  • Add MiniMax-H3 latent super-resolution, high-resolution refinement and res_multistep sampling; avoid unnecessarily enlarging reference inputs. (#8322, #8378, #8253)
  • Add Boogu-Image request-level batching, SenseNova online FP8, and bounded long-video restoration with SeedVR2. (#7227, #7955, #8102)
  • Improve MammothModa2 with TeaCache, continuous batching, fused QK normalization/RoPE and text-to-image serving fixes. (#5357, #7954, #7969, #7199, #7293, #7482)
  • Add HSDP pre-sharded loading, prototype FA4 diffusion attention execution contracts, and fix layerwise-offload shutdown, Cosmos3 FP32 sampling state and washed-out LingBot Video output. (#7948, #7379, #8453, #7592, #7037)

Compatibility & Packaging

  • Output-type names now require canonical spellings: text, image, audio, latent and token_ids. Legacy aliases and malformed values are rejected; existing +/, combinations and None defaults remain supported. (#7843)
  • Unsupported async-chunk configurations are disabled with a warning, and async chunk defaults depend on pipeline support. (#5099)
  • Explicitly package runtime deployment files, chat templates, kernel sources and model assets, with source-archive/wheel checks. Fix remaining Python 3.10 compatibility issues. (#7956, #7985)

Validation & Known Issues

  • The release is pinned to 6dd0d1f9310f7598b773c796b434c8000c2816ec, the merge commit of #8500, containing the vLLM 0.31.0 alignment and routed-expert test correction.
  • The completed CUDA nightly build #16714 was reused for this release. Its recorded job states are 54 passed, 3 failed and 4 broken; this is not an entirely passing nightly run.
  • The Omni single-GPU job reported 20 passed, 1 failed and 3 skipped. The failing case is MiniCPM-o test_text_to_audio_long_form_001[async_chunk], which failed its audio/text content assertion. This release does not claim that issue is resolved.
  • Two Wan2.2 jobs (T2V function and device-postprocess equivalence) failed with exit status -1, without an assigned agent, start time or test log. Their test execution remains unverified. The broken steps include ready/merge/weekly uploads gated off for this nightly configuration and the email-distribution step.

Release Artifacts

  • Install from PyPI: vllm-omni 0.31.0rc1. The published vllm_omni-0.31.0rc1-py3-none-any.whl matches the validated CI artifact; version metadata, installation, twine check, and all 136 runtime package-data files were checked.
  • CUDA images: vllm/vllm-omni:v0.31.0rc1 and vllm/vllm-omni:latest, with amd64 and arm64 manifests, based on vLLM 0.31.0. See DockerHub tags.
  • Release artifacts were built and published by omni-release #3462, pinned to the release commit above.

What's Changed

  • [Bugfix] Apply function module marker per-model instead of hardcoding diffusion by @khairulkabir1661 in #5875
  • Add Intel Arc BMG recipes for validated P0 diffusion models. by @Joshna-Medisetty in #7537
  • [CI][Diffusion] Add tiny model builder for Krea2Pipeline by @khairulkabir1661 in #6421
  • [CI][Diffusion] Restore Krea2 model type marker by @andyluo7 in #8164
  • [Refactor][Diffusion] Read the resolved offload policy outside the config layer by @specture724 in #7327
  • [Core][Diffusion] Add byte backpressure to chunk transfers by @specture724 in #7916
  • [Model] Optimize Qwen3-TTS streaming throughput and first audio on CUDA by @Sy0307 in #8163
  • feat(magi2): add fused SwiGLU7 activation by @yeahdongcn in #7252
  • [Diffusion] Enable request-level batching for Boogu-Image (TI2I) by @ShengleiFu in #7227
  • [Misc] Remove stale offline word_timestamps example resurrected by #5146 by @twu3202 in #7233
  • feat: configure Torch Dynamo recompile limit from environment by @yeahdongcn in #7243
  • [Doc] Update README and docs for v0.30.0 by @hsliuustc0106 in #8175
  • [Core][Model] Fit PersonaPlex into the Unified Full-duplex Framework (RFC #7181 PR 2) by @chickeyton in #7695
  • ci: extend AMD Z-Image cold-start timeout by @andyluo7 in #8195
  • [Model] Support SenseNova-U1-A3B-MoT (MoE) by @nussejzz in #8172
  • [Hardware][Ascend] Add FastH3 four-step deployment to MiniMax-H3 NPU … by @kinnara1989 in #7149
  • [Performance][MiniCPM-o] Zero-copy block-aligned KV sliding window with in-place Re-RoPE for duplex Stage-0 by @BeatSeat in #7821
  • test: keep Hunyuan LoRA CPU path on native linear by @andyluo7 in #8166
  • [Perf][PersonaPlex] Serve 16 concurrent realtime duplex sessions on one GPU by @linyueqian in #8192
  • [Misc] Sort OMNI_PIPELINES registry alphabetically by @HelloWorldU in #8208
  • [Bugfix][MiniCPM-o] Forward a speech unit closed by LISTEN after turn_eos to the Talker by @chickeyton in #8227
  • [TTS] Improve voice delete error message by @clodaghwalsh17 in #7795
  • feat(cosyvoice3): GPU-resident F0 predictor & pure-PyTorch vocoder (Task C2) by @BeatSeat in #7518
  • [Perf][Diffusion] support Sensenova online fp8 by @Dmaner in #7955
  • feat(magi2): add fused BF16 routed MoE path by @yeahdongcn in #7206
  • [BugFix] Add field constraints to /v1/omni/sleep and /v1/omni/wakeup by @Shaun-Walsh in #4740
  • [Kernel][MammothModa2] Use shared RMSNorm in the DiT by @Jerry2423 in #7489
  • [Bugfix] Raise when a diffusion stage fails sleep or wake_up by @cr-gao in #7811
  • [Perf][TTS] Run talker_mtp eagerly when a batch carries explicit seeds by @jingchengtian in #8009
  • [Tests] Use get_file_store_init_method helper from vllm by @NickCao in #7528
  • [Bugfix] Avoid enlarging MiniMax H3 reference images and videos by @lishunyang12 in #8253
  • [Frontend] OpenAI-compliant fullduplex /v1/realtime with Qwen3-Omni support by @vraiti in #7285
  • [Kernel] Improve 2 backends importing FlashInfer: FLASHINFER_ATTN and TRTLLM_ATTN by @xrq-phys in #7431
  • [Performance][MiniCPM-o] Whole-Euler CUDA graphs for Code2Wav with a shared attention arena by @BeatSeat in #8007
  • [Feature][TTS] Port static chunk ramp to Moss-TTS streaming codec decode by @Wallbreazzz in #8124
  • [Bugfix][BAGEL] Keep replicated prefill and cache-update phases off the SP strategy by @nussejzz in #8177
  • [Frontend] Tune reference cache limits and registered MOSS voice requests by @gcanlin in #7883
  • SeedVR2: bounded long-video restoration by @0z5a in #8102
  • [Bugfix] Fix Qwen3-Omni realtime routing with typed stage configs by @xuhuan51 in #8279
  • [Core] Extend the shared multi-stage execution framework for mixed V1/MRv2 pipelines by @Sy0307 in #8184
  • test: define Wan fused BF16 parity budget by @andyluo7 in #8188
  • [Bugfix] Break duplex OpenAI import cycle by @NickCao in #8287
  • [Model] Add opt-in packed Flow and batched HiFT inference for CosyVoice3 by @Sy0307 in #8224
  • [Config] Disable Async Chunk for Unsupported Pipelines by @alex-jw-brooks in #5099
  • [Perf][TTS] Add per-row generator support to MOSS-TTS talker to eliminate serial fallback for seeded batched requests by @jingchengtian in #7922
  • [Model] Add MiniCPM-o 4.5 turn-mode MRv2 and batched audio inference by @Sy0307 in #8222
  • [Tests] Fix Qwen3-Omni E2E LiveKit Test by @vraiti in #8294
  • [Perf][Auk] VAE cuda graph and tiling by @BruceLoveDecimal in #7881
  • [Feat][Perf][AuK]Support Online Serving and DiT CUDAGraph by @sphinxkkkbc in #7469
  • [Hardware][Ascend] Accelerate VoxCPM2 LocDiT with NPUGraph by @Big2Wheel in #7275
  • [Perf][AuK] Compile the DiT step, use cuDNN attention and warm up at startup by @linyueqian in #8300
  • [Bugfix] Do not block diffusion client shutdown on a dead subprocess by @cr-gao in #7812
  • [Model] Make VoxCPM2 generation settings configurable by @lolyhop in #7010
  • [Experimental][Perf] Eliminate redundant host copies from generated-video transport by @mglyn in #7355
  • [Benchmark] Add MammothModa2 startup/loading breakdown by @BuckMulligan99 in #7312
  • [Bugfix] Fix the remaining Python 3.10 breakages by @armaanamatya in #7985
  • [Doc] Update vLLM-Omni WeChat QR code by @david6666666 in #8301
  • [Bugfix][Helios] Guard cross-attn KV cache against address reuse by @yancaocn in #8064
  • [CI/Build] Split long weekly TTS jobs and route slow full-model cases to nightly by @yenuo26 in #8239
  • [CI/Build] Refresh A3 and MiniCPM-o H100 performance baselines by @congw729 in #6538
  • [Bugfix] Fix Qwen3-Omni audio-input mean_e2el_ms regression by @psv666 in #8306
  • [Model] Enable Higgs Audio v3 MRV2 streaming and fix mixed-prefill capture bounds by @Sy0307 in #8226
  • [Bugfix] Defer VoxCPM2 NPU patch to avoid platform import cycle by @amy-why-3459 in #8330
  • [Model][Cosmos3] Keep diffusion sampling state in FP32 by @yzhautouskay in #7592
  • [Model] Add MiniMax-H3 latent super-resolution and hi-res refinement by @princepride in #8322
  • [Model] Add YuE2-3B text-to-music (model + /v1/audio/speech adapter) by @Darcy-Lee in #7886
  • [CI/Build] Package runtime data explicitly by @dougbtv in #7956
  • fix(ming): build ISTFT spectrogram in float32 by @Joshna-Medisetty in #7538
  • fix: repair ROCm test routing, CPU timeouts, and Breeze autotuning by @andyluo7 in #8342
  • [Bugfix] Fix washed-out LingBot Video outputs by @wtz2333 in #7037
  • [Model] Optimize Qwen3-Omni MRv2 handoff, code prediction and audio decoding by @Sy0307 in #8223
  • test: measure offload memory inside diffusion workers by @andyluo7 in #8189
  • [CI/Build] Restore streaming executor fixture contract by @princepride in #8418
  • [feature] Add HSDP pre-sharded option by @MaciejBalaNV in #7948
  • [Core] Prototype diffusion attention execution contracts with FA4 by @rahul-steiger-nv in #7379
  • [Model] Add single-stage Qwen3-TTS with in-Talker streaming codec decode by @Sy0307 in #8259
  • [AR-Diffusion] Page in kernel-legal blocks so the default resolution can run by @linzhenpl07 in #6481
  • [Model] YuE2: async single-stage synthesis and CUDA graph acceleration by @Sy0307 in #8407
  • [Bugfix] Restore Fish Speech and CosyVoice3 audio quality at high concurrency by @Sy0307 in #8422
  • [FEAT][Fronted] Support incremental prefill for realtime video streaming by @ZhengWG in #5894
  • [Bugfix] Initialize eager state in Higgs test fixture by @NickCao in #8426
  • [Perf][AuK] Precompute adaLN modulations per request and run the conv position embedding as one GEMM by @linyueqian in #8328
  • [Perf][AuK] Graph the encoder prefill, cache reference latents and fuse the codec activation by @linyueqian in #8305
  • [Model] Optimize CosyVoice3 packed streaming with bounded HiFT graphs by @Sy0307 in #8420
  • [Bugfix] Skip async-chunk SHM put for stage-0-final requests by @ZhengWG in #7245
  • [Model] Optimize MOSS Local 1.5 MRV2 serving and progressive chunks by @Sy0307 in #8213
  • [CI/Build] Fix MOSS CPU tests on ROCm workers by @andyluo7 in #8447
  • [Bugfix] Route MOSS reference attention through ROCm FlashAttention by @andyluo7 in #8463
  • fix: preserve Torch seeded sampling on ROCm by @andyluo7 in #8341
  • fix: preserve CosyVoice full payload and live conditioning by @andyluo7 in #8343
  • [Model] Select CosyVoice3 streaming defaults by device capability by @Sy0307 in #8433
  • [Model] Batch Qwen3-TTS Talker prompt projection and settled-row preprocessing by @Sy0307 in #8477
  • test: bound CPU vocoder reference threads by @andyluo7 in #8347
  • [Bugfix][CI/Build] Fix layerwise offload shutdown and enable it for the LTX-2 distilled T2V accuracy test by @nikiwang-125 in #8453
  • [Feature] Add MiniMax H3 res_multistep sampler by @nodeeeeee in #8378
  • [Bugfix] Add TRTLLM attention custom op and execution contract by @rahul-steiger-nv in #7214
  • ci: route AR paged attention tests to GPU lanes by @andyluo7 in #8346
  • [Perf] Support TeaCache for MammothModa2 by @nagisa-kunhah in #5357
  • [Perf] Support continuous batching for MammothModa2 by @nagisa-kunhah in #7954
  • [Bugfix][Examples][MammothModa2] Serve T2I via LLM-typed DiT stage (#7199) + benchmark harness and profiler wiring by @quan-yi-ai in #7293
  • [Bugfix] Validate MammothModa2 text-to-image dimensions by @TonybotNi in #7482
  • [Model] Fuse MammothModa2 QK norm and RoPE by @Levius-Fubuki in #7969
  • [Core][Frontend] Bound duplex output delivery and reject cancelled audio by @NolenLiang in #7644
  • [Bugfix] Wire output buffers into duplex model test harnesses by @linyueqian in #8489
  • [feature] Enable multi-node inference through Ray backend by @MaciejBalaNV in #7938
  • [Model] Add SheetSage2 preprocessing and YuE2 score handoff by @princepride in #8371
  • [Model][Core] Avoid unused CosyVoice3 payloads and redundant AR output work by @Sorbett in #8473
  • [Rebase] Align vLLM-Omni with vLLM v0.31.0 by @tzhouam in #8459
  • [Core] Add token_ids output type and enforce canonical names by @Jungle430 in #7843
  • [Bugfix] Update routed-expert tests for native auxiliary output by @linyueqian in #8500

New Contributors

Full Changelog: v0.30.0...v0.31.0rc1

Don't miss a new vllm-omni release

NewReleases is sending notifications on new releases.