github microsoft/onnxruntime plugin-ep-cuda/v0.3.0
ONNX Runtime CUDA Plugin EP 0.3.0

4 hours ago

This release brings sparse attention and packed continuous-batching support, INT2 and mixed-width QMoE execution,
decode-kernel tuning, and reliability fixes to the ONNX Runtime CUDA Plugin EP. Highlights include new QSA/CSA
attention operators, expanded paged speculative decoding, opt-in MatMul auto-tuning, and improved MXFP4 support
on Blackwell consumer GPUs.

Compatibility

The minimum ONNX Runtime core version remains 1.24.4, but meeting that minimum alone does not guarantee
compatibility for contributed operators. Core and the plugin must use the same contributed-operator schemas;
building both from the same ONNX Runtime revision is recommended. Schema mismatches can cause incorrect execution
or crashes and may not be detected when registering the plugin (#32855).

The matched ONNX Runtime version is 1.31.0 for this release.

Highlights

Plugin Runtime and CUDA Graphs

  • Added checked GroupQueryAttention workspace recipes for KV-cache preparation and XQA, FlashAttention, memory-efficient, and unfused routes, plus conservative workspace estimation for resource-constrained graph partitioning. cuDNN workspace is reported as unavailable rather than incorrectly estimated as zero (#32446, #32453, #32454, #32617).
  • Bounded sliding-window GQA workspace by the resident or staged KV extent rather than the unbounded absolute sequence length, and consolidated runtime scratch sizing and sequence-length-independent route eligibility into shared helpers (#32602, #32945, #33001).
  • Replaced eligible decode-path memset operations with small clearing kernels for fpA_intB split-K semaphores, paged XQA semaphores, and compact convolution state updates, reducing CUDA Graph dependency overhead (#32884).
  • Replaced shared ones-buffer bias broadcasts in Gemm and DecoderAttention with direct broadcasts on the consuming CUDA stream, removed the constant-ones cache, and corrected zero-K Gemm bias scaling (#32852).

Attention and KV Caching

  • Added a cuDNN paged SDPA decode tier to PagedAttention for eligible FP16/BF16 causal workloads with unquantized KV caches. It is selected ahead of FlashAttention when XQA is not the winner, automatically considered on SM90+, and can be disabled with ORT_ENABLE_CUDNN_FLASH_ATTENTION=0. Follow-up fixes stabilize default-scale cache keys and backend selection during CUDA Graph execution (#32493, #32624).
  • Expanded non-causal PagedAttention beyond FlashAttention, including portable paths for small cache pages and supported XQA speculative paths; added coverage for head size 128 with 16/32/64-token pages and ragged batches (#32619).
  • Added causal paged speculative XQA for FP16 query/output, INT8 KV caches, head size 128, GQA group size 6, and query widths 2-8. Unsupported widths and non-causal H128 workloads retain existing fallbacks (#32740).
  • Documented the block-eviction layout contract for sliding-window GQA caches, including how to index resident tokens after multi-token speculative or chunked-prefill steps (#32419).
  • Allowed GQA to borrow unquantized past K/V without allocating or copying present caches when new K/V is empty and both present-cache outputs are omitted, and corrected the unfused sliding-window bound to include the current token (#33175).
  • Skipped redundant present-cache copies in ONNX Attention when the outputs alias the external KV-cache inputs (#31150).

Sparse Attention and Continuous Batching

  • Added CUDA SparsePagedAttention and DynamicSparseAttention for externally selected attention indices over paged and contiguous KV caches. These support Qwen4-Exp QSA selected-main-cache attention and DeepSeek V4 CSA joint local/selected-auxiliary attention, with one softmax across participating entries and optional attention sinks (#32524, #32525).
  • Added SparseAttentionIndexer for QSA token selection and CSA compressed-block selection, including rotary preparation, scoring, deterministic TopK, and cache updates (#32526).
  • Added PackedSparseAttentionIndexer for packed, variable-length continuous batching, with per-request metadata, fixed-capacity state, and selection outputs compatible with SparsePagedAttention (#32618).
  • Optimized dense and packed QSA indexing, added packed QK inputs and broader mask/rotary-cache support, and added optional transactional QSA state-update capture to the packed indexer so speculative decoding can replay an accepted prefix (#32656).
  • Parallelized DynamicSparseAttention across selected-token tiles, bounded split workspace, replaced quadratic duplicate validation with a GPU bitmap, and added single-split fast paths for large-context inference (#32671).

New Operator and Model Support

  • Extended EngramGate and NGramHashMapping for Qwen4-Exp, including normalized gated-value output, autoregressive history, per-head offsets, and EOS/segment resets (#32285).
  • Added VarlenNGramHashMapping for packed DeepSeek Engram batching, preserving request-local history and preventing n-gram mixing across request boundaries, then extended it with Qwen4-Exp hashing semantics (#32358, #32467).
  • Added an activation attribute to GatedRMSNorm for SiLU, its Swish alias, and Sigmoid gating. SiLU remains the default (#32512).
  • Added the weightless hyper-connection fusion operators BranchwiseRMSNorm, ScaledSiLU, HyperConnectionPreMix, and HyperConnectionPostMix (#32687).
  • Extended GatherBlockQuantized to FP8 and FP4 data, including full-axis blocks and scale broadcasting. Behavior change: out-of-range indices now produce zero output slices for both integer and floating-point quantized formats, rather than following ONNX Gather error semantics (#32480).

Quantized MatMul and MoE

  • Added per-projection QMoE weight-width attributes and CUDA execution for uniform INT2 and mixed INT2/INT4 FC1/FC2 combinations (2,2), (2,4), and (4,2), with packed decode paths and a scratch-bounded dequantization fallback for supported configurations (#32697, #32743, #32761).
  • Enabled eligible packed INT2/mixed-width QMoE prefill by default on SM80+ with FP16/BF16 activations, symmetric quantization, and interleaved fused SwiGLU. Support covers block sizes 32/64/128 and the three FC1/FC2 combinations above, subject to alignment, expert-count, and workspace restrictions (#32963, #33045, #33005, #33092).
  • Added block-scaled FP8 QMoE support with routed-expert compaction, alongside the existing per-expert global-scale mode (#32887).
  • Made NVFP4 QMoE decode consume schema-native packed weights and scales directly, avoiding decode-only repacking while retaining grouped-GEMM paths for larger workloads (#32325).
  • Fixed MXFP4 W4A16 QMoE session creation and dispatch on SM120/SM121 and other supported SM80+ configurations, corrected the plugin adapter's FLOAT8E8M0 type mapping, and accelerated FP4 conversion in decode and grouped GEMM (#33067).
  • Added BF16 routing support to CUDA MoE and documented/tested packed token inputs for MoE and QMoE (#32659, #32522).
  • Added opt-in FP16/BF16 MatMul auto-tuning via ep.cuda.enable_gemm_auto_tune=1 or ORT_CUDA_GEMM_AUTO_TUNE=1, choosing among cuBLAS, small-N GEMV, and eligible tinygemm2 kernels on SM90+. cuBLAS remains the default. Graph-enabled sessions validate candidates with CUDA Graph replay before caching the winner; warm-up inference is required before capture (#32876, #32901, #33085).
  • Improved MatMulNBits small-M GEMV tiling, biased near-tie tactic selection toward GEMV for small decode batches, and enabled the W8 M=6-8 fast path for the qualified SM121 FP16 shape K=3584, N=200064, block size 64 (#32876, #32884, #32742).
  • Added an opt-in SM90 DeepGEMM path for eligible MatMulBlockQuantizedFp8Weight workloads with a_scale and block size 128, avoiding the FP16/BF16 weight-dequantization buffer. Enable with ORT_FP8_MATMUL_DEEPGEMM=1 in a supported non-Windows DeepGEMM build; it is disabled by default and native FP8 rounding requires model-accuracy validation (#32599).
  • Added opt-in MatMulNBits GEMV experiments: ep.cuda.fpa_intb_gemv_wave_aware=1 for eligible SM12x FP16/INT4 M=8 tiles, and ep.cuda.fpa_intb_gemv_paired_k=1 (or ORT_FPA_INTB_GEMV_PAIRED_K=1) for eligible FP16/INT4 block-32 M=5-8 profiling buckets. Both are disabled by default. The paired-K tactic is experimental: FP16 partial sums can overflow even when the default GEMV remains finite, so numerical validation is required; tactic profiling checks speed, not numerical safety (#33186, #33188).

General Kernel Performance

  • Added a small innermost-axis Split fast path for 2-4 outputs, passing metadata by value and eliminating per-invocation pinned-memory allocations and metadata uploads (#32706).
  • Simplified CUDA RMSNorm reduction, used FP32 arithmetic for paired FP16 output, and checked actual half2 buffer alignment before selecting vectorized output paths, retaining scalar fallbacks for offset external buffers (#33178).

Reliability and Correctness

  • Fixed CUDA plugin MatMulNBits optional-input detection, including omitted edges and empty placeholders, and restored zero-point/scale type comparisons (#32970).
  • Handled missing optional tensors in MatMulInteger, Conv/FusedConv/ConvTranspose, Squeeze, and GemmFloat8, and corrected the GemmFloat8 C-operand descriptor type (#32971, #31967).
  • Added CUDA attention sequence-start/length validation and release-mode GatherND index bounds checks (#31641, #31646).
  • Validated PackedMultiHeadAttention and RestorePadding token offsets before invalid values can participate in device address calculations (#33098, #33099).
  • Checked CUDA Split dimension products and rejected generation-buffer expansion when the input sequence exceeds the maximum sequence length (#32551, #32553).
  • Hardened antialiased Resize handling for out-of-range coordinates and bounded coefficient windows (#33056, #33095).
  • Fixed GatedDeltaNet shape inference for rank-2/rank-3 packed QKV with omitted key/value inputs, while preserving separate-Q/K/V behavior (#32708).
  • Bound external-data path validation to opened files and strengthened file-handle-based access and mapping (#33063).

Build and Packaging

  • Restricted XQA CUDA compilation to SM80+ targets, avoiding mixed-architecture CUDA 13.3 link failures while preserving fallback builds (#32547).
  • Honored onnxruntime_USE_FLASH_ATTENTION=OFF by excluding FlashAttention CUDA sources, and fixed builds with memory-efficient attention disabled (#32580, #32588).
  • Fixed GCC compilation of SM120 FP4 QMoE TMA kernels and tinygemm2 compilation in Python wheel builds targeting older architectures (#32584, #33111).
  • Fixed CUDA 13+ plugin warning-as-error failures from Protobuf and Windows CUDA/CCCL headers, plus NVCC-generated-object linker warnings (#32854, #33149).
  • Removed the Windows CUDA internal-test CMake dependency cycle and fixed MoE/GQA integration build regressions (#32805, #32932, #33179).
  • Updated CUDA 13 Java packaging and package-test integration, and added Windows vcpkg binary caching (#33131, #32933).

Contributors

Thanks to the 17 human contributors who contributed to this release:

@apsonawane,
@baijumeswani,
@bheu,
@chilo-ms,
@edgchen1,
@eserscor,
@gramalingam,
@hanbitmyths,
@hariharans29,
@jiafatom,
@justinchuby,
@kunal-vaishnavi,
@SIDDARTHAREDDY8,
@skottmckay,
@tianleiwu,
@titaiwangms,
@xadupre


This release covers changes from origin/plugin-ep-cuda/rel-0.2 to origin/plugin-ep-cuda/rel-0.3 affecting CUDA Plugin
EP code, shared CUDA kernels, build integration, and packaging. The highlights focus on user-facing changes rather
than listing every infrastructure-only update.

This summary was drafted with AI assistance from commit history and PR metadata.

Don't miss a new onnxruntime release

NewReleases is sending notifications on new releases.