This release brings sparse attention and packed continuous-batching support, INT2 and mixed-width QMoE execution,
decode-kernel tuning, and reliability fixes to the ONNX Runtime CUDA Plugin EP. Highlights include new QSA/CSA
attention operators, expanded paged speculative decoding, opt-in MatMul auto-tuning, and improved MXFP4 support
on Blackwell consumer GPUs.
Compatibility
The minimum ONNX Runtime core version remains 1.24.4, but meeting that minimum alone does not guarantee
compatibility for contributed operators. Core and the plugin must use the same contributed-operator schemas;
building both from the same ONNX Runtime revision is recommended. Schema mismatches can cause incorrect execution
or crashes and may not be detected when registering the plugin (#32855).
The matched ONNX Runtime version is 1.31.0 for this release.
Highlights
Plugin Runtime and CUDA Graphs
- Added checked
GroupQueryAttentionworkspace recipes for KV-cache preparation and XQA, FlashAttention, memory-efficient, and unfused routes, plus conservative workspace estimation for resource-constrained graph partitioning. cuDNN workspace is reported as unavailable rather than incorrectly estimated as zero (#32446, #32453, #32454, #32617). - Bounded sliding-window GQA workspace by the resident or staged KV extent rather than the unbounded absolute sequence length, and consolidated runtime scratch sizing and sequence-length-independent route eligibility into shared helpers (#32602, #32945, #33001).
- Replaced eligible decode-path memset operations with small clearing kernels for
fpA_intBsplit-K semaphores, paged XQA semaphores, and compact convolution state updates, reducing CUDA Graph dependency overhead (#32884). - Replaced shared ones-buffer bias broadcasts in
GemmandDecoderAttentionwith direct broadcasts on the consuming CUDA stream, removed the constant-ones cache, and corrected zero-KGemmbias scaling (#32852).
Attention and KV Caching
- Added a cuDNN paged SDPA decode tier to
PagedAttentionfor eligible FP16/BF16 causal workloads with unquantized KV caches. It is selected ahead of FlashAttention when XQA is not the winner, automatically considered on SM90+, and can be disabled withORT_ENABLE_CUDNN_FLASH_ATTENTION=0. Follow-up fixes stabilize default-scale cache keys and backend selection during CUDA Graph execution (#32493, #32624). - Expanded non-causal
PagedAttentionbeyond FlashAttention, including portable paths for small cache pages and supported XQA speculative paths; added coverage for head size 128 with 16/32/64-token pages and ragged batches (#32619). - Added causal paged speculative XQA for FP16 query/output, INT8 KV caches, head size 128, GQA group size 6, and query widths 2-8. Unsupported widths and non-causal H128 workloads retain existing fallbacks (#32740).
- Documented the block-eviction layout contract for sliding-window GQA caches, including how to index resident tokens after multi-token speculative or chunked-prefill steps (#32419).
- Allowed GQA to borrow unquantized past K/V without allocating or copying present caches when new K/V is empty and both present-cache outputs are omitted, and corrected the unfused sliding-window bound to include the current token (#33175).
- Skipped redundant present-cache copies in ONNX
Attentionwhen the outputs alias the external KV-cache inputs (#31150).
Sparse Attention and Continuous Batching
- Added CUDA
SparsePagedAttentionandDynamicSparseAttentionfor externally selected attention indices over paged and contiguous KV caches. These support Qwen4-Exp QSA selected-main-cache attention and DeepSeek V4 CSA joint local/selected-auxiliary attention, with one softmax across participating entries and optional attention sinks (#32524, #32525). - Added
SparseAttentionIndexerfor QSA token selection and CSA compressed-block selection, including rotary preparation, scoring, deterministic TopK, and cache updates (#32526). - Added
PackedSparseAttentionIndexerfor packed, variable-length continuous batching, with per-request metadata, fixed-capacity state, and selection outputs compatible withSparsePagedAttention(#32618). - Optimized dense and packed QSA indexing, added packed QK inputs and broader mask/rotary-cache support, and added optional transactional QSA state-update capture to the packed indexer so speculative decoding can replay an accepted prefix (#32656).
- Parallelized
DynamicSparseAttentionacross selected-token tiles, bounded split workspace, replaced quadratic duplicate validation with a GPU bitmap, and added single-split fast paths for large-context inference (#32671).
New Operator and Model Support
- Extended
EngramGateandNGramHashMappingfor Qwen4-Exp, including normalized gated-value output, autoregressive history, per-head offsets, and EOS/segment resets (#32285). - Added
VarlenNGramHashMappingfor packed DeepSeek Engram batching, preserving request-local history and preventing n-gram mixing across request boundaries, then extended it with Qwen4-Exp hashing semantics (#32358, #32467). - Added an
activationattribute toGatedRMSNormfor SiLU, its Swish alias, and Sigmoid gating. SiLU remains the default (#32512). - Added the weightless hyper-connection fusion operators
BranchwiseRMSNorm,ScaledSiLU,HyperConnectionPreMix, andHyperConnectionPostMix(#32687). - Extended
GatherBlockQuantizedto FP8 and FP4 data, including full-axis blocks and scale broadcasting. Behavior change: out-of-range indices now produce zero output slices for both integer and floating-point quantized formats, rather than following ONNXGathererror semantics (#32480).
Quantized MatMul and MoE
- Added per-projection QMoE weight-width attributes and CUDA execution for uniform INT2 and mixed INT2/INT4 FC1/FC2 combinations
(2,2),(2,4), and(4,2), with packed decode paths and a scratch-bounded dequantization fallback for supported configurations (#32697, #32743, #32761). - Enabled eligible packed INT2/mixed-width QMoE prefill by default on SM80+ with FP16/BF16 activations, symmetric quantization, and interleaved fused SwiGLU. Support covers block sizes 32/64/128 and the three FC1/FC2 combinations above, subject to alignment, expert-count, and workspace restrictions (#32963, #33045, #33005, #33092).
- Added block-scaled FP8 QMoE support with routed-expert compaction, alongside the existing per-expert global-scale mode (#32887).
- Made NVFP4 QMoE decode consume schema-native packed weights and scales directly, avoiding decode-only repacking while retaining grouped-GEMM paths for larger workloads (#32325).
- Fixed MXFP4 W4A16 QMoE session creation and dispatch on SM120/SM121 and other supported SM80+ configurations, corrected the plugin adapter's
FLOAT8E8M0type mapping, and accelerated FP4 conversion in decode and grouped GEMM (#33067). - Added BF16 routing support to CUDA MoE and documented/tested packed token inputs for MoE and QMoE (#32659, #32522).
- Added opt-in FP16/BF16
MatMulauto-tuning viaep.cuda.enable_gemm_auto_tune=1orORT_CUDA_GEMM_AUTO_TUNE=1, choosing among cuBLAS, small-N GEMV, and eligible tinygemm2 kernels on SM90+. cuBLAS remains the default. Graph-enabled sessions validate candidates with CUDA Graph replay before caching the winner; warm-up inference is required before capture (#32876, #32901, #33085). - Improved
MatMulNBitssmall-M GEMV tiling, biased near-tie tactic selection toward GEMV for small decode batches, and enabled the W8 M=6-8 fast path for the qualified SM121 FP16 shapeK=3584,N=200064, block size 64 (#32876, #32884, #32742). - Added an opt-in SM90 DeepGEMM path for eligible
MatMulBlockQuantizedFp8Weightworkloads witha_scaleand block size 128, avoiding the FP16/BF16 weight-dequantization buffer. Enable withORT_FP8_MATMUL_DEEPGEMM=1in a supported non-Windows DeepGEMM build; it is disabled by default and native FP8 rounding requires model-accuracy validation (#32599). - Added opt-in
MatMulNBitsGEMV experiments:ep.cuda.fpa_intb_gemv_wave_aware=1for eligible SM12x FP16/INT4 M=8 tiles, andep.cuda.fpa_intb_gemv_paired_k=1(orORT_FPA_INTB_GEMV_PAIRED_K=1) for eligible FP16/INT4 block-32 M=5-8 profiling buckets. Both are disabled by default. The paired-K tactic is experimental: FP16 partial sums can overflow even when the default GEMV remains finite, so numerical validation is required; tactic profiling checks speed, not numerical safety (#33186, #33188).
General Kernel Performance
- Added a small innermost-axis
Splitfast path for 2-4 outputs, passing metadata by value and eliminating per-invocation pinned-memory allocations and metadata uploads (#32706). - Simplified CUDA RMSNorm reduction, used FP32 arithmetic for paired FP16 output, and checked actual half2 buffer alignment before selecting vectorized output paths, retaining scalar fallbacks for offset external buffers (#33178).
Reliability and Correctness
- Fixed CUDA plugin
MatMulNBitsoptional-input detection, including omitted edges and empty placeholders, and restored zero-point/scale type comparisons (#32970). - Handled missing optional tensors in
MatMulInteger,Conv/FusedConv/ConvTranspose,Squeeze, andGemmFloat8, and corrected theGemmFloat8C-operand descriptor type (#32971, #31967). - Added CUDA attention sequence-start/length validation and release-mode
GatherNDindex bounds checks (#31641, #31646). - Validated
PackedMultiHeadAttentionandRestorePaddingtoken offsets before invalid values can participate in device address calculations (#33098, #33099). - Checked CUDA
Splitdimension products and rejected generation-buffer expansion when the input sequence exceeds the maximum sequence length (#32551, #32553). - Hardened antialiased
Resizehandling for out-of-range coordinates and bounded coefficient windows (#33056, #33095). - Fixed
GatedDeltaNetshape inference for rank-2/rank-3 packed QKV with omitted key/value inputs, while preserving separate-Q/K/V behavior (#32708). - Bound external-data path validation to opened files and strengthened file-handle-based access and mapping (#33063).
Build and Packaging
- Restricted XQA CUDA compilation to SM80+ targets, avoiding mixed-architecture CUDA 13.3 link failures while preserving fallback builds (#32547).
- Honored
onnxruntime_USE_FLASH_ATTENTION=OFFby excluding FlashAttention CUDA sources, and fixed builds with memory-efficient attention disabled (#32580, #32588). - Fixed GCC compilation of SM120 FP4 QMoE TMA kernels and tinygemm2 compilation in Python wheel builds targeting older architectures (#32584, #33111).
- Fixed CUDA 13+ plugin warning-as-error failures from Protobuf and Windows CUDA/CCCL headers, plus NVCC-generated-object linker warnings (#32854, #33149).
- Removed the Windows CUDA internal-test CMake dependency cycle and fixed MoE/GQA integration build regressions (#32805, #32932, #33179).
- Updated CUDA 13 Java packaging and package-test integration, and added Windows vcpkg binary caching (#33131, #32933).
Contributors
Thanks to the 17 human contributors who contributed to this release:
@apsonawane,
@baijumeswani,
@bheu,
@chilo-ms,
@edgchen1,
@eserscor,
@gramalingam,
@hanbitmyths,
@hariharans29,
@jiafatom,
@justinchuby,
@kunal-vaishnavi,
@SIDDARTHAREDDY8,
@skottmckay,
@tianleiwu,
@titaiwangms,
@xadupre
This release covers changes from origin/plugin-ep-cuda/rel-0.2 to origin/plugin-ep-cuda/rel-0.3 affecting CUDA Plugin
EP code, shared CUDA kernels, build integration, and packaging. The highlights focus on user-facing changes rather
than listing every infrastructure-only update.
This summary was drafted with AI assistance from commit history and PR metadata.