This release brings expanded quantized inference and attention support, kernel performance improvements, and
reliability fixes to the ONNX Runtime CUDA Plugin EP. Highlights include 2-bit MatMulNBits, INT4 paged KV caches,
speculative decoding improvements, compact recurrent-state updates, and safer CUDA Graph execution across sessions.
Please refer to the CUDA Plugin EP Quick Start for usage and installation guidance. The matched ONNX Runtime version is 1.30.0.
Highlights
Plugin Runtime and CUDA Graphs
- Gave each plugin EP session its own device arena, preventing cross-session reuse of CUDA Graph buffers and honoring per-session arena options (#32807).
- Added workspace-memory accounting and reporting for resource-constrained graph partitioning, including packed attention workspace recipes and estimates, and optional-input-aware shape handling (#31962, #32283, #32312, #32321).
- Made
GatherNDcapture-safe and keptCudaAsyncBufferstaging buffers alive across graph replay (#32120, #32121). - Fixed shared-cache scratch lifetime in CUDA MHA, reset plugin EP stream chunks before release, and corrected the user-stream test lifetime (#31968, #31983, #32449).
- Fixed CUDA plugin device discovery on WSL, with improved hardware-identity matching and handling of devices missing from platform discovery (#32517).
Attention and KV Caching
- Added bidirectional
GroupQueryAttentionsupport and anis_causalattribute toPagedAttention(#31704, #32225). - Added INT4 paged KV caches with per-channel scales, including FP16 XQA decode and speculative-decode paths. Packed INT4 cache storage uses half the bytes of INT8 cache storage; availability is controlled by the
onnxruntime_USE_INT4_KV_CACHEbuild option, which defaults to enabled (#32515). - Enabled split-KV for paged FlashAttention decode and improved
PagedAttentiondispatch diagnostics (#32102, #32099). - Expanded paged XQA with group-size-6 support, head size 256, FP16-cache decode at head size 256, and native block tables for 128-token pages (#32108, #32229, #32263, #32127).
- Added paged XQA speculative verification for metadata-bounded query groups of 2-8 tokens at head size 256 and group size 6, including ragged batches and INT8, FP8, and native FP16/BF16 KV caches (#32340).
- Added the
session.gqa_value_layoutsession option for applications using a BNHSGroupQueryAttentionValue cache, with graph-inserted layout conversions; BNSH remains the default (#32139).
New Operator and Model Support
- Added compact
GatedDeltaNetwith transition capture for replaying accepted speculative prefixes, plus BF16 support (#32282, #32307). - Added
VarlenCausalConvWithStatefor packed, variable-length continuous batching, and compact state updates that avoid duplicating full convolution-state checkpoints for speculative prefixes (#32168, #32290). - Added CUDA kernels for the DeepSeek Engram contrib operators
EngramGateandNGramHashMapping(#32268). - Registered BF16
ReduceMeankernels (#32326).
Quantized MatMul and MoE
- Added 2-bit
MatMulNBitssupport, including dequantization and specialized GEMV/batched paths, and fused floating-point-activation/integer-weight (fpA_intB) GEMM/GEMV kernels for eligible FP16/BF16 workloads (#32693, #32699). - Made compact
fpA_intBkernels the default build configuration, extended compactMatMulNBitsto BF16, and refined GEMV eligibility checks. The complete kernel matrix remains available throughonnxruntime_USE_FPA_INTB_GEMM_FULL(#32324, #32721, #32338). - Enabled FP4 QMoE kernels by default in CUDA builds and added Windows support for Blackwell SM120 (#32096, #32163).
- Added an opt-in FP8 DeepGEMM decode path for eligible SM90 QMoE workloads, controlled by
ORT_QMOE_FP4_DEEPGEMMand disabled by default. This path is disabled on Windows because of upstream header requirements (#32122, #32485). - Bounded QMoE workspace with configurable row tiling and bounded FP8 weight-dequantization scratch by tiling over N (#32097, #32129).
- Vectorized NVFP4 weight dequantization for prefill, tuned NVFP4 GEMV tiling for Qwen multi-token prediction, and extended speculative-decode GEMVs to 64 rows (#32128, #32140, #32289).
- Tuned FP4/FP8 GEMV scheduling for 48-SM SM121 GPUs and improved FP8 GEMV residency for grids just past two blocks per SM (#32408, #32409, #32433).
- Reduced expensive
fpA_intBtactic-profiling work and added optional, size-gated M-row chunking for largeMatMulNBitsworkloads. Chunking is disabled by default and can be configured withep.cuda.matmul_nbits_m_chunk_sizeorORT_MATMULNBITS_M_CHUNK_SIZE(#32758, #32810).
General Kernel Performance
- Parallelized
ArgMax/ArgMinover wide last axes, optimized wide-last-axisTopK, and accelerated low-lane INT64CumSum(#32092, #32404, #32238). - Added a single-memcpy
Slicefast path for contiguous subregions (#28902). - Passed
Concatper-input metadata by value and avoided pinned buffers in CUDASplitandConcatfast paths (#32119, #32410).
Reliability and Correctness
- Fixed
ScatterElementsreduction dispatch by element type andAbssigned-zero handling (#29879, #31477). - Handled zero-element inputs/bias in
BiasGeluandFastGeluas no-ops and zero-sized outputs in CUDA random generator kernels (#31698, #31997). - Added bounds and shape validation for per-element
Splitsizes, 8-bitMatMulNBitsg_idx,GatherElementscounts, andScatterNDindex depth (#29461, #31643, #32030, #32034). - Validated CUDA NMS mask sizes, QDQ element counts, per-channel
ImageScalerbias, andCropinputs (#32014, #32029, #32002, #32157). - Bounded
RemovePaddingsequence-token counts and Whisper beam-search cross-QK layer/head indices (#31994, #31998). - Hardened integer arithmetic in
RotaryEmbedding,SparseAttention, CUDA reduction scans, and softmax offsets (#31995, #31996, #32137, #32330).
Build, Packaging, and Dependencies
- Upgraded CUTLASS to 4.7, cuDNN frontend to 1.27, and Protobuf to 33.6 (#32111, #29906).
- Updated CUDA architecture selections and plugin package-test pipelines, and adjusted Linux AArch64 build parallelism (#31989, #32072, #32165).
- Fixed Windows CUDA 12.9 SM120 compilation and Windows ARM64 CUDA plugin packaging (#32114, #32355).
- Fixed
PagedAttentionbuilds without FlashAttention and added the CUDA 13 CCCL include path to plugin builds (#32327, #32392). - Used authenticated package feeds, pinned GitHub Actions to full commit SHAs, and updated artifact upload/download actions (#32005, #32176, #32207, #32209).
Contributors
Thanks to the 17 human contributors who contributed to this release:
@apsonawane, @arnej27959, @baijumeswani, @chilo-ms, @danfiedler-msft, @DKAIN-py, @edgchen1, @eserscor, @javier-intel, @jiafatom, @justinchuby, @kunal-vaishnavi, @MohamedElashri, @Noperi0r, @sanaa-hamel-microsoft,
@tianleiwu, @titaiwangms
This release covers changes since CUDA Plugin EP v0.1.0 affecting CUDA Plugin EP code, CUDA kernels, build integration,
and packaging. The highlights focus on user-facing changes rather than listing every infrastructure-only update.
This summary was drafted with AI assistance from commit history and PR metadata.
Full Changelog: plugin-ep-cuda/v0.1.0...plugin-ep-cuda/v0.2.0