cuDNN Frontend v1.26.0 Release Notes
cuDNN Frontend v1.26.0 is the recommended version for cuDNN 9.24.0 and later releases.
Updates to Graph API 🚀 🚀
SDPA
- Added unified-engine FP8 and MXFP8 forward SDPA support (#301). Requires cuDNN 9.25.0 or later (#339).
- SDPA backward with head dimension d=256 now uses the cuDNN-native path on cuDNN 9.23+, bypassing the OSS kernel path (#335).
Data types
- Added a new
BYTE_BOOLEANfrontend data type (#302). Boolean tensors automatically map toBYTE_BOOLEANwhen running against cuDNN 9.25 and later (#339).
Serialization and plan management
Graph::deserializenow accepts anenforce_precompiledoption to require precompiled engine plans during deserialization (#323).- Added a
run_warmupopt-out and a reuse-parsed-json overload toGraph::deserialize, reducing repeated parsing overhead (#329).
Open-Source Kernels (CuTe DSL) 🚀 🚀
Block-sparse attention - Video Sparse Attention
- Added block-sparse attention CuTe DSL kernels for Hopper and Blackwell (#333).
DSA Deepseek Sparse Attention
- Added q causal offsets and SM100F support (#316).
- Optimized the DSA backward SM100 kernel (#318).
- Aligned DSA indexer kernels and fixed dense score-gradient clipping (#297).
Grouped GEMM
grouped_gemm_quant_wrapper_sm100now accepts an optional caller-provided output tensor (#338).- Added SReLU support in the grouped GEMM Hadamard fusion (#315).
- CuTe DSL 4.5 compatibility: migrated
cute.core.ThrMmaandcute.make_fragmentusage (#321) and switched the dGLU dbias reduction to a constexpr loop to fix a DSL 4.5 regression (#322).
Benchmarks, Samples and Documentation ✨✨
- SDPA benchmark now computes SOL% using the sampled SM clock and per-architecture MMA throughput (#314), with refreshed benchmarking artifacts for cuDNN 9.24.0.27 (#306).
- Added a CuTe DSL fusion-kernel benchmark suite with initial B200 and B300 results (#303).
- Test and sample robustness improvements from fuzzer mining across cuDNN 9.18–9.24, including block-scale and SDPA fixes (#330).
- Updated the conv get-plan sample heuristic configuration count (#278).
Bug Fixes 🐛
- Fixed a grid-dimension overflow in the DSA backward convert kernel on SM100 (#331).
- Fixed an illegal memory access in
indexer_topk_wrapper(#312). - Fixed SM100 sparse score recompute compact top-k code generation (#317).
- Fixed the
reduce_dKVvalidity guard incorrectly comparing the top-k column position (#298). - Fixed the scale-factor sort order in
block_scale_quantize.h(#319). - Fixed a synchronization issue in MXFP8 tests (#325).
- Fixed SDPA attribute handling in the flash attention node (#326).
- Fixed an unused ragged-offset version-error variable (#299).
Acknowledgements 🙏
Thanks to everyone who contributed to this release:
@dimitar-asenov, @HollowMan6, @Hyaloid, @jiayus-nvidia, @Jie-Fang, @jiemingz, @NVIDIA-JerryChen, @phu0ngng, @shraiysh, @sraman-rgb, @szluyu99, @take-cheeze, @vincejhan, @Vinnie6167, and @zianglih.