github NVIDIA/cudnn-frontend v1.26.0
v1.26.0 release

3 months ago

cuDNN Frontend v1.26.0 Release Notes

cuDNN Frontend v1.26.0 is the recommended version for cuDNN 9.24.0 and later releases.

Updates to Graph API 🚀 🚀

SDPA

  • Added unified-engine FP8 and MXFP8 forward SDPA support (#301). Requires cuDNN 9.25.0 or later (#339).
  • SDPA backward with head dimension d=256 now uses the cuDNN-native path on cuDNN 9.23+, bypassing the OSS kernel path (#335).

Data types

  • Added a new BYTE_BOOLEAN frontend data type (#302). Boolean tensors automatically map to BYTE_BOOLEAN when running against cuDNN 9.25 and later (#339).

Serialization and plan management

  • Graph::deserialize now accepts an enforce_precompiled option to require precompiled engine plans during deserialization (#323).
  • Added a run_warmup opt-out and a reuse-parsed-json overload to Graph::deserialize, reducing repeated parsing overhead (#329).

Open-Source Kernels (CuTe DSL) 🚀 🚀

Block-sparse attention - Video Sparse Attention

  • Added block-sparse attention CuTe DSL kernels for Hopper and Blackwell (#333).

DSA Deepseek Sparse Attention

  • Added q causal offsets and SM100F support (#316).
  • Optimized the DSA backward SM100 kernel (#318).
  • Aligned DSA indexer kernels and fixed dense score-gradient clipping (#297).

Grouped GEMM

  • grouped_gemm_quant_wrapper_sm100 now accepts an optional caller-provided output tensor (#338).
  • Added SReLU support in the grouped GEMM Hadamard fusion (#315).
  • CuTe DSL 4.5 compatibility: migrated cute.core.ThrMma and cute.make_fragment usage (#321) and switched the dGLU dbias reduction to a constexpr loop to fix a DSL 4.5 regression (#322).

Benchmarks, Samples and Documentation ✨✨

  • SDPA benchmark now computes SOL% using the sampled SM clock and per-architecture MMA throughput (#314), with refreshed benchmarking artifacts for cuDNN 9.24.0.27 (#306).
  • Added a CuTe DSL fusion-kernel benchmark suite with initial B200 and B300 results (#303).
  • Test and sample robustness improvements from fuzzer mining across cuDNN 9.18–9.24, including block-scale and SDPA fixes (#330).
  • Updated the conv get-plan sample heuristic configuration count (#278).

Bug Fixes 🐛

  • Fixed a grid-dimension overflow in the DSA backward convert kernel on SM100 (#331).
  • Fixed an illegal memory access in indexer_topk_wrapper (#312).
  • Fixed SM100 sparse score recompute compact top-k code generation (#317).
  • Fixed the reduce_dKV validity guard incorrectly comparing the top-k column position (#298).
  • Fixed the scale-factor sort order in block_scale_quantize.h (#319).
  • Fixed a synchronization issue in MXFP8 tests (#325).
  • Fixed SDPA attribute handling in the flash attention node (#326).
  • Fixed an unused ragged-offset version-error variable (#299).

Acknowledgements 🙏

Thanks to everyone who contributed to this release:

@dimitar-asenov, @HollowMan6, @Hyaloid, @jiayus-nvidia, @Jie-Fang, @jiemingz, @NVIDIA-JerryChen, @phu0ngng, @shraiysh, @sraman-rgb, @szluyu99, @take-cheeze, @vincejhan, @Vinnie6167, and @zianglih.

Don't miss a new cudnn-frontend release

NewReleases is sending notifications on new releases.