github Dao-AILab/flash-attention fa4-v4.0.0.beta26

pre-release6 hours ago

What's Changed

  • [CuTe,Bwd,Sm90] Fix: wait for bwd_preprocess on the block-sparse path, matching the dense path by @Fugoes in #2756
  • [Cute, bwd, sm90/100/110] Support learnable sink in backward by @henrylhtsang in #2706
  • [CuTe, FA4] Preserve first-tile flag during scheduler reconstruction by @dongxiao92 in #2705
  • Fix duplicated word in layer norm comment by @cupkk in #2744
  • [CuTe, SM100] Fix deadlock in varlen + block-sparse + SplitKV forward by @JiaxuanBai in #2761
  • [CuTe, Fwd] Fix forward compile key churn when max_seqlen is a tensor by @eamonn-zh in #2762
  • [CuTe] Fix forward dynamic-shape correctness by @drisspg in #2745
  • Fix CLC fuzz scheduler expectations by @JiaxuanBai in #2766
  • Fix removed Quack packed subtraction API by @JiaxuanBai in #2787

New Contributors

Full Changelog: fa4-v4.0.0.beta25...fa4-v4.0.0.beta26

Don't miss a new flash-attention release

NewReleases is sending notifications on new releases.