github Dao-AILab/flash-attention fa4-v4.0.0.beta34

pre-release5 hours ago

What's Changed

  • Fix Windows builds with CUDA 13 and PyTorch 2.15 by @tuananhlfc in #2872
  • [CuTe] Allow apache-tvm-ffi 0.1.11 by @felipemello1 in #2948
  • [CuTe, SM100] S/P ping-pong to overlap QK with softmax in decode by @drisspg in #2928
  • [CuTe, SM100] Unify hd256 forward and recover the remaining stack by @drisspg in #2962
  • [CuTe, SM100] Fix hd128 decode crash on SM103: cap other-WG registers at the launch count by @felipemello1 in #2950
  • [Cute,Fwd,Sm100] hd256: S ping-pong on SM103, 2CTA for varlen Q and decode by @felipemello1 in #2951
  • [CuTe, SM100] Allow deterministic=True in the hd256 backward by @felipemello1 in #2952
  • [Cute, Sm100] bwd postprocess: load 2-CTA dQaccum directly into registers by @pchen7e2 in #2960
  • [CuTe, SM100] Unify hd256 backward (1/5): tests first, named 2CTA schedules and TMEM widths by @Johnsonms in #2964
  • [CuTe, SM100] Key MLA forward compilation on the KV head count by @drisspg in #2978

New Contributors

Full Changelog: fa4-v4.0.0.beta33...fa4-v4.0.0.beta34

Don't miss a new flash-attention release

NewReleases is sending notifications on new releases.