github Dao-AILab/flash-attention fa4-v4.0.0.beta33

pre-release3 hours ago

What's Changed

  • [CuTe, SM100] Sparse MLA bwd: keep P normalization and dS product in fp32 by @abcdabcd987 in #2893
  • [CuTe, SM100] Sparse MLA training fwd: exact running softmax max by @abcdabcd987 in #2907
  • [CuTe, SM100] Sparse MLA training: dpsum from an O residual by @abcdabcd987 in #2908
  • [CuTe, SM100] Sparse MLA recompute-P: allow 64 Q heads by @abcdabcd987 in #2911
  • [CuTe, SM100] Sparse MLA bwd preprocess: stop the lse_log2 store at the sequence's rows by @abcdabcd987 in #2912
  • [CuTe, SM100] Write real zeros for block-sparse empty tiles by @pashu-cohere in #2906
  • [CuTe, SM90] Fit forward tile to the block-sparse block size by @LiRunGuo in #2903
  • [CuTe, SM90] Skip the causal/local mask on unmasked KV blocks in forward by @LiRunGuo in #2904
  • [ROCm/CK] Fix num_splits heuristic, which could never return more than 1 by @Johnsonms in #2901
  • Support seqused_q/seqused_k in SM100 HD256 backward by @Mellonta in #2891
  • [CuTe, SM100] hd256 fwd: derive lengths and KV ranges from BlockInfo/SeqlenInfoQK by @drisspg in #2917
  • [Cute, Sm100] 1CTA MLA forward (dense + sparse top-k), MLA dispatch heuristic, sparse MLA backward at any head count by @jayhshah in #2938
  • Fix lint failure by @jayhshah in #2940

New Contributors

Full Changelog: fa4-v4.0.0.beta32...fa4-v4.0.0.beta33

Don't miss a new flash-attention release

NewReleases is sending notifications on new releases.