github Dao-AILab/flash-attention fa4-v4.0.0.beta32

pre-release5 hours ago

What's Changed

  • [CuTe, SM100] Sparse MLA bwd: in-kernel recompute-P + token-chunked backward by @abcdabcd987 in #2816
  • Build CUDA extension with C++20 on PyTorch 2.13+ by @zhang-keliang in #2879
  • ci: add PyTorch 2.14 to the wheel build matrix by @Johnsonms in #2892
  • ci: registry-free GPU job (runner-local SIF) + verify the provisioned overlay from a fresh session by @Johnsonms in #2894
  • Compile with c++20 for pytorch 2.13+ by @cih9088 in #2899
  • [ROCM] add FLASH_ATTENTION_USE_SYSTEM_AITER flag and commit bump by @micmelesse in #2900
  • [CuTe, SM100] Support 1..128 Q heads in sparse MLA via in-kernel TMA padding by @drisspg in #2883

New Contributors

Full Changelog: fa4-v4.0.0.beta31...fa4-v4.0.0.beta32

Don't miss a new flash-attention release

NewReleases is sending notifications on new releases.