What's Changed
- Fix Windows builds with CUDA 13 and PyTorch 2.15 by @tuananhlfc in #2872
- [CuTe] Allow apache-tvm-ffi 0.1.11 by @felipemello1 in #2948
- [CuTe, SM100] S/P ping-pong to overlap QK with softmax in decode by @drisspg in #2928
- [CuTe, SM100] Unify hd256 forward and recover the remaining stack by @drisspg in #2962
- [CuTe, SM100] Fix hd128 decode crash on SM103: cap other-WG registers at the launch count by @felipemello1 in #2950
- [Cute,Fwd,Sm100] hd256: S ping-pong on SM103, 2CTA for varlen Q and decode by @felipemello1 in #2951
- [CuTe, SM100] Allow deterministic=True in the hd256 backward by @felipemello1 in #2952
- [Cute, Sm100] bwd postprocess: load 2-CTA dQaccum directly into registers by @pchen7e2 in #2960
- [CuTe, SM100] Unify hd256 backward (1/5): tests first, named 2CTA schedules and TMEM widths by @Johnsonms in #2964
- [CuTe, SM100] Key MLA forward compilation on the KV head count by @drisspg in #2978
New Contributors
- @tuananhlfc made their first contribution in #2872
- @felipemello1 made their first contribution in #2948
- @pchen7e2 made their first contribution in #2960
Full Changelog: fa4-v4.0.0.beta33...fa4-v4.0.0.beta34