github flashinfer-ai/flashinfer v0.6.16rc4
Release v0.6.16rc4

latest release: nightly-v0.6.16-20260729
6 hours ago

What's Changed

  • fix(moe): reject incompatible output-scale cubins by @aleozlx in #4213
  • fix(jit): pass map_sm107_to_100f in gen_moe_utils_module by @aleozlx in #4215
  • test: fix #4191 regressions in autotuner_core and symlink race test by @kahyunnam in #4225
  • cherry-pick: #4145 test(sm103): fix FP4 autotuner cache inspection by @kahyunnam in #4231
  • cherry-pick: #4199 fix(xqa): PDL load ordering and SM90 fp8 draft-mask dispatch by @kahyunnam in #4232
  • cherry-pick: #3882 fix mxfp8 gemm quantization / scale-layout warnings by @kahyunnam in #4233
  • fix(trtllm): restrict routed-MoE backends to supported architectures (release port for #4107) by @kahyunnam in #4230
  • fix(moe): name-filter output-scale-incompatible cubins until pinned packages carry mDtypeSfC by @Vinnie6167 in #4235
  • Make CuTe-DSL arch guard env-aware and gate norm's DSL dispatch on it by @Vinnie6167 in #4226
  • chore: bump version to 0.6.16rc4 by @kahyunnam in #4236

Full Changelog: v0.6.16rc3...v0.6.16rc4

Don't miss a new flashinfer release

NewReleases is sending notifications on new releases.