github flashinfer-ai/flashinfer v0.7.1rc1
Release v0.7.1rc1

pre-release3 hours ago

What's Changed

  • feat(comm): copy-engine based kernel for the PCIe IPC all-reduce by @qsang-nv in #4870
  • feat(cake_warp_decode): Add SM100 NVFP4 warp decode and standalone SiLU by @yyihuang in #5036
  • feat(cake_nvfp4_svdquant): add NVFP4 SVDQuant GEMM SM100 and SM103 backend by @yyihuang in #5076
  • fix(moe): bound routing_replay_out dim0 from below in both validators by @aleozlx in #5072
  • fix: guard BF16 warp Split-K availability and prune tactics by @jiahanc in #5089
  • refactor(kda): remove the unreleased recurrent KDA training/backward family by @YangXu1990uiuc in #4965
  • fix(activation): pick a vector width that divides hidden_size in act_and_mul (fixes d=3420) by @Wint3rNight in #5013
  • docs: document MiniMax-H3 prepared run parameters by @kangbintNV in #5087
  • perf(moe): improve SM90 mixed-input CUTLASS MoE backend performance by @StudyingShao in #5005
  • refactor(kda): stop 0.7 advertising unreleased KDA surface (mark prefill wrapper experimental, trim kda_kernels.all) by @kahyunnam in #5040
  • bench: compare DeepSeek MoE against pure TRTLLM BF16 by @zianglih in #4985
  • perf(topk_varlen): gvr_2 dispatch rungs for 4K-8K rows, SM-count-aware register band, hint-free gvr_2 in auto by @dhiraj113 in #4986
  • feat(cake_mega_moe): add SM100 BF16 rank-major MegaMoE backend by @yyihuang in #4810
  • fix(attention): default skip_all_rows_active_check to true by @saltyminty in #5039
  • feat(cake_fmha): add sm100/103 paged-attention by @yyihuang in #4980
  • fix: preserve both generated-source pre-commit exclusions by @saltyminty in #5113
  • fix(fp4_gemm autotune) SM107 kernel family excluded from autotune candidates and heuristics update by @Victor49152 in #4856
  • fix(mla): halve sparse-MLA cpb L2 guard-rail window on integrated GPUs (Spark) by @bkryu in #5097
  • fix(cake_sage_attention): support non-contiguous Sage attention with partial KV tiles by @yyihuang in #5083
  • Fix MoE Tile Mapping Zeroing by @Vinnie6167 in #4735
  • fix(test): pin the GDN disk-cache round-trip test to the flashinfer backend by @bkryu in #5104
  • feat(cake_kda): add fused decode backend for B200 and B300 by @yyihuang in #4967
  • feat(gemm): cuTile alpha-beta / masked-bmm / ragged-bmm / block-scaled FP8 GEMM by @yifeis-nv in #4020
  • feat(cake_comm): add opt-in Blackwell MNNVL MoE all-to-all backend by @yyihuang in #4784
  • feat(cake_diffusion): add SM103a BF16 pre-attention by @yyihuang in #4690
  • fix(moe): Fix per-token quantization fast-math INF scale when activation contains all-zero row by @xuantengh in #5031
  • Fix bmm_fp8 cublasLt handle usage in autotuned cublas runner #26381 by @baonudesifeizhai in #2808
  • build: raise the nvidia-cudnn-frontend floor to 1.28.0 by @Anerudhan in #5131
  • fix(moe): avoid FP8 scale overflow in the Humming MXFP4 path by @yilin-void in #5024
  • chore: add codebase-wise approvers to every CODEOWNERS rule by @bkryu in #5146
  • Add CUTLASS Primitives and Task Scheduling Blackwell MoE kernels by @nekorobov in #4361
  • fix[mla]: add Q8 floor for tile selection in trtllm-gen sparse mla kernels by @jimmyzho in #5106
  • feat(quantization): support fp32 input in the CuTe-DSL mxfp8_quantize kernel by @bkryu in #5112
  • feat(sm120): runtime KV row stride + canonical 528B GLM53_NOPE rows + NaN-safe masked gathers in sparse MLA by @lucifer1004 in #5075
  • Revert "build: raise the nvidia-cudnn-frontend floor to 1.28.0" by @aleozlx in #5159
  • feat(cake_diffusion): add SM100a support to MiniMax-H3 BF16 pre-attention by @yyihuang in #5137
  • fix(moe): accept mxfp8's flat linear activation-scale layout in the MoE autotuner by @aleozlx in #4943
  • fix(comm): make TorchDistBackend.Split() synchronizing, like MPI_Comm_split by @aleozlx in #4206
  • fix(sampling): preserve softmax normalization at low temperatures by @Stelath in #5088
  • fix(infra): isolate distributed rendezvous ports by test worker by @yongwww in #5153
  • feat(cake_bf16_fp4_gemm): add SM100/SM103 BF16 x FP4 GEMM kernels by @yyihuang in #5138
  • feat(cake_mega_moe): add reusable MegaMoE workspaces and an SM100a reducer by @yyihuang in #4819
  • perf(cake_mega_moe): optimize BF16 x MXFP8 MegaMoE EP16 backend by @hzfan in #5148
  • Add cuDNN backend for linear attention by @jhjpark in #4968
  • feat(moe): remove QuantVariant and dispatch QuantConfig by MMA pair by @feih-nv in #5061
  • fix(cake_sage_attn): cover Sage combinations and pass tensor maps by value by @yyihuang in #5127
  • fix(cutile): stabilize autotune config selection in select GEMMs by @elwhyjay in #4459
  • test(comm): report every rank's fate when a dist worker dies by @aleozlx in #4205
  • fix[GEMM]: WAR for specific cuTile autotune candidate in mm_bf16 by @jimmyzho in #5155
  • ci: apply --prerelease to rc in release workflow by @aleozlx in #4472
  • fix(moe): validate TMA warp-specialized configs in queryOccupancyForConfig by @SSHdotCodes in #4847
  • test(cake_comm): initialize NVSHMEM for BF16 rank-major GPU regression by @yyihuang in #5126
  • feat(cake_comm): add Cake MoE finalize all-reduce backend by @yyihuang in #4875
  • feat(cake_alpha_moe): add optimized Blackwell W8A8 expert up/down compute by @xslingcn in #4287
  • perf(cake_comm): reduce PCIe CE ring publication overhead on SM120 by @yyihuang in #5169
  • feat(prims-ts): add non-absorbed MLA prefill by @PerkzZheng in #5059
  • BF16 cutedsl prune [2] + cache bf16 cutedsl kernels by @jiahanc in #5140
  • docs: fix 0.7.0 documentation check findings by @kangbintNV in #5199
  • build: raise the nvidia-cudnn-frontend floor to 1.29.0 by @Anerudhan in #5182
  • fix(cake_attention): enable DCP on SM107 and skip unsupported VSA tests by @saltyminty in #5111
  • feat(cake_comm): support SM100 and SM103 TP8 packed-QKV all-gather matmul by @yyihuang in #5116
  • fix: add discoverable MLA prefill API by @saltyminty in #5158
  • feat(moe): tag the unified MoE API as official with @flashinfer_api by @aleozlx in #5200
  • perf(attention): specialize paged FA2 equal-stride KV by @saltyminty in #4736
  • feat(comm): add allocation-stable head-chunk Ulysses primitives by @tiffany940107 in #5027
  • test(kda): xfail CuTe DSL recurrent prefill tests on nvidia-cutlass-dsl<4.7 by @bkryu in #5219
  • Paged-MQA logits: Rubin (SM107) enablement, FP4 next_n=4, and DKG epilogue/scheduler optimizations by @dhiraj113 in #4737
  • feat(gemm): add SM10x contiguous grouped FP8 CuTe backend by @samodi-nv in #4734
  • feat: enable large-head attention on SM80, for VLLM to run FP8 kv by @jhaotingc in #5044
  • feat(cake_bf16_bmm): add optimized SM100a and SM103a backend by @yyihuang in #4370
  • fix(comm): fence DCP receiver unpack before async output copy by @klshuster in #5172
  • feat(prims-ts): add QToken-KvBlock-Sparse-Attention by @PerkzZheng in #4996
  • ci(jit-cache): enable provider-only wheel releases by @dierksen in #5096
  • fix(autotuner): recover from builtin memory errors by @ZenAlexa in #4920
  • feat(mla): add planned TRTLLM-GEN, XQA, and CuTe DSL backends by @saltyminty in #5070
  • fix(topk): handle padded short rows in CUB transforms by @zianglih in #5187
  • fix(comm): match the library filename when scanning /proc/self/maps by @JacobHelwig in #5198
  • feat: add low-latency Frost DSv4.1 sparse MLA decode by @YangXu1990uiuc in #5119
  • fix(ci): support installed BF16 rank-major session tests by @XFDG in #5171
  • fix: restore default GDN prefill dispatch to CuTe by @yyihuang in #5225
  • refactor(moe): update and refactor the SM90 CuTe-DSL MoE tactics by @Aneureka in #5004
  • feat(prefill): let backend="auto" reach cuDNN/CUTLASS on Blackwell ragged prefill (up to 3.3x) by @Anerudhan in #5133
  • feat(prims-ts): support QK-BF16/PV-FP8 in FMHA context by @harrisonzhy in #4879
  • ci: cache checksum-verified cubin downloads by @dierksen in #5240
  • fix(jit): refresh stale arch detection and allow SM120 on CUDA 12.8 by @zhougit86 in #3633
  • fix(gdn): compile and tune CuTe-DSL decode kernels for the operand device by @kahyunnam in #4507
  • chore: add @jhjpark as CODEOWNER for GDN, Mamba, and KDA by @kahyunnam in #5247
  • fix(topk): fall back to CUB for graph-safe page-table transforms on SM120 by @bkryu in #5226
  • feat(moe): add cuTile MXFP4 and W4A16 support by @bkryu in #5099
  • perf(cake_gdn): optimize prefill kernels for SM100 and SM103 by @yyihuang in #5243
  • perf(cake_alpha_moe): accelerate AlphaMoE fused router on Blackwell by @yyihuang in #4339
  • fix(gdn): force min_blocks_per_mp=1 on SM107 to avoid DSL auto-carveout fault by @Vinnie6167 in #5227
  • perf(gdn): optimize SM100 CP kernels by @guangyunh-nv in #4917
  • chore: remove deprecated APIs eligible at the 0.7.0 boundary by @aleozlx in #5218
  • fix(moe): write fused-shared routing replay as routed-only [T, K] by @feih-nv in #4894
  • refactor(attention): use fixed page tables for paged block-sparse attention by @heyuhhh in #5144
  • feat(cake_comm): add optimized Blackwell MoE all-reduce fusion backend by @yyihuang in #4730
  • feat(decode): backend="cudnn" for BatchDecodeWithPagedKVCacheWrapper; cudnn decode honors q.dtype + base-2 LSE by @vedaanta in #4625
  • fix(moe): prune SM120 fused-MoE tile candidates that can never run by @yichengj0 in #4008
  • chore: clean up stale comments and vendoring notes by @bkryu in #5223
  • perf(prefill): build the cuDNN ragged graph once per length class via execute-time shape override; keep single-token GQA rows off cuDNN (NVBug 6783545) by @YangXu1990uiuc in #5245
  • fix: perf regression from missing trtllm moe tactics by @rosenrodt in #5207
  • [None][refactor] Mark PrimTS attention APIs as experimental by @yuxianq in #5082
  • test(autotuner): time both timers on the same run by @kahyunnam in #4188
  • BF16 CuTeDSL Dense GEMM: use runtime M and revert kernel disk caching by @jiahanc in #5231
  • fix(tests): skip cuDNN GDN kernel-oracle test when CUDA < 13 by @kahyunnam in #5246
  • fix(comm): accept MNNVL fabric UUIDs with leading zero bytes by @sshleifer in #5143
  • test(cudnn): evidence coverage for large-page and non-causal multi-token decode by @vedaanta in #4626
  • feat(gemm): enable mm_bf16 cute-dsl backend on SM107 (Rubin) by @gracehonv in #5222
  • Remove the SM100+ gate on the NVFP4 KV tile repack by @lesj0610 in #5239
  • feat(sparse): add SM120 Sage (QK-INT8/PV-FP8 and ALL FP8) block-sparse attention (b… by @hsr1234563 in #4691
  • feat(prefill): lse_base and lse_layout on the prefill wrappers (defaults unchanged; "ln" / "HN" for LSE-merge consumers) by @YangXu1990uiuc in #5257
  • feat(cake_comm): add generated Blackwell Ulysses all-to-all kernels by @yyihuang in #5263
  • fix(topk): add missing __syncthreads before AdvanceRadixGroupBarrier by @henrylhtsang in #5251
  • ci: link nightly releases to source commit by @dierksen in #5277
  • Sm90 megamoe optimization by @inocsin in #4688
  • test(kda): cover dense-state auto fallback for PR5037 by @cindyzxq in #5136
  • feat(cake_sparse_mla): support SM100 and SM103 by @yyihuang in #4573
  • Enable PDL for intermediate quantization by @wookjeHan in #5281
  • fix(kda): force min_blocks_per_mp=1 on SM107 in WY output-only decode by @kahyunnam in #5253
  • fix(tests): skip tests/prims_ts instead of erroring when CUTLASS DSL < 4.7 by @aleozlx in #5215
  • fix(attention): size prefill partial outputs by query rows, not packed indices by @gf239 in #5177
  • fix(cli): use uv for kernel wheel installs in uv environments by @mgoin in #5299
  • [Bugfix] Consume CPU seq-len mirrors on the cute-dsl ragged prefill backend by @zhang-keliang in #4816
  • feat(cake_gqa): add experimental prepared SM110 GQA decode by @yyihuang in #5302
  • perf(quantization): SM107 dispatch tuning and 128x4 tile kernels for nvfp4/mxfp4/mxfp8 by @bkryu in #5300
  • feat(moe_ep): MXFP8 x BF16 integration by @djns99 in #4604
  • feat(attention): add skip-softmax to SM120 PRIMS by @lishunyang12 in #4859
  • feat(prims-ts): extend QToken-KvBlock-Sparse-Attention shapes and patterns by @PerkzZheng in #5229
  • feat(MSA): Add NVFP4 paged-KV MSA sparse decode and prefill for SM100/SM103 by @nv-yunzheq in #4982
  • perf: reduce split W4A16 routing and output overhead by @zianglih in #5186
  • refactor(sparse-mla): unify SM120 execution, calibration and DSv4.1 support by @lucifer1004 in #5197
  • feat(moe): CUTLASS FP8 per-tensor runner takes the TRT-LLM static-scale activation pack by @feih-nv in #5294
  • feat(attention): add causal + bidirectional-ranges batch prefill by @lesj0610 in #5189
  • docs(moe): state the canonical [up, gate] row order of the gated GEMM1 weight (w13) by @feih-nv in #5291
  • feat(moe): add more activation types to the SM90 CuTe-DSL MoE backend by @Aneureka in #5295
  • fix(tests): skip cuDNN KDA kernel-oracle arm where the oracle has no SM107 by @kahyunnam in #5298
  • perf(norm): tune CuTe-DSL qk_rmsnorm and fused_add_rmsnorm_quant tiling for SM100/103/107 by @bkryu in #5305
  • perf(cute-dsl): skip inactive pages in windowed decode by @kzos in #4825
  • feat(cake_gdn): add opt-in SM100/SM103 context-parallel prefill backend by @yyihuang in #5320
  • feat(cudnn): decode backend forwards q_len_per_req > 1, sliding window and attention sinks by @vedaanta in #5327
  • perf(topk_varlen, attn_scores): upstream sync — gvr_2 hint-free engines (TRT-LLM #18410) + paged-MQA relu epilogue (DKG !27837) by @dhiraj113 in #5288
  • feat(moe): CUTLASS unified runners consume the TRT-LLM canonical activation packs by @feih-nv in #5230
  • feat(cake_nvfp4_attn): add experimental NVFP4 attention on SM103 by @yyihuang in #5283
  • feat(cake_kda): add TF32 and BF16 prefill kernels for SM100a and SM103a by @yyihuang in #5278
  • perf(moe-a2a): reuse dispatch receive views instead of rebuilding them per call by @Anerudhan in #5336
  • perf(mega_moe): SM90 fp8 megamoe optimization by @inocsin in #5338
  • feat(moe_ep): Support MiniMax-M3 MegaMoE by @jiahanc in #5360
  • perf(cake_kda): route one-wave BF16 grids to the M64 value split by @yyihuang in #5363
  • fix: guard K-tail loads in Blackwell gather GEMM by @murphymatt in #5108
  • fix(attention): materialize block_valid_mask for padded CUDA-graph prefill launches by @gf239 in #5176
  • feat(comm): add fused decode CP A2A + LSE reduce by @kwen2501 in #4929
  • feat(mla): uniform LSE base semantics via return_lse_base_on_e by @3xela in #4650
  • refactor(moe_ep): move the SM100 vendor package under sm100 by @akaashrp in #5371
  • perf(cake_kda): regenerate the BF16 KDA modules from the current Cake revision by @yyihuang in #5370
  • feat(prims-ts): support QK-BF16/PV-FP8 in FMHA decode by @harrisonzhy in #5150
  • refactor(kda): unify the two recurrent_kda facades onto one decode dispatcher by @kahyunnam in #5248
  • feat(cake_fmha): add BF16 head-dim-256 small-M speculative decode route by @yyihuang in #5369
  • ci(docs): never deploy a release candidate to the stable docs site by @aleozlx in #5383
  • Test Pruning Infrastructure to Separate Regular/Full Test Suites by @righthandabacus in #5380
  • Update CODEOWNERS for communication directories by @aleozlx in #5425
  • fix(gdn): honor Q/K normalization in GDN prefill by @akaashrp in #5255
  • test(gdn): prune duplicate kernel specializations in test_decode_delta_rule by @kahyunnam in #5379
  • feat(moe): relax the alignment constraints of the SM90 CuTe-DSL MoE by @Aneureka in #5372
  • feat(cake_gdn): add Qwen3.5 TP2 seven-token decode rows for SM100 and SM103 by @yyihuang in #5378
  • fix(cute-dsl): use has_side_effects to prevent wrong llvm optimisation by @namonakimono in #5125
  • perf(attention): use Tensor.max() for the longest lengths in plan() by @gf239 in #5043
  • ci: reuse cubin caches after key changes by @dierksen in #5273
  • docs(gemm): use Vera Rubin NVL72 nomenclature by @dierksen in #5457
  • fix(gemm): clamping sfa load row index for blackwell fp8 grouped gemm by @jimmyzho in #5387
  • ci: diagnose stalled JIT-cache provider builds by @dierksen in #5274
  • perf(moe): feed cuTile BF16 GEMM2 from an expert-sorted buffer at large M by @elwhyjay in #5129
  • feat(cake_sage): add SM100 and SM103 Sage-FP8 block-sparse attention with fused quantizer by @yyihuang in #5442
  • feat(moe_ep): CUDA-graph support for the split layer (nccl_ep + nixl_ep) by @Anerudhan in #5287
  • test: prune top-k parameter matrix by @righthandabacus in #5429
  • feat(moe): add SM12x (SM120/SM121) fp8 + mxfp8_mxfp4 grouped-MoE kernels by @CarstyYou in #4720
  • test: prune sampling parameter matrix by @righthandabacus in #5428
  • feat(moe): add cuTile FP8 and MXFP8 precision support by @bkryu in #5332
  • ci: require PR testing for .txt changes (requirements.txt, version.txt) by @bkryu in #5469
  • feat(comm): add H5120 MNNVL CuTe DSL profiles (TP4/TP8) and an apply_rms_norm switch by @zyongye in #5178
  • Harden JIT build watchdog cleanup by @dierksen in #5466
  • test(comm): cover the moe-a2a receive-view cache, and key it on coerced FFI ints by @aleozlx in #5406
  • feat(cake_sampling): fused radix top-k / top-p sampling pipeline for SM100/SM103 (single source; SM90/SM107/SM110 compile targets) by @yyihuang in #5439
  • fix(test): defer CUDA seed initialization by @dierksen in #5454
  • test: prune norm parameter matrices by @righthandabacus in #5427
  • feat(cake_gqa): experimental on-device load-balanced BF16 paged GQA decode (balanced_gqa_decode) on SM100/SM103 by @yyihuang in #5474
  • [None][refactor] Unify PrimTS paged decode entry points by @yuxianq in #5084
  • feat(moe_ep): add SM107 block-scaled MegaMoE backends by @akaashrp in #5384
  • feat(cake_softmax): add optimized Blackwell softmax by @yyihuang in #4282
  • perf(cake_sampling): qualify the radix sampling pipeline on SM90 (H100) and SM107 (Rubin R200); select JIT targets through CompilationContext by @yyihuang in #5482
  • feat(kda): add fused speculative decode kernel by @djmmoss in #4809
  • feat(kda): add CuTe DSL small-BH KDA prefill backend for Blackwell by @qiangxu1996 in #5032
  • feat: add is_deterministic mode to top_k_renorm_probs by @zpeng-xai in #4831
  • bench: sm107 whitelist and updates by @jimmyzho in #4981
  • ci: intersect targeted A10G/T4 scopes with their bare-run shard lists by @bkryu in #5476
  • docs: add self-review skill and fix contributor entry links by @bkryu in #5501
  • feat(cake_diffusion): add MiniMax-H3 fused RMSNorm + AdaLN + FC1 + SwiGLU for SM100 and SM103 by @yyihuang in #5491
  • [moe_ep] Batch NVFP4 epilogue staging copies by @ZenAlexa in #5132
  • test: prune attention sink parameter matrix by @righthandabacus in #5407
  • test: prune batch attention parameter matrix by @righthandabacus in #5409
  • chore: add jimmyzho to gemm codeowners by @jimmyzho in #5505
  • feat(cake_mega_moe): add CuTe DSL backend for experimental MXFP8 MegaMoE EP16 by @hzfan in #5433
  • feat(cake_minimax_h3): SM120 FP8/NVFP4 fused MiniMax-H3 pre-attention (RTX 5090 / RTX PRO 6000) by @yyihuang in #5493
  • Native-layout NVFP4 W4A16 in CuTe DSL on SM12x by @stecasta in #5242
  • perf(cake_megamoe_topk_reduce): regenerate the SM100a reducer from the current Cake exporter by @yyihuang in #5431
  • feat(cake_backend): add MiniMax-H3 NVFP4 (W4A4) pre-attention for SM100a/SM103a by @yyihuang in #5496
  • perf(cake_diffusion): 256-column quantized GEMM tiles for the MiniMax-H3 fused FC1 + SwiGLU on SM100 and SM103 by @yyihuang in #5506
  • feat(cake_minimax_h3): experimental MiniMax-H3 packed-varlen BF16 + NVFP4 attention (SM100/SM103) by @yyihuang in #5499
  • feat(cake_comm): add generated Blackwell DCP all-to-all kernels by @yyihuang in #5512
  • feat(kda): support bf16 state in packed fused KDA decode for T>1 by @djmmoss in #5503
  • perf(cake_backend): packed-row MTP fast path for the experimental balanced_gqa_decode (q_len_per_req 3..8) on SM100/SM103 by @yyihuang in #5490
  • refactor(tvm_ffi_utils): add named tensor-argument check helpers by @yzh119 in #5408
  • feat(cake_nvfp4_mla_decode): add experimental NVFP4 DeepSeek-V4 decode attention on SM100/SM103 by @yyihuang in #5443
  • Reuse cuDNN attention plans and fix metadata lifetime and layout by @YangXu1990uiuc in #5350
  • test(cake_sampling): pin the nvcc sm_107 probe in the build-target test by @aleozlx in #5511
  • feat(cake_megamoe_topk_reduce): publish the SM103a reducer export and enable the native reducer on B300 by @yyihuang in #5509
  • feat(mla): MLA SM100 auto mode by @saltyminty in #5463
  • feat(cake_fmha): MiniMax-H3 dense BF16 self-attention for SM120 (Cake kernel route + benchmark) by @yyihuang in #5472
  • feat(cake_fused_norm_combine): SM100 eight-peer fused residual-add + RMSNorm + BF16 combine by @yyihuang in #5528
  • test: prune batch prefill parameter matrices by @righthandabacus in #5411
  • test: prune batch decode parameter matrix by @righthandabacus in #5410
  • test: prune XQA parameter matrix by @righthandabacus in #5420
  • test: prune TensorRT-LLM decode parameter matrices by @righthandabacus in #5418
  • perf(cake_diffusion): 256x448 GEMM tile for the MiniMax-H3 FC1+SwiGLU NVFP4 operator (SM100a/SM103a) by @yyihuang in #5513
  • feat(cake_backend): MiniMax-H3 one-pass QKV quantize-and-pack helpers (NVFP4 + MXFP8, SM100a/SM103a) by @yyihuang in #5527
  • test: prune DeepSeek MLA parameter matrix by @righthandabacus in #5412
  • test: prune sliding-window parameter matrices by @righthandabacus in #5417
  • test: prune logits-cap parameter matrix by @righthandabacus in #5415
  • feat(cake_minimax_h3): SM120 FP8/NVFP4 fused MiniMax-H3 attention output projection with indexed gate + residual (RTX 5090 / RTX PRO 6000) by @yyihuang in #5524
  • fix(cake_kda): regenerate the BF16 KDA prefill modules with tile-anchored unbounded-gate decay (SM100/SM103) by @yyihuang in #5440
  • test: prune Hopper prefill parameter matrix by @righthandabacus in #5414
  • test: prune FMHA v2 prefill parameter matrices by @righthandabacus in #5413
  • feat(cake_kda): prepared BF16 KDA prefill plan cache, FP32 intermediate states and cached affine split for unbounded FP32-state prefill (SM100/SM103) by @yyihuang in #5452
  • feat(cake_minimax_h3): SM120 FP8 packed-varlen attention for MiniMax-H3 (RTX 5090 / RTX PRO 6000) by @yyihuang in #5539
  • feat(cake_backend): Kimi-K3 fused MoE router (896 experts, top-16) generated programs for SM100/SM103 by @yyihuang in #5531
  • feat(cake_diffusion): MiniMax-H3 SM120 fused quantized FC1 + SwiGLU (RTX 5090 / RTX PRO 6000) by @yyihuang in #5521
  • feat(cake_mla): add generated Blackwell TRT-LLM MLA decode backend for SM100 and SM103 by @yyihuang in #4557
  • perf(cake_mla): split the FP8 page-64 producer o_done handshake per V stage (SM100 + SM103) by @yyihuang in #5547
  • perf(cake_backend): drop the redundant device fence before the grid join in the Kimi-K3 fused router (SM100/SM103) by @yyihuang in #5548
  • feat(cake_diffusion): add MiniMax-H3 direct-layout out-proj + gated residual for SM100 and SM103 by @yyihuang in #5525
  • feat(cake_kda): dense-beta unbounded coverage, measured affine break-even, sm_103a constants by @yyihuang in #5543
  • fix(cake_warp_decode): extend NVFP4 warp-decode coverage and synchronization by @yyihuang in #5134
  • add-cudnn-frost-as-moe-backend by @yanqinz2 in #5249
  • feat(cake_kernel): epilogue-fused NVFP4 QKV GEMM for MiniMax-H3 pre-attention (SM100a/SM103a) by @yyihuang in #5529
  • feat(cake_comm): export the SM100 world-size-4 MoE all-reduce fusion union by @yyihuang in #5514
  • feat(cake_kda): Kimi K3 TP12 H=8 fused KDA decode in the Cake backend on SM100/SM103 by @yyihuang in #5551
  • feat(cake_vision_backend): Kimi-K3 vision tower Cake backend (SM100a/SM103a) with generated programs by @yyihuang in #5554
  • feat(cake_gemm): prepared contiguous grouped FP8 GEMM program family for SM100a by @yyihuang in #5500
  • perf(cake_diffusion): MiniMax-H3 out-proj NVFP4 quantization pass 14x2 loads in flight, raster group 8, 64-column TMEM drain (SM100/SM103) by @yyihuang in #5562
  • perf(cake_kernel): unified 16-warp epilogue role for the epilogue-fused NVFP4 QKV GEMM on SM103a (MiniMax-H3 pre-attention) by @yyihuang in #5563
  • feat(moe): wire dynamic FP8 FC2 quantization for per-channel MoE by @raayandhar in #5149
  • perf(cake_minimax_h3): K/V-split the partial-wave units of the packed-varlen attention (SM100/SM103) by @yyihuang in #5530
  • perf(cake_backend): warp-per-row selection arm for the largest Kimi-K3 fused router batches (SM100/SM103) by @yyihuang in #5564
  • perf(cake_kda): faster fused M128 BF16 prefill body (checkpoint rows off the chain, packed gate scan), page64 phase fix by @yyihuang in #5565
  • perf(cake_minimax_h3): fp8-PV split-program schedule fix and sm_103a split-cost recalibration (SM100/SM103) by @yyihuang in #5571
  • perf(cake_warp_decode): route the SM103 Qwen3-30B-A3B num_tokens=8 warp-decode row onto the fused route-pack K256 path by @yyihuang in #5574
  • feat(cake_vsa_sm90): SM90 BF16 variable block-sparse attention backend for VariableBlockSparseAttentionWrapper by @yyihuang in #5569
  • perf(cake_vision_backend): Kimi-K3 vision tower Cake backend round 2 (SM100a/SM103a): packed epilogues, f16x2 RoPE table, PDL by @yyihuang in #5570
  • feat(cake_backend): Kimi-K3 FP8_PB_WO KDA/MLA projection GEMMs (SM100/SM103) by @yyihuang in #5572
  • feat(cake_fused_norm_combine): add SM103 (B300) eight-peer fused norm-combine beside SM100 and a pipelined persistent owner reduce for 1024+ tokens by @yyihuang in #5540
  • perf(cake_backend): split cooperative joins and a warp-level round-0 shortcut for the mid and large Kimi-K3 fused router batches (SM100/SM103) by @yyihuang in #5568
  • chore(cake_kernel): regenerate the epilogue-fused NVFP4 QKV GEMM export for MiniMax-H3 pre-attention (SM100a/SM103a) by @yyihuang in #5566
  • perf(cake_kda): fused apply route for the unbounded affine BF16 prefill composite (operator export, pair-map prefix, fused apply kernel), H12 beta export fix by @yyihuang in #5573
  • feat(cake_backend): Kimi-K3 Stable LatentMoE front/tail projections (SM100a/SM103a) with generated programs by @yyihuang in #5575
  • feat(cake_msa): NVFP4 paged-KV MSA decode program for SM100/SM103 + layout contract by @yyihuang in #5402
  • feat(cake_kimi_k3_mla): Cake-generated Kimi-K3 MLA FP8 paged attention backend for SM100/SM103 (backend="cake") by @yyihuang in #5552
  • feat(cake_gemm): fused grouped FP8 gate_up GEMM + SwiGLU + FP8 quantization for SM100a by @yyihuang in #5567
  • feat(cake_nvfp4_attn): MiniMax-H3 SM120 ragged NVFP4 attention with the SageAttention3 FP4 recipe (experimental, #4532 candidate 5090-K2) by @yyihuang in #5545
  • perf(cake_vsa_sm90): split-KV small-selection route for sub-occupancy grids by @yyihuang in #5581
  • feat(cake_xqa): add experimental SM110 XQA attention (tcgen05 and register-MMA) by @yyihuang in #5293
  • feat(cake_mla_varq_dcp_decode): Cake MLA variable-query decode with DCP for SM100 / SM103 by @yyihuang in #5577
  • perf(cake_nvfp4_attn): MiniMax-H3 SM120 ragged NVFP4 attention (SageAttention3 recipe): block-max-first softmax, 4 partial row sums by @yyihuang in #5583
  • perf(cake_vsa_sm90): cluster-merge small route for uniform sliced selections by @yyihuang in #5582
  • Per-token alpha for NVFP4 mm_fp4 (CuTe-DSL SM100/SM103) and out_scale fold in per-token quantize by @aleozlx in #5504
  • feat(moe_ep): Support Kimi K3 SiTU in NVFP4 MegaMoE by @wzhao18 in #5455
  • bench(cake_mla_varq_dcp_decode): drop the package README and add a query-pattern benchmark option (second-round attribution, no module change) by @yyihuang in #5588
  • perf(cake_warp_decode): evict-first weight streams for the sm_103a Qwen3-30B-A3B fused route-pack rows (num_tokens 8-32) by @yyihuang in #5593
  • perf(cake_xqa): add the split-KV register-MMA FP16 page128 route to experimental SM110 XQA by @yyihuang in #5597
  • perf(cake_vsa_sm90): two-warpgroup small kernels, deferred cluster rendezvous, persistent first-tile header by @yyihuang in #5586
  • perf(cake_sampling): Rubin R200 (212 SM) single-wave capacity table and re-qualification by @yyihuang in #5585
  • fix(cake_sparse_mla): harden the DSv4 sparse-MLA Cake backend for #4671 (padded Q, separate SWA/compressed tables, caller-owned workspace, length offsets) and regenerate the SM100/SM103 programs by @yyihuang in #5591
  • perf(cake_kda): shorter unbounded prepare chain for the BF16 prefill programs (sm_100a, sm_103a), correction buffer only on the chain route by @yyihuang in #5598
  • feat(cake_latent_moe): Kimi-K3 TP12 fused LatentMoE tail for GB200 / GB300 NVL72, kimi_k3_tp12_tail) by @yyihuang in #5603
  • perf(cake_mm_fp4): faster NVFP4 per-token path (register-resident per-token quantize, per-tile alpha epilogue, low-M split-K, L2 eviction policies) by @yyihuang in #5609
  • perf(cake_sampling): round-3 kernels (candidate-list stage 1, composite bitonic stage 2/3) and re-fitted dispatch by @yyihuang in #5607
  • feat(cake_fused_moe): add Kimi-K3 W4A8 MXFP4 SiTU routed MoE for SM100/SM103 by @yyihuang in #5430
  • perf(cake_fp8_projection): Kimi-K3 KDA/MLA projection GEMM (per-token FP8 activations x per-block FP8 weights, SM100/SM103) round 2: TMA-store epilogue and fused-quant decode route for large-M single-N-tile layers by @yyihuang in #5590
  • feat(cake_comm): extend the MoE all-reduce fusion union export to world sizes 2, 4 and 8 on SM100 and SM103 by @yyihuang in #5599
  • perf(cake_fp8_projection): Kimi-K3 KDA/MLA projection GEMM (per-token FP8 activations x per-block FP8 weights, SM100/SM103) round 3: narrow quantization units and a decoupled BF16 token ring in the fused-quant decode kernel by @yyihuang in #5604
  • feat(cake_alpha_moe): add Blackwell AlphaMoE NVFP4 expert up/down compute by @yyihuang in #4340
  • perf(cake_mla): kimi_k3 shape, regenerate both architectures for the r52 kernel (deferred O rescale, wider long-request threshold) by @yyihuang in #5602
  • perf(cake_sparse_mla): warp-uniform FP8/H128 TMA gathers and 32-byte epilogue stores in the DSv4 sparse-MLA Cake programs (SM100/SM103) by @yyihuang in #5610
  • perf(cake_nvfp4_attn): MiniMax-H3 SM120 NVFP4 varlen attention: fused statistics + quantization pre-processing for mid-size plans by @yyihuang in #5595
  • feat(cake_msa): nvfp4_decode, compute capability 10.7 programs under a three-architecture export lock by @yyihuang in #5594
  • perf(cake_kimi_k3_latent_moe): fused-norm single-launch prefill tail for single-wave rows + 6-deep ring instance for the TP1 tail at T <= 256 (SM100a/SM103a), regenerated programs by @yyihuang in #5584
  • feat(cake_vsa_sm90): CuTe DSL build of the Hopper VSA kernels (backend="cake_cute") by @yyihuang in #5596
  • perf(cake_kda): scan-owned state epilogue on the apply route, apply-route descriptor re-encode after a plan-cache rebind (round 4, sm_100a, sm_103a) by @yyihuang in #5622
  • perf(cake_kimi_k3_fused_router): Kimi-K3 router register-prefetch L arm on the 32-128-token shapes it wins, per-architecture route tables, regenerated programs by @yyihuang in #5621
  • perf(cake_mla): kimi_k3, regenerate both architectures for the r54 kernel (redundant p_full barrier removed) by @yyihuang in #5627
  • feat(cake_gemm): odd-tail kernel, routing-aware route rule and wide activation kernel for the fused grouped FP8 gate_up + SwiGLU + quant program (SM100a) by @yyihuang in #5592
  • perf(cake_fp8_projection): Kimi-K3 KDA/MLA projection GEMM (per-token FP8 activations x per-block FP8 weights, SM100/SM103) round 4: balanced single-split 128-CTA dispatch for the M=16384 fused-decode bucket by @yyihuang in #5626
  • perf(cake_latent_moe): Kimi-K3 TP12 fused LatentMoE tail round 4: persistent K3 pipeline from 256 tokens, weight-streaming K2 for M <= 4 (GB200 / GB300 NVL72) by @yyihuang in #5624
  • perf(cake_vision_backend): round 3 — attention score ring, GEMM tile boundaries, TMA-out epilogue, pointer TMA ABI (sm_100a/sm_103a) by @yyihuang in #5623
  • ci: add CUDA 13.4 CI images by @dierksen in #5456
  • feat(moe): add CuTe DSL NVFP4 W4A16 MegaMoE by @zianglih in #5019
  • perf(cake_sparse_mla): round-3 DSv4 sparse-MLA programs for SM100/SM103 (persistent-item epilogue, index ring, retained-KV decode body, SwapsAb V alias, routing) by @yyihuang in #5630
  • optimize-cudnn-frost-moe-decoding by @yanqinz2 in #5628
  • perf(cake_nvfp4_attn): MiniMax-H3 SM120 NVFP4 varlen attention round 9: attention-kernel ceiling study on RTX 5090 / RTX PRO 6000 (kernel unchanged, route docs) by @yyihuang in #5629
  • perf(cake_kimi_k3_latent_moe): packed global-y norm, landing-zone alias, deferred cluster wait and a 128-wide TP8 tail tile (SM100a/SM103a), regenerated programs by @yyihuang in #5647
  • [moe_ep] Fix EXPERT_MAJOR unnecessary computation of padded elements by @x41lakazam in #5265
  • perf(cake_kda): round 5 — grouped TMEM loads with the state decay off the ready chain, strided q/k/v read in place, ping-pong prefix chain by @yyihuang in #5641
  • perf(cake_latent_moe): Kimi-K3 TP12 fused LatentMoE tail round 5: early shared scatter K1, fused K2+K3 for M <= 4, bulk reduce-scatter K3 pipeline above 256 tokens (GB200 / GB300 NVL72) by @yyihuang in #5646
  • perf(cake_comm): faster SM100/SM103 MoE all-reduce fusion union programs by @yyihuang in #5650
  • perf(cake_sparse_mla): round-4 DSv4 sparse-MLA Cake programs for SM100/SM103 (stacked on #5630) by @yyihuang in #5644
  • perf(cake_fp8_projection): Kimi-K3 KDA/MLA projection GEMM (per-token FP8 activations x per-block FP8 weights, SM100/SM103) round 5: cluster split-K decode routes and staged register GEMM epilogue by @yyihuang in #5642
  • test(topk_varlen): skip the radix_filter all-backends case when the installed DSL cannot run it by @dhiraj113 in #5656
  • perf(cake_xqa): add the tcgen05/TMEM cluster-multicast D512 tree routes to experimental SM110 XQA by @yyihuang in #5658
  • perf(cake_sampling): round-4 kernels (streaming candidate-list stage 1, fused small-k tail, host-decided PDL trigger) and k-aware dispatch by @yyihuang in #5636
  • perf(cake_vsa_sm90): round-8 planner load model, odd-tile store guard and acquire-release split merge on the CuTe route by @yyihuang in #5638
  • feat(comm): add PCIe IPC all-gather and reduce-scatter by @yilin-void in #5023
  • perf(cake_vision_backend): round 4 — stream-K tail for the N=1024 projector GEMMs, per-arch attention split cost model (sm_100a/sm_103a) by @yyihuang in #5643
  • ci: use released sccache v0.18.0 for CUDA 13.4 builds by @dierksen in #5653
  • feat(cake_gemm): paired BF16 conversions and a mixed-schedule route for the fused grouped FP8 gate_up + SwiGLU + quant program (SM100a) by @yyihuang in #5645
  • perf(cake_latent_moe): Kimi-K3 TP12 fused LatentMoE tail round 6: fused K23 up to eight tokens (capacity ladder), fp64-gated numerics by @yyihuang in #5668
  • test: prune TensorRT-LLM XQA parameter matrix by @righthandabacus in #5419
  • test: prune XQA batch-decode parameter matrix by @righthandabacus in #5421
  • perf(cake_xqa): per-GQA-ratio tmem kernels and the cta_group::2 pair routes for experimental SM110 XQA by @yyihuang in #5674
  • perf(cake_sparse_mla): round-5 DSv4 sparse-MLA Cake programs for SM100/SM103 (exact numerics; stacked on #5644) by @yyihuang in #5686
  • perf(prims_ts): mask partial last K/V tile outside softmax loop to prevent register spills in dense context fmha by @harrisonzhy in #5280
  • perf: specialize persistent BatchAttention for equal KV strides by @saltyminty in #5329
  • [prims-ts] Mixed precision Fmha decode by @IwakuraRein in #4414
  • fix(aot): ship SM100-family modules in the SM103 JIT-cache provider by @mmangkad in #5544
  • docs: resolve blocking documentation checks by @kangbintNV in #5483
  • ci: add @yyihuang to every CODEOWNERS rule by @yyihuang in #5654
  • feat(moe): add opt-in Prims-TS backend to the unified MoE API by @feih-nv in #5289
  • feat(quantization): expose NVFP4 4over6 recipe through the public API by @aleozlx in #5152
  • perf(cake_vsa_sm90): queue-scheduled persistent stage for long uniform plans by @yyihuang in #5687
  • perf(cake_gemm): evict_first weight stream and descriptor prefetch for the fused grouped FP8 gate_up + SwiGLU + quant program by @yyihuang in #5692
  • perf(jit): disable pre-RA instruction scheduling for the trtllm-gen MoE manifest by @taylor-yb-lee in #5435
  • perf(cake_vision_backend): round 5 — PDL prologue overlap (PDL_EARLY) census windows, m_e8 / m_e8_cs tiles, fc1 stream-K twins (sm_100a/sm_103a) by @yyihuang in #5691
  • perf(cake_kimi_k3_latent_moe): round 8 -- single-wave aligned K halves for the front prefill trailing wave, evict_first front instance for T <= 512, coalesced stream-K partial slots (SM100a/SM103a) by @yyihuang in #5671
  • fix(moe): forward enable_pdl to CuTeDSL MoE routing by @suiyoubi in #5497
  • perf(cake_comm): converge the SM100/SM103 MoE all-reduce fusion union programs (world sizes 2/4/8) by @yyihuang in #5694
  • feat(cake_fused_moe): add Kimi-K3 NVFP4 SiTU experts for SM100/SM103 by @yyihuang in #5183
  • fix(comm): support Kimi K3 shapes in MNNVL HT allreduce_fusion by @syuoni in #5632
  • feat(cake_deepgemm): add generated DeepGEMM-family kernels for SM100a/SM103a (sparse MQA indexer, routing gate, mHC, MoE, FP8/FP4 GEMM) by @yyihuang in #5523
  • feat(moe): expert parallelism for the SM12x W4A16 fused MoE by @yichengj0 in #4302
  • feat(moe_bgmv): direct decode fast-path for shrink (1.3-3.8x vs sliced kernel) by @aws-jiadingg in #3535
  • perf(moe_bgmv): coalesced 128-bit loads for expand kernel (1.8–3× at batch≥256) by @aws-jiadingg in #3542
  • ci: add a retry window with capped backoff to cubin downloads by @aleozlx with @Copilot in #5256
  • bump version to 0.7.1 by @jimmyzho in #5711

New Contributors

  • @Wint3rNight made their first contribution in #5013
  • @Victor49152 made their first contribution in #4856
  • @yilin-void made their first contribution in #5024
  • @Stelath made their first contribution in #5088
  • @SSHdotCodes made their first contribution in #4847
  • @samodi-nv made their first contribution in #4734
  • @klshuster made their first contribution in #5172
  • @ZenAlexa made their first contribution in #4920
  • @JacobHelwig made their first contribution in #5198
  • @XFDG made their first contribution in #5171
  • @harrisonzhy made their first contribution in #4879
  • @zhougit86 made their first contribution in #3633
  • @sshleifer made their first contribution in #5143
  • @gracehonv made their first contribution in #5222
  • @henrylhtsang made their first contribution in #5251
  • @inocsin made their first contribution in #4688
  • @wookjeHan made their first contribution in #5281
  • @gf239 made their first contribution in #5177
  • @lishunyang12 made their first contribution in #4859
  • @kzos made their first contribution in #4825
  • @3xela made their first contribution in #4650
  • @akaashrp made their first contribution in #5371
  • @namonakimono made their first contribution in #5125
  • @zyongye made their first contribution in #5178
  • @qiangxu1996 made their first contribution in #5032
  • @zpeng-xai made their first contribution in #4831
  • @stecasta made their first contribution in #5242
  • @mmangkad made their first contribution in #5544
  • @suiyoubi made their first contribution in #5497

Full Changelog: v0.7.0rc4...v0.7.1rc1

Don't miss a new flashinfer release

NewReleases is sending notifications on new releases.