github flashinfer-ai/flashinfer v0.7.1
Release v0.7.1

4 hours ago

These highlights are also published at flashinfer.ai/releases.

v0.7.1 Highlights

FlashInfer v0.7.1 enables quantized expert-parallel MoE on Rubin and Blackwell SM12x, adds fused Kimi K3 speculative verification, and adds cuDNN linear-attention prefill on Blackwell. Sparse MLA gains DeepSeek-V4.1 dual-cache support and compact GLM53-NoPE caches, alongside MLA API improvements and per-model NVFP4 quantization recipes.

Kimi K3 adds packed speculative decoding and NVFP4 SiTU experts

Kimi K3 integrations can use packed_fused_kda_decode on Blackwell SM10x to fuse short convolution, Kimi Delta Attention, and gated RMS normalization during speculative verification. The API supports ragged batches, accepted-token cache updates, CUDA Graph capture, and FP32 or BF16 recurrent state. NVFP4 MegaMoE also supports Kimi K3's SiTU activation on Blackwell SM10x.

  • #4809 feat(kda): add fused speculative decode kernel
  • #5503 feat(kda): support bf16 state in packed fused KDA decode for T>1
  • #5455 feat(moe_ep): Support Kimi K3 SiTU in NVFP4 MegaMoE

Run quantized expert-parallel MoE on Rubin and Blackwell SM12x

Rubin SM107 supports NVFP4 and MXFP8 MegaMoE through MoEEpLayer, with BF16 combine and output. Blackwell SM120/SM121 W4A16 fused MoE supports expert sharding through expert_map, with the caller distributing the full batch and summing per-rank outputs. The split MoE layer supports CUDA Graph capture with both NCCL-EP and NIXL-EP through persistent graph state.

  • #5384 feat(moe_ep): add SM107 block-scaled MegaMoE backends
  • #4302 feat(moe): expert parallelism for the SM12x W4A16 fused MoE
  • #5287 feat(moe_ep): CUDA-graph support for the split layer (nccl_ep + nixl_ep)

cuDNN runs Kimi Delta Attention and Gated DeltaNet prefill on Blackwell

Kimi Delta Attention and Gated DeltaNet prefill run on cuDNN's open-source fused linear-attention engine on Blackwell SM100, selected with recurrent_kda(backend="cudnn") and chunk_gated_delta_rule(backend="cudnn"). The cuDNN path also runs Gated DeltaNet prefill on Blackwell CUDA 12 builds. New chunk_gated_delta_product and chunk_gated_delta_rule2 entry points add Gated Delta Product and Gated Delta Rule 2 prefill on the same engine.

  • #4968 Add cuDNN backend for linear attention

Call MLA prefill directly and reuse execution plans

trtllm_prefill_with_kv_cache_mla, available through both flashinfer.mla and flashinfer.prefill, gives integrations a named prefill entry point with the existing MLA decode API's signature and behavior. Multi-token prefill callers must explicitly select a compatible backend; XQA supports one query token. BatchMLAPagedAttentionWrapper prepares TRTLLM-GEN, XQA, and monolithic or modular CuTe DSL backends in plan() and reuses them in run(), giving integrations a shared lifecycle for repeated execution and CUDA Graph capture. On SM100, callers can pass backend="auto" and let plan() choose among all available backends for the planned shapes and features.

  • #5158 fix: add discoverable MLA prefill API
  • #5070 feat(mla): add planned TRTLLM-GEN, XQA, and CuTe DSL backends
  • #5463 feat(mla): MLA SM100 auto mode

Run DeepSeek-V4.1 dual-cache attention and shrink GLM53-NoPE caches on SM12x

DeepSeek-V4.1 sparse MLA runs on SM120/SM121 with a primary FP8 cache and an optional FP4 or FP8 extra cache, independent page sizes and top-k selections, and FP8 or BF16 compute. GLM53-NoPE supports compact 528-byte cache rows, reducing per-token KV storage by 19.5% for callers switching from 656-byte rows. Existing padded 656-byte pools remain supported through strided 3D/4D cache views.

  • #5197 refactor(sparse-mla): unify SM120 execution, calibration and DSv4.1 support
  • #5075 feat(sm120): runtime KV row stride + canonical 528B GLM53_NOPE rows + NaN-safe masked gathers in sparse MLA

NVFP4 recipes can be configured per model

NVFP4 quantization APIs accept an explicit NVFP44Over6Config, allowing models in one process to use different 4over6 recipes. The unified MoE QuantConfig carries the same setting for supporting runners, including the CuTe DSL per-token path. Backend support is checked when an explicit recipe is requested.

  • #5152 feat(quantization): expose NVFP4 4over6 recipe through the public API

Upgrade notices

  • Install nvidia-cudnn-frontend>=1.30.0. For the cuDNN linear-attention backend, include the cutedsl extra: pip install -U 'nvidia-cudnn-frontend[cutedsl]>=1.30.0'.
  • Update calls to fp8_paged_mqa_logits and fp4_paged_mqa_logits: use block_tables, seq_lens, and max_seq_len; pass the page table before sequence lengths in positional calls; remove num_epi_subtiles. FP4 callers must also rename sf_q to q_sf.
  • Update direct SM90 CuTe DSL MoE integrations for the current signatures: cute_dsl_fused_moe_bf16 uses tactic for tuning, and CuteDslBf16MoEWrapper has a changed constructor argument order. Pass output_dtype, enable_pdl, and use_fused_finalize by keyword and remove legacy tile arguments.
  • Pass recv_view_cache by keyword to moe_a2a_dispatch.
  • Review PrimTS integrations against the experimental API contract. Migrate callers of the removed prims_ts_batch_decode_with_kv_cache and prims_ts_batch_mla_decode_with_kv_cache entry points to the corresponding paged decode wrappers.

What's Changed

  • feat(comm): copy-engine based kernel for the PCIe IPC all-reduce by @qsang-nv in #4870
  • feat(cake_warp_decode): Add SM100 NVFP4 warp decode and standalone SiLU by @yyihuang in #5036
  • feat(cake_nvfp4_svdquant): add NVFP4 SVDQuant GEMM SM100 and SM103 backend by @yyihuang in #5076
  • fix(moe): bound routing_replay_out dim0 from below in both validators by @aleozlx in #5072
  • fix: guard BF16 warp Split-K availability and prune tactics by @jiahanc in #5089
  • refactor(kda): remove the unreleased recurrent KDA training/backward family by @YangXu1990uiuc in #4965
  • fix(activation): pick a vector width that divides hidden_size in act_and_mul (fixes d=3420) by @Wint3rNight in #5013
  • docs: document MiniMax-H3 prepared run parameters by @kangbintNV in #5087
  • perf(moe): improve SM90 mixed-input CUTLASS MoE backend performance by @StudyingShao in #5005
  • refactor(kda): stop 0.7 advertising unreleased KDA surface (mark prefill wrapper experimental, trim kda_kernels.all) by @kahyunnam in #5040
  • bench: compare DeepSeek MoE against pure TRTLLM BF16 by @zianglih in #4985
  • perf(topk_varlen): gvr_2 dispatch rungs for 4K-8K rows, SM-count-aware register band, hint-free gvr_2 in auto by @dhiraj113 in #4986
  • feat(cake_mega_moe): add SM100 BF16 rank-major MegaMoE backend by @yyihuang in #4810
  • fix(attention): default skip_all_rows_active_check to true by @saltyminty in #5039
  • feat(cake_fmha): add sm100/103 paged-attention by @yyihuang in #4980
  • fix: preserve both generated-source pre-commit exclusions by @saltyminty in #5113
  • fix(fp4_gemm autotune) SM107 kernel family excluded from autotune candidates and heuristics update by @Victor49152 in #4856
  • fix(mla): halve sparse-MLA cpb L2 guard-rail window on integrated GPUs (Spark) by @bkryu in #5097
  • fix(cake_sage_attention): support non-contiguous Sage attention with partial KV tiles by @yyihuang in #5083
  • Fix MoE Tile Mapping Zeroing by @Vinnie6167 in #4735
  • fix(test): pin the GDN disk-cache round-trip test to the flashinfer backend by @bkryu in #5104
  • feat(cake_kda): add fused decode backend for B200 and B300 by @yyihuang in #4967
  • feat(gemm): cuTile alpha-beta / masked-bmm / ragged-bmm / block-scaled FP8 GEMM by @yifeis-nv in #4020
  • feat(cake_comm): add opt-in Blackwell MNNVL MoE all-to-all backend by @yyihuang in #4784
  • feat(cake_diffusion): add SM103a BF16 pre-attention by @yyihuang in #4690
  • fix(moe): Fix per-token quantization fast-math INF scale when activation contains all-zero row by @xuantengh in #5031
  • Fix bmm_fp8 cublasLt handle usage in autotuned cublas runner #26381 by @baonudesifeizhai in #2808
  • build: raise the nvidia-cudnn-frontend floor to 1.28.0 by @Anerudhan in #5131
  • fix(moe): avoid FP8 scale overflow in the Humming MXFP4 path by @yilin-void in #5024
  • chore: add codebase-wise approvers to every CODEOWNERS rule by @bkryu in #5146
  • Add CUTLASS Primitives and Task Scheduling Blackwell MoE kernels by @nekorobov in #4361
  • fix[mla]: add Q8 floor for tile selection in trtllm-gen sparse mla kernels by @jimmyzho in #5106
  • feat(quantization): support fp32 input in the CuTe-DSL mxfp8_quantize kernel by @bkryu in #5112
  • feat(sm120): runtime KV row stride + canonical 528B GLM53_NOPE rows + NaN-safe masked gathers in sparse MLA by @lucifer1004 in #5075
  • Revert "build: raise the nvidia-cudnn-frontend floor to 1.28.0" by @aleozlx in #5159
  • feat(cake_diffusion): add SM100a support to MiniMax-H3 BF16 pre-attention by @yyihuang in #5137
  • fix(moe): accept mxfp8's flat linear activation-scale layout in the MoE autotuner by @aleozlx in #4943
  • fix(comm): make TorchDistBackend.Split() synchronizing, like MPI_Comm_split by @aleozlx in #4206
  • fix(sampling): preserve softmax normalization at low temperatures by @Stelath in #5088
  • fix(infra): isolate distributed rendezvous ports by test worker by @yongwww in #5153
  • feat(cake_bf16_fp4_gemm): add SM100/SM103 BF16 x FP4 GEMM kernels by @yyihuang in #5138
  • feat(cake_mega_moe): add reusable MegaMoE workspaces and an SM100a reducer by @yyihuang in #4819
  • perf(cake_mega_moe): optimize BF16 x MXFP8 MegaMoE EP16 backend by @hzfan in #5148
  • Add cuDNN backend for linear attention by @jhjpark in #4968
  • feat(moe): remove QuantVariant and dispatch QuantConfig by MMA pair by @feih-nv in #5061
  • fix(cake_sage_attn): cover Sage combinations and pass tensor maps by value by @yyihuang in #5127
  • fix(cutile): stabilize autotune config selection in select GEMMs by @elwhyjay in #4459
  • test(comm): report every rank's fate when a dist worker dies by @aleozlx in #4205
  • fix[GEMM]: WAR for specific cuTile autotune candidate in mm_bf16 by @jimmyzho in #5155
  • ci: apply --prerelease to rc in release workflow by @aleozlx in #4472
  • fix(moe): validate TMA warp-specialized configs in queryOccupancyForConfig by @SSHdotCodes in #4847
  • test(cake_comm): initialize NVSHMEM for BF16 rank-major GPU regression by @yyihuang in #5126
  • feat(cake_comm): add Cake MoE finalize all-reduce backend by @yyihuang in #4875
  • feat(cake_alpha_moe): add optimized Blackwell W8A8 expert up/down compute by @xslingcn in #4287
  • perf(cake_comm): reduce PCIe CE ring publication overhead on SM120 by @yyihuang in #5169
  • feat(prims-ts): add non-absorbed MLA prefill by @PerkzZheng in #5059
  • BF16 cutedsl prune [2] + cache bf16 cutedsl kernels by @jiahanc in #5140
  • docs: fix 0.7.0 documentation check findings by @kangbintNV in #5199
  • build: raise the nvidia-cudnn-frontend floor to 1.29.0 by @Anerudhan in #5182
  • fix(cake_attention): enable DCP on SM107 and skip unsupported VSA tests by @saltyminty in #5111
  • feat(cake_comm): support SM100 and SM103 TP8 packed-QKV all-gather matmul by @yyihuang in #5116
  • fix: add discoverable MLA prefill API by @saltyminty in #5158
  • feat(moe): tag the unified MoE API as official with @flashinfer_api by @aleozlx in #5200
  • perf(attention): specialize paged FA2 equal-stride KV by @saltyminty in #4736
  • feat(comm): add allocation-stable head-chunk Ulysses primitives by @tiffany940107 in #5027
  • test(kda): xfail CuTe DSL recurrent prefill tests on nvidia-cutlass-dsl<4.7 by @bkryu in #5219
  • Paged-MQA logits: Rubin (SM107) enablement, FP4 next_n=4, and DKG epilogue/scheduler optimizations by @dhiraj113 in #4737
  • feat(gemm): add SM10x contiguous grouped FP8 CuTe backend by @samodi-nv in #4734
  • feat: enable large-head attention on SM80, for VLLM to run FP8 kv by @jhaotingc in #5044
  • feat(cake_bf16_bmm): add optimized SM100a and SM103a backend by @yyihuang in #4370
  • fix(comm): fence DCP receiver unpack before async output copy by @klshuster in #5172
  • feat(prims-ts): add QToken-KvBlock-Sparse-Attention by @PerkzZheng in #4996
  • ci(jit-cache): enable provider-only wheel releases by @dierksen in #5096
  • fix(autotuner): recover from builtin memory errors by @ZenAlexa in #4920
  • feat(mla): add planned TRTLLM-GEN, XQA, and CuTe DSL backends by @saltyminty in #5070
  • fix(topk): handle padded short rows in CUB transforms by @zianglih in #5187
  • fix(comm): match the library filename when scanning /proc/self/maps by @JacobHelwig in #5198
  • feat: add low-latency Frost DSv4.1 sparse MLA decode by @YangXu1990uiuc in #5119
  • fix(ci): support installed BF16 rank-major session tests by @XFDG in #5171
  • fix: restore default GDN prefill dispatch to CuTe by @yyihuang in #5225
  • refactor(moe): update and refactor the SM90 CuTe-DSL MoE tactics by @Aneureka in #5004
  • feat(prefill): let backend="auto" reach cuDNN/CUTLASS on Blackwell ragged prefill (up to 3.3x) by @Anerudhan in #5133
  • feat(prims-ts): support QK-BF16/PV-FP8 in FMHA context by @harrisonzhy in #4879
  • ci: cache checksum-verified cubin downloads by @dierksen in #5240
  • fix(jit): refresh stale arch detection and allow SM120 on CUDA 12.8 by @zhougit86 in #3633
  • fix(gdn): compile and tune CuTe-DSL decode kernels for the operand device by @kahyunnam in #4507
  • chore: add @jhjpark as CODEOWNER for GDN, Mamba, and KDA by @kahyunnam in #5247
  • fix(topk): fall back to CUB for graph-safe page-table transforms on SM120 by @bkryu in #5226
  • feat(moe): add cuTile MXFP4 and W4A16 support by @bkryu in #5099
  • perf(cake_gdn): optimize prefill kernels for SM100 and SM103 by @yyihuang in #5243
  • perf(cake_alpha_moe): accelerate AlphaMoE fused router on Blackwell by @yyihuang in #4339
  • fix(gdn): force min_blocks_per_mp=1 on SM107 to avoid DSL auto-carveout fault by @Vinnie6167 in #5227
  • perf(gdn): optimize SM100 CP kernels by @guangyunh-nv in #4917
  • chore: remove deprecated APIs eligible at the 0.7.0 boundary by @aleozlx in #5218
  • fix(moe): write fused-shared routing replay as routed-only [T, K] by @feih-nv in #4894
  • refactor(attention): use fixed page tables for paged block-sparse attention by @heyuhhh in #5144
  • feat(cake_comm): add optimized Blackwell MoE all-reduce fusion backend by @yyihuang in #4730
  • feat(decode): backend="cudnn" for BatchDecodeWithPagedKVCacheWrapper; cudnn decode honors q.dtype + base-2 LSE by @vedaanta in #4625
  • fix(moe): prune SM120 fused-MoE tile candidates that can never run by @yichengj0 in #4008
  • chore: clean up stale comments and vendoring notes by @bkryu in #5223
  • perf(prefill): build the cuDNN ragged graph once per length class via execute-time shape override; keep single-token GQA rows off cuDNN (NVBug 6783545) by @YangXu1990uiuc in #5245
  • fix: perf regression from missing trtllm moe tactics by @rosenrodt in #5207
  • [None][refactor] Mark PrimTS attention APIs as experimental by @yuxianq in #5082
  • test(autotuner): time both timers on the same run by @kahyunnam in #4188
  • BF16 CuTeDSL Dense GEMM: use runtime M and revert kernel disk caching by @jiahanc in #5231
  • fix(tests): skip cuDNN GDN kernel-oracle test when CUDA < 13 by @kahyunnam in #5246
  • fix(comm): accept MNNVL fabric UUIDs with leading zero bytes by @sshleifer in #5143
  • test(cudnn): evidence coverage for large-page and non-causal multi-token decode by @vedaanta in #4626
  • feat(gemm): enable mm_bf16 cute-dsl backend on SM107 (Rubin) by @gracehonv in #5222
  • Remove the SM100+ gate on the NVFP4 KV tile repack by @lesj0610 in #5239
  • feat(sparse): add SM120 Sage (QK-INT8/PV-FP8 and ALL FP8) block-sparse attention (b… by @hsr1234563 in #4691
  • feat(prefill): lse_base and lse_layout on the prefill wrappers (defaults unchanged; "ln" / "HN" for LSE-merge consumers) by @YangXu1990uiuc in #5257
  • feat(cake_comm): add generated Blackwell Ulysses all-to-all kernels by @yyihuang in #5263
  • fix(topk): add missing __syncthreads before AdvanceRadixGroupBarrier by @henrylhtsang in #5251
  • ci: link nightly releases to source commit by @dierksen in #5277
  • Sm90 megamoe optimization by @inocsin in #4688
  • test(kda): cover dense-state auto fallback for PR5037 by @cindyzxq in #5136
  • feat(cake_sparse_mla): support SM100 and SM103 by @yyihuang in #4573
  • Enable PDL for intermediate quantization by @wookjeHan in #5281
  • fix(kda): force min_blocks_per_mp=1 on SM107 in WY output-only decode by @kahyunnam in #5253
  • fix(tests): skip tests/prims_ts instead of erroring when CUTLASS DSL < 4.7 by @aleozlx in #5215
  • fix(attention): size prefill partial outputs by query rows, not packed indices by @gf239 in #5177
  • fix(cli): use uv for kernel wheel installs in uv environments by @mgoin in #5299
  • [Bugfix] Consume CPU seq-len mirrors on the cute-dsl ragged prefill backend by @zhang-keliang in #4816
  • feat(cake_gqa): add experimental prepared SM110 GQA decode by @yyihuang in #5302
  • perf(quantization): SM107 dispatch tuning and 128x4 tile kernels for nvfp4/mxfp4/mxfp8 by @bkryu in #5300
  • feat(moe_ep): MXFP8 x BF16 integration by @djns99 in #4604
  • feat(attention): add skip-softmax to SM120 PRIMS by @lishunyang12 in #4859
  • feat(prims-ts): extend QToken-KvBlock-Sparse-Attention shapes and patterns by @PerkzZheng in #5229
  • feat(MSA): Add NVFP4 paged-KV MSA sparse decode and prefill for SM100/SM103 by @nv-yunzheq in #4982
  • perf: reduce split W4A16 routing and output overhead by @zianglih in #5186
  • refactor(sparse-mla): unify SM120 execution, calibration and DSv4.1 support by @lucifer1004 in #5197
  • feat(moe): CUTLASS FP8 per-tensor runner takes the TRT-LLM static-scale activation pack by @feih-nv in #5294
  • feat(attention): add causal + bidirectional-ranges batch prefill by @lesj0610 in #5189
  • docs(moe): state the canonical [up, gate] row order of the gated GEMM1 weight (w13) by @feih-nv in #5291
  • feat(moe): add more activation types to the SM90 CuTe-DSL MoE backend by @Aneureka in #5295
  • fix(tests): skip cuDNN KDA kernel-oracle arm where the oracle has no SM107 by @kahyunnam in #5298
  • perf(norm): tune CuTe-DSL qk_rmsnorm and fused_add_rmsnorm_quant tiling for SM100/103/107 by @bkryu in #5305
  • perf(cute-dsl): skip inactive pages in windowed decode by @kzos in #4825
  • feat(cake_gdn): add opt-in SM100/SM103 context-parallel prefill backend by @yyihuang in #5320
  • feat(cudnn): decode backend forwards q_len_per_req > 1, sliding window and attention sinks by @vedaanta in #5327
  • perf(topk_varlen, attn_scores): upstream sync — gvr_2 hint-free engines (TRT-LLM #18410) + paged-MQA relu epilogue (DKG !27837) by @dhiraj113 in #5288
  • feat(moe): CUTLASS unified runners consume the TRT-LLM canonical activation packs by @feih-nv in #5230
  • feat(cake_nvfp4_attn): add experimental NVFP4 attention on SM103 by @yyihuang in #5283
  • feat(cake_kda): add TF32 and BF16 prefill kernels for SM100a and SM103a by @yyihuang in #5278
  • perf(moe-a2a): reuse dispatch receive views instead of rebuilding them per call by @Anerudhan in #5336
  • perf(mega_moe): SM90 fp8 megamoe optimization by @inocsin in #5338
  • feat(moe_ep): Support MiniMax-M3 MegaMoE by @jiahanc in #5360
  • perf(cake_kda): route one-wave BF16 grids to the M64 value split by @yyihuang in #5363
  • fix: guard K-tail loads in Blackwell gather GEMM by @murphymatt in #5108
  • fix(attention): materialize block_valid_mask for padded CUDA-graph prefill launches by @gf239 in #5176
  • feat(comm): add fused decode CP A2A + LSE reduce by @kwen2501 in #4929
  • feat(mla): uniform LSE base semantics via return_lse_base_on_e by @3xela in #4650
  • refactor(moe_ep): move the SM100 vendor package under sm100 by @akaashrp in #5371
  • perf(cake_kda): regenerate the BF16 KDA modules from the current Cake revision by @yyihuang in #5370
  • feat(prims-ts): support QK-BF16/PV-FP8 in FMHA decode by @harrisonzhy in #5150
  • refactor(kda): unify the two recurrent_kda facades onto one decode dispatcher by @kahyunnam in #5248
  • feat(cake_fmha): add BF16 head-dim-256 small-M speculative decode route by @yyihuang in #5369
  • ci(docs): never deploy a release candidate to the stable docs site by @aleozlx in #5383
  • Test Pruning Infrastructure to Separate Regular/Full Test Suites by @righthandabacus in #5380
  • Update CODEOWNERS for communication directories by @aleozlx in #5425
  • fix(gdn): honor Q/K normalization in GDN prefill by @akaashrp in #5255
  • test(gdn): prune duplicate kernel specializations in test_decode_delta_rule by @kahyunnam in #5379
  • feat(moe): relax the alignment constraints of the SM90 CuTe-DSL MoE by @Aneureka in #5372
  • feat(cake_gdn): add Qwen3.5 TP2 seven-token decode rows for SM100 and SM103 by @yyihuang in #5378
  • fix(cute-dsl): use has_side_effects to prevent wrong llvm optimisation by @namonakimono in #5125
  • perf(attention): use Tensor.max() for the longest lengths in plan() by @gf239 in #5043
  • ci: reuse cubin caches after key changes by @dierksen in #5273
  • docs(gemm): use Vera Rubin NVL72 nomenclature by @dierksen in #5457
  • fix(gemm): clamping sfa load row index for blackwell fp8 grouped gemm by @jimmyzho in #5387
  • ci: diagnose stalled JIT-cache provider builds by @dierksen in #5274
  • perf(moe): feed cuTile BF16 GEMM2 from an expert-sorted buffer at large M by @elwhyjay in #5129
  • feat(cake_sage): add SM100 and SM103 Sage-FP8 block-sparse attention with fused quantizer by @yyihuang in #5442
  • feat(moe_ep): CUDA-graph support for the split layer (nccl_ep + nixl_ep) by @Anerudhan in #5287
  • test: prune top-k parameter matrix by @righthandabacus in #5429
  • feat(moe): add SM12x (SM120/SM121) fp8 + mxfp8_mxfp4 grouped-MoE kernels by @CarstyYou in #4720
  • test: prune sampling parameter matrix by @righthandabacus in #5428
  • feat(moe): add cuTile FP8 and MXFP8 precision support by @bkryu in #5332
  • ci: require PR testing for .txt changes (requirements.txt, version.txt) by @bkryu in #5469
  • feat(comm): add H5120 MNNVL CuTe DSL profiles (TP4/TP8) and an apply_rms_norm switch by @zyongye in #5178
  • Harden JIT build watchdog cleanup by @dierksen in #5466
  • test(comm): cover the moe-a2a receive-view cache, and key it on coerced FFI ints by @aleozlx in #5406
  • feat(cake_sampling): fused radix top-k / top-p sampling pipeline for SM100/SM103 (single source; SM90/SM107/SM110 compile targets) by @yyihuang in #5439
  • fix(test): defer CUDA seed initialization by @dierksen in #5454
  • test: prune norm parameter matrices by @righthandabacus in #5427
  • feat(cake_gqa): experimental on-device load-balanced BF16 paged GQA decode (balanced_gqa_decode) on SM100/SM103 by @yyihuang in #5474
  • [None][refactor] Unify PrimTS paged decode entry points by @yuxianq in #5084
  • feat(moe_ep): add SM107 block-scaled MegaMoE backends by @akaashrp in #5384
  • feat(cake_softmax): add optimized Blackwell softmax by @yyihuang in #4282
  • perf(cake_sampling): qualify the radix sampling pipeline on SM90 (H100) and SM107 (Rubin R200); select JIT targets through CompilationContext by @yyihuang in #5482
  • feat(kda): add fused speculative decode kernel by @djmmoss in #4809
  • feat(kda): add CuTe DSL small-BH KDA prefill backend for Blackwell by @qiangxu1996 in #5032
  • feat: add is_deterministic mode to top_k_renorm_probs by @zpeng-xai in #4831
  • bench: sm107 whitelist and updates by @jimmyzho in #4981
  • ci: intersect targeted A10G/T4 scopes with their bare-run shard lists by @bkryu in #5476
  • docs: add self-review skill and fix contributor entry links by @bkryu in #5501
  • feat(cake_diffusion): add MiniMax-H3 fused RMSNorm + AdaLN + FC1 + SwiGLU for SM100 and SM103 by @yyihuang in #5491
  • [moe_ep] Batch NVFP4 epilogue staging copies by @ZenAlexa in #5132
  • test: prune attention sink parameter matrix by @righthandabacus in #5407
  • test: prune batch attention parameter matrix by @righthandabacus in #5409
  • chore: add jimmyzho to gemm codeowners by @jimmyzho in #5505
  • feat(cake_mega_moe): add CuTe DSL backend for experimental MXFP8 MegaMoE EP16 by @hzfan in #5433
  • feat(cake_minimax_h3): SM120 FP8/NVFP4 fused MiniMax-H3 pre-attention (RTX 5090 / RTX PRO 6000) by @yyihuang in #5493
  • Native-layout NVFP4 W4A16 in CuTe DSL on SM12x by @stecasta in #5242
  • perf(cake_megamoe_topk_reduce): regenerate the SM100a reducer from the current Cake exporter by @yyihuang in #5431
  • feat(cake_backend): add MiniMax-H3 NVFP4 (W4A4) pre-attention for SM100a/SM103a by @yyihuang in #5496
  • perf(cake_diffusion): 256-column quantized GEMM tiles for the MiniMax-H3 fused FC1 + SwiGLU on SM100 and SM103 by @yyihuang in #5506
  • feat(cake_minimax_h3): experimental MiniMax-H3 packed-varlen BF16 + NVFP4 attention (SM100/SM103) by @yyihuang in #5499
  • feat(cake_comm): add generated Blackwell DCP all-to-all kernels by @yyihuang in #5512
  • feat(kda): support bf16 state in packed fused KDA decode for T>1 by @djmmoss in #5503
  • perf(cake_backend): packed-row MTP fast path for the experimental balanced_gqa_decode (q_len_per_req 3..8) on SM100/SM103 by @yyihuang in #5490
  • refactor(tvm_ffi_utils): add named tensor-argument check helpers by @yzh119 in #5408
  • feat(cake_nvfp4_mla_decode): add experimental NVFP4 DeepSeek-V4 decode attention on SM100/SM103 by @yyihuang in #5443
  • Reuse cuDNN attention plans and fix metadata lifetime and layout by @YangXu1990uiuc in #5350
  • test(cake_sampling): pin the nvcc sm_107 probe in the build-target test by @aleozlx in #5511
  • feat(cake_megamoe_topk_reduce): publish the SM103a reducer export and enable the native reducer on B300 by @yyihuang in #5509
  • feat(mla): MLA SM100 auto mode by @saltyminty in #5463
  • feat(cake_fmha): MiniMax-H3 dense BF16 self-attention for SM120 (Cake kernel route + benchmark) by @yyihuang in #5472
  • feat(cake_fused_norm_combine): SM100 eight-peer fused residual-add + RMSNorm + BF16 combine by @yyihuang in #5528
  • test: prune batch prefill parameter matrices by @righthandabacus in #5411
  • test: prune batch decode parameter matrix by @righthandabacus in #5410
  • test: prune XQA parameter matrix by @righthandabacus in #5420
  • test: prune TensorRT-LLM decode parameter matrices by @righthandabacus in #5418
  • perf(cake_diffusion): 256x448 GEMM tile for the MiniMax-H3 FC1+SwiGLU NVFP4 operator (SM100a/SM103a) by @yyihuang in #5513
  • feat(cake_backend): MiniMax-H3 one-pass QKV quantize-and-pack helpers (NVFP4 + MXFP8, SM100a/SM103a) by @yyihuang in #5527
  • test: prune DeepSeek MLA parameter matrix by @righthandabacus in #5412
  • test: prune sliding-window parameter matrices by @righthandabacus in #5417
  • test: prune logits-cap parameter matrix by @righthandabacus in #5415
  • feat(cake_minimax_h3): SM120 FP8/NVFP4 fused MiniMax-H3 attention output projection with indexed gate + residual (RTX 5090 / RTX PRO 6000) by @yyihuang in #5524
  • fix(cake_kda): regenerate the BF16 KDA prefill modules with tile-anchored unbounded-gate decay (SM100/SM103) by @yyihuang in #5440
  • test: prune Hopper prefill parameter matrix by @righthandabacus in #5414
  • test: prune FMHA v2 prefill parameter matrices by @righthandabacus in #5413
  • feat(cake_kda): prepared BF16 KDA prefill plan cache, FP32 intermediate states and cached affine split for unbounded FP32-state prefill (SM100/SM103) by @yyihuang in #5452
  • feat(cake_minimax_h3): SM120 FP8 packed-varlen attention for MiniMax-H3 (RTX 5090 / RTX PRO 6000) by @yyihuang in #5539
  • feat(cake_backend): Kimi-K3 fused MoE router (896 experts, top-16) generated programs for SM100/SM103 by @yyihuang in #5531
  • feat(cake_diffusion): MiniMax-H3 SM120 fused quantized FC1 + SwiGLU (RTX 5090 / RTX PRO 6000) by @yyihuang in #5521
  • feat(cake_mla): add generated Blackwell TRT-LLM MLA decode backend for SM100 and SM103 by @yyihuang in #4557
  • perf(cake_mla): split the FP8 page-64 producer o_done handshake per V stage (SM100 + SM103) by @yyihuang in #5547
  • perf(cake_backend): drop the redundant device fence before the grid join in the Kimi-K3 fused router (SM100/SM103) by @yyihuang in #5548
  • feat(cake_diffusion): add MiniMax-H3 direct-layout out-proj + gated residual for SM100 and SM103 by @yyihuang in #5525
  • feat(cake_kda): dense-beta unbounded coverage, measured affine break-even, sm_103a constants by @yyihuang in #5543
  • fix(cake_warp_decode): extend NVFP4 warp-decode coverage and synchronization by @yyihuang in #5134
  • add-cudnn-frost-as-moe-backend by @yanqinz2 in #5249
  • feat(cake_kernel): epilogue-fused NVFP4 QKV GEMM for MiniMax-H3 pre-attention (SM100a/SM103a) by @yyihuang in #5529
  • feat(cake_comm): export the SM100 world-size-4 MoE all-reduce fusion union by @yyihuang in #5514
  • feat(cake_kda): Kimi K3 TP12 H=8 fused KDA decode in the Cake backend on SM100/SM103 by @yyihuang in #5551
  • feat(cake_vision_backend): Kimi-K3 vision tower Cake backend (SM100a/SM103a) with generated programs by @yyihuang in #5554
  • feat(cake_gemm): prepared contiguous grouped FP8 GEMM program family for SM100a by @yyihuang in #5500
  • perf(cake_diffusion): MiniMax-H3 out-proj NVFP4 quantization pass 14x2 loads in flight, raster group 8, 64-column TMEM drain (SM100/SM103) by @yyihuang in #5562
  • perf(cake_kernel): unified 16-warp epilogue role for the epilogue-fused NVFP4 QKV GEMM on SM103a (MiniMax-H3 pre-attention) by @yyihuang in #5563
  • feat(moe): wire dynamic FP8 FC2 quantization for per-channel MoE by @raayandhar in #5149
  • perf(cake_minimax_h3): K/V-split the partial-wave units of the packed-varlen attention (SM100/SM103) by @yyihuang in #5530
  • perf(cake_backend): warp-per-row selection arm for the largest Kimi-K3 fused router batches (SM100/SM103) by @yyihuang in #5564
  • perf(cake_kda): faster fused M128 BF16 prefill body (checkpoint rows off the chain, packed gate scan), page64 phase fix by @yyihuang in #5565
  • perf(cake_minimax_h3): fp8-PV split-program schedule fix and sm_103a split-cost recalibration (SM100/SM103) by @yyihuang in #5571
  • perf(cake_warp_decode): route the SM103 Qwen3-30B-A3B num_tokens=8 warp-decode row onto the fused route-pack K256 path by @yyihuang in #5574
  • feat(cake_vsa_sm90): SM90 BF16 variable block-sparse attention backend for VariableBlockSparseAttentionWrapper by @yyihuang in #5569
  • perf(cake_vision_backend): Kimi-K3 vision tower Cake backend round 2 (SM100a/SM103a): packed epilogues, f16x2 RoPE table, PDL by @yyihuang in #5570
  • feat(cake_backend): Kimi-K3 FP8_PB_WO KDA/MLA projection GEMMs (SM100/SM103) by @yyihuang in #5572
  • feat(cake_fused_norm_combine): add SM103 (B300) eight-peer fused norm-combine beside SM100 and a pipelined persistent owner reduce for 1024+ tokens by @yyihuang in #5540
  • perf(cake_backend): split cooperative joins and a warp-level round-0 shortcut for the mid and large Kimi-K3 fused router batches (SM100/SM103) by @yyihuang in #5568
  • chore(cake_kernel): regenerate the epilogue-fused NVFP4 QKV GEMM export for MiniMax-H3 pre-attention (SM100a/SM103a) by @yyihuang in #5566
  • perf(cake_kda): fused apply route for the unbounded affine BF16 prefill composite (operator export, pair-map prefix, fused apply kernel), H12 beta export fix by @yyihuang in #5573
  • feat(cake_backend): Kimi-K3 Stable LatentMoE front/tail projections (SM100a/SM103a) with generated programs by @yyihuang in #5575
  • feat(cake_msa): NVFP4 paged-KV MSA decode program for SM100/SM103 + layout contract by @yyihuang in #5402
  • feat(cake_kimi_k3_mla): Cake-generated Kimi-K3 MLA FP8 paged attention backend for SM100/SM103 (backend="cake") by @yyihuang in #5552
  • feat(cake_gemm): fused grouped FP8 gate_up GEMM + SwiGLU + FP8 quantization for SM100a by @yyihuang in #5567
  • feat(cake_nvfp4_attn): MiniMax-H3 SM120 ragged NVFP4 attention with the SageAttention3 FP4 recipe (experimental, #4532 candidate 5090-K2) by @yyihuang in #5545
  • perf(cake_vsa_sm90): split-KV small-selection route for sub-occupancy grids by @yyihuang in #5581
  • feat(cake_xqa): add experimental SM110 XQA attention (tcgen05 and register-MMA) by @yyihuang in #5293
  • feat(cake_mla_varq_dcp_decode): Cake MLA variable-query decode with DCP for SM100 / SM103 by @yyihuang in #5577
  • perf(cake_nvfp4_attn): MiniMax-H3 SM120 ragged NVFP4 attention (SageAttention3 recipe): block-max-first softmax, 4 partial row sums by @yyihuang in #5583
  • perf(cake_vsa_sm90): cluster-merge small route for uniform sliced selections by @yyihuang in #5582
  • Per-token alpha for NVFP4 mm_fp4 (CuTe-DSL SM100/SM103) and out_scale fold in per-token quantize by @aleozlx in #5504
  • feat(moe_ep): Support Kimi K3 SiTU in NVFP4 MegaMoE by @wzhao18 in #5455
  • bench(cake_mla_varq_dcp_decode): drop the package README and add a query-pattern benchmark option (second-round attribution, no module change) by @yyihuang in #5588
  • perf(cake_warp_decode): evict-first weight streams for the sm_103a Qwen3-30B-A3B fused route-pack rows (num_tokens 8-32) by @yyihuang in #5593
  • perf(cake_xqa): add the split-KV register-MMA FP16 page128 route to experimental SM110 XQA by @yyihuang in #5597
  • perf(cake_vsa_sm90): two-warpgroup small kernels, deferred cluster rendezvous, persistent first-tile header by @yyihuang in #5586
  • perf(cake_sampling): Rubin R200 (212 SM) single-wave capacity table and re-qualification by @yyihuang in #5585
  • fix(cake_sparse_mla): harden the DSv4 sparse-MLA Cake backend for #4671 (padded Q, separate SWA/compressed tables, caller-owned workspace, length offsets) and regenerate the SM100/SM103 programs by @yyihuang in #5591
  • perf(cake_kda): shorter unbounded prepare chain for the BF16 prefill programs (sm_100a, sm_103a), correction buffer only on the chain route by @yyihuang in #5598
  • feat(cake_latent_moe): Kimi-K3 TP12 fused LatentMoE tail for GB200 / GB300 NVL72, kimi_k3_tp12_tail) by @yyihuang in #5603
  • perf(cake_mm_fp4): faster NVFP4 per-token path (register-resident per-token quantize, per-tile alpha epilogue, low-M split-K, L2 eviction policies) by @yyihuang in #5609
  • perf(cake_sampling): round-3 kernels (candidate-list stage 1, composite bitonic stage 2/3) and re-fitted dispatch by @yyihuang in #5607
  • feat(cake_fused_moe): add Kimi-K3 W4A8 MXFP4 SiTU routed MoE for SM100/SM103 by @yyihuang in #5430
  • perf(cake_fp8_projection): Kimi-K3 KDA/MLA projection GEMM (per-token FP8 activations x per-block FP8 weights, SM100/SM103) round 2: TMA-store epilogue and fused-quant decode route for large-M single-N-tile layers by @yyihuang in #5590
  • feat(cake_comm): extend the MoE all-reduce fusion union export to world sizes 2, 4 and 8 on SM100 and SM103 by @yyihuang in #5599
  • perf(cake_fp8_projection): Kimi-K3 KDA/MLA projection GEMM (per-token FP8 activations x per-block FP8 weights, SM100/SM103) round 3: narrow quantization units and a decoupled BF16 token ring in the fused-quant decode kernel by @yyihuang in #5604
  • feat(cake_alpha_moe): add Blackwell AlphaMoE NVFP4 expert up/down compute by @yyihuang in #4340
  • perf(cake_mla): kimi_k3 shape, regenerate both architectures for the r52 kernel (deferred O rescale, wider long-request threshold) by @yyihuang in #5602
  • perf(cake_sparse_mla): warp-uniform FP8/H128 TMA gathers and 32-byte epilogue stores in the DSv4 sparse-MLA Cake programs (SM100/SM103) by @yyihuang in #5610
  • perf(cake_nvfp4_attn): MiniMax-H3 SM120 NVFP4 varlen attention: fused statistics + quantization pre-processing for mid-size plans by @yyihuang in #5595
  • feat(cake_msa): nvfp4_decode, compute capability 10.7 programs under a three-architecture export lock by @yyihuang in #5594
  • perf(cake_kimi_k3_latent_moe): fused-norm single-launch prefill tail for single-wave rows + 6-deep ring instance for the TP1 tail at T <= 256 (SM100a/SM103a), regenerated programs by @yyihuang in #5584
  • feat(cake_vsa_sm90): CuTe DSL build of the Hopper VSA kernels (backend="cake_cute") by @yyihuang in #5596
  • perf(cake_kda): scan-owned state epilogue on the apply route, apply-route descriptor re-encode after a plan-cache rebind (round 4, sm_100a, sm_103a) by @yyihuang in #5622
  • perf(cake_kimi_k3_fused_router): Kimi-K3 router register-prefetch L arm on the 32-128-token shapes it wins, per-architecture route tables, regenerated programs by @yyihuang in #5621
  • perf(cake_mla): kimi_k3, regenerate both architectures for the r54 kernel (redundant p_full barrier removed) by @yyihuang in #5627
  • feat(cake_gemm): odd-tail kernel, routing-aware route rule and wide activation kernel for the fused grouped FP8 gate_up + SwiGLU + quant program (SM100a) by @yyihuang in #5592
  • perf(cake_fp8_projection): Kimi-K3 KDA/MLA projection GEMM (per-token FP8 activations x per-block FP8 weights, SM100/SM103) round 4: balanced single-split 128-CTA dispatch for the M=16384 fused-decode bucket by @yyihuang in #5626
  • perf(cake_latent_moe): Kimi-K3 TP12 fused LatentMoE tail round 4: persistent K3 pipeline from 256 tokens, weight-streaming K2 for M <= 4 (GB200 / GB300 NVL72) by @yyihuang in #5624
  • perf(cake_vision_backend): round 3 — attention score ring, GEMM tile boundaries, TMA-out epilogue, pointer TMA ABI (sm_100a/sm_103a) by @yyihuang in #5623
  • ci: add CUDA 13.4 CI images by @dierksen in #5456
  • feat(moe): add CuTe DSL NVFP4 W4A16 MegaMoE by @zianglih in #5019
  • perf(cake_sparse_mla): round-3 DSv4 sparse-MLA programs for SM100/SM103 (persistent-item epilogue, index ring, retained-KV decode body, SwapsAb V alias, routing) by @yyihuang in #5630
  • optimize-cudnn-frost-moe-decoding by @yanqinz2 in #5628
  • perf(cake_nvfp4_attn): MiniMax-H3 SM120 NVFP4 varlen attention round 9: attention-kernel ceiling study on RTX 5090 / RTX PRO 6000 (kernel unchanged, route docs) by @yyihuang in #5629
  • perf(cake_kimi_k3_latent_moe): packed global-y norm, landing-zone alias, deferred cluster wait and a 128-wide TP8 tail tile (SM100a/SM103a), regenerated programs by @yyihuang in #5647
  • [moe_ep] Fix EXPERT_MAJOR unnecessary computation of padded elements by @x41lakazam in #5265
  • perf(cake_kda): round 5 — grouped TMEM loads with the state decay off the ready chain, strided q/k/v read in place, ping-pong prefix chain by @yyihuang in #5641
  • perf(cake_latent_moe): Kimi-K3 TP12 fused LatentMoE tail round 5: early shared scatter K1, fused K2+K3 for M <= 4, bulk reduce-scatter K3 pipeline above 256 tokens (GB200 / GB300 NVL72) by @yyihuang in #5646
  • perf(cake_comm): faster SM100/SM103 MoE all-reduce fusion union programs by @yyihuang in #5650
  • perf(cake_sparse_mla): round-4 DSv4 sparse-MLA Cake programs for SM100/SM103 (stacked on #5630) by @yyihuang in #5644
  • perf(cake_fp8_projection): Kimi-K3 KDA/MLA projection GEMM (per-token FP8 activations x per-block FP8 weights, SM100/SM103) round 5: cluster split-K decode routes and staged register GEMM epilogue by @yyihuang in #5642
  • test(topk_varlen): skip the radix_filter all-backends case when the installed DSL cannot run it by @dhiraj113 in #5656
  • perf(cake_xqa): add the tcgen05/TMEM cluster-multicast D512 tree routes to experimental SM110 XQA by @yyihuang in #5658
  • perf(cake_sampling): round-4 kernels (streaming candidate-list stage 1, fused small-k tail, host-decided PDL trigger) and k-aware dispatch by @yyihuang in #5636
  • perf(cake_vsa_sm90): round-8 planner load model, odd-tile store guard and acquire-release split merge on the CuTe route by @yyihuang in #5638
  • feat(comm): add PCIe IPC all-gather and reduce-scatter by @yilin-void in #5023
  • perf(cake_vision_backend): round 4 — stream-K tail for the N=1024 projector GEMMs, per-arch attention split cost model (sm_100a/sm_103a) by @yyihuang in #5643
  • ci: use released sccache v0.18.0 for CUDA 13.4 builds by @dierksen in #5653
  • feat(cake_gemm): paired BF16 conversions and a mixed-schedule route for the fused grouped FP8 gate_up + SwiGLU + quant program (SM100a) by @yyihuang in #5645
  • perf(cake_latent_moe): Kimi-K3 TP12 fused LatentMoE tail round 6: fused K23 up to eight tokens (capacity ladder), fp64-gated numerics by @yyihuang in #5668
  • test: prune TensorRT-LLM XQA parameter matrix by @righthandabacus in #5419
  • test: prune XQA batch-decode parameter matrix by @righthandabacus in #5421
  • perf(cake_xqa): per-GQA-ratio tmem kernels and the cta_group::2 pair routes for experimental SM110 XQA by @yyihuang in #5674
  • perf(cake_sparse_mla): round-5 DSv4 sparse-MLA Cake programs for SM100/SM103 (exact numerics; stacked on #5644) by @yyihuang in #5686
  • perf(prims_ts): mask partial last K/V tile outside softmax loop to prevent register spills in dense context fmha by @harrisonzhy in #5280
  • perf: specialize persistent BatchAttention for equal KV strides by @saltyminty in #5329
  • [prims-ts] Mixed precision Fmha decode by @IwakuraRein in #4414
  • fix(aot): ship SM100-family modules in the SM103 JIT-cache provider by @mmangkad in #5544
  • docs: resolve blocking documentation checks by @kangbintNV in #5483
  • ci: add @yyihuang to every CODEOWNERS rule by @yyihuang in #5654
  • feat(moe): add opt-in Prims-TS backend to the unified MoE API by @feih-nv in #5289
  • feat(quantization): expose NVFP4 4over6 recipe through the public API by @aleozlx in #5152
  • perf(cake_vsa_sm90): queue-scheduled persistent stage for long uniform plans by @yyihuang in #5687
  • perf(cake_gemm): evict_first weight stream and descriptor prefetch for the fused grouped FP8 gate_up + SwiGLU + quant program by @yyihuang in #5692
  • perf(jit): disable pre-RA instruction scheduling for the trtllm-gen MoE manifest by @taylor-yb-lee in #5435
  • perf(cake_vision_backend): round 5 — PDL prologue overlap (PDL_EARLY) census windows, m_e8 / m_e8_cs tiles, fc1 stream-K twins (sm_100a/sm_103a) by @yyihuang in #5691
  • perf(cake_kimi_k3_latent_moe): round 8 -- single-wave aligned K halves for the front prefill trailing wave, evict_first front instance for T <= 512, coalesced stream-K partial slots (SM100a/SM103a) by @yyihuang in #5671
  • fix(moe): forward enable_pdl to CuTeDSL MoE routing by @suiyoubi in #5497
  • perf(cake_comm): converge the SM100/SM103 MoE all-reduce fusion union programs (world sizes 2/4/8) by @yyihuang in #5694
  • feat(cake_fused_moe): add Kimi-K3 NVFP4 SiTU experts for SM100/SM103 by @yyihuang in #5183
  • fix(comm): support Kimi K3 shapes in MNNVL HT allreduce_fusion by @syuoni in #5632
  • feat(cake_deepgemm): add generated DeepGEMM-family kernels for SM100a/SM103a (sparse MQA indexer, routing gate, mHC, MoE, FP8/FP4 GEMM) by @yyihuang in #5523
  • feat(moe): expert parallelism for the SM12x W4A16 fused MoE by @yichengj0 in #4302
  • feat(moe_bgmv): direct decode fast-path for shrink (1.3-3.8x vs sliced kernel) by @aws-jiadingg in #3535
  • perf(moe_bgmv): coalesced 128-bit loads for expand kernel (1.8–3× at batch≥256) by @aws-jiadingg in #3542
  • ci: add a retry window with capped backoff to cubin downloads by @aleozlx with @Copilot in #5256
  • bump version to 0.7.1 by @jimmyzho in #5711
  • [release-v0.7.1] Revert "feat(cake_fmha): add BF16 head-dim-256 small-M speculative decode route (#5369)" — fixes intermittent B200 GPU hang (#5724) by @aleozlx in #5978

New Contributors

  • @Wint3rNight made their first contribution in #5013
  • @Victor49152 made their first contribution in #4856
  • @yilin-void made their first contribution in #5024
  • @Stelath made their first contribution in #5088
  • @SSHdotCodes made their first contribution in #4847
  • @samodi-nv made their first contribution in #4734
  • @klshuster made their first contribution in #5172
  • @ZenAlexa made their first contribution in #4920
  • @JacobHelwig made their first contribution in #5198
  • @XFDG made their first contribution in #5171
  • @zhougit86 made their first contribution in #3633
  • @sshleifer made their first contribution in #5143
  • @gracehonv made their first contribution in #5222
  • @henrylhtsang made their first contribution in #5251
  • @inocsin made their first contribution in #4688
  • @wookjeHan made their first contribution in #5281
  • @gf239 made their first contribution in #5177
  • @lishunyang12 made their first contribution in #4859
  • @kzos made their first contribution in #4825
  • @3xela made their first contribution in #4650
  • @namonakimono made their first contribution in #5125
  • @zyongye made their first contribution in #5178
  • @qiangxu1996 made their first contribution in #5032
  • @zpeng-xai made their first contribution in #4831
  • @mmangkad made their first contribution in #5544
  • @suiyoubi made their first contribution in #5497

Full Changelog: v0.7.0rc4...v0.7.1

Don't miss a new flashinfer release

NewReleases is sending notifications on new releases.