Highlights
-
Model Support
- Support Qwen3.5 checkpoints with global FP8 scales across model loading paths #19519
- Optimize Gemma4 vision rotary embeddings without a GEMM operation #19571
- Enable the CuTeDSL MoE backend for Kimi K3 NVFP4 SiTU #19003
- Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
- Extend MiniMax-M3 piecewise CUDA graphs through model and executor paths #19423
-
API
- Remove the AutoDeploy integration and its public entry points (BREAKING) #19028
- Remove the obsolete cpp_only build option from CMake (BREAKING) #19421
- Accept parallel_tool_calls and cap tool enum values in OpenAI requests #19459
- Treat explicit null optional fields as unset in the Responses API #19454
- Validate media data URIs and require opt-in for private URL fetches #19458
-
Feature
- Add an MLA-backed standalone DSpark drafter for Inferact and Kimi K3 #19040
- Enable Helix speculative verification for FP8, FP4 MLA, and DSpark groups #19273
- Add a non-MTP KDA decode kernel and dispatch path #18938
- Add an NVFP4 MLA residual switch for DeepSeek V4 #19305
- Support MNNVL all-reduce through the Ray orchestrator #18231
- Add opt-in cached KV token counts and per-rank routing logs #19375, #16323
- Add quiet PyExecutor hang diagnostics for stalled inference requests #19045
- Support backend-independent agent tools and a required-tool policy #19236
- Update scaffolding MCP transport and KV-cache truncation behavior #19055
- Support W4A16 NVFP4 group size 32 with explicit validation #19091
- Add Triton top-k combine for Marlin NVFP4 MoE #18999
- Fuse MLA context FP8 Q quantization into the absorb BMM epilogue #19082
- Enable PDL on additional decode edges and FlashInfer GDN prefill on SM120/121 #19446, #19376
- Defer fused-GELU MLP NVFP4 scale finalization until needed #18906
- Avoid KV-cache V2 prefix probes when context chunk budget is insufficient #19366
- Publish MoE all-to-all counters before dispatch and consolidate C++ MoE ops #19425, #19216
-
Fix
- Compute Cosmos3 rotary tables with broadcast multiplication to avoid K=1 matmul failures #19532
- Restore Qwen3.5 gather and scatter behavior after a performance regression #19578
- Preserve disaggregated KV transfer outcomes and synchronize receive destinations #19377, #19444
- Authenticate disaggregated subagent affinity and reject unbindable IPv6 listener addresses #19175, #19544
- Make disaggregated load-balancing routing global and repair cache-reuse adapter behavior #19399, #19480
- Extend startup and warmup timeouts for Blackwell pipeline-parallel tests #19402
- Keep one-sided MoE combine offsets and FP4 permute buffers correctly sized #19312, #19251
- Reject paged-context attention when no fused kernel exists #19353
- Share speculative CUDA-graph capture buffers and correct RNG corner cases #19350, #19505
- Correct DFlash buffer sizing, positions, and FlashInfer speculation gating #19466
- Guard PyExecutor model-engine access while handling responses #19559
- Route zero-token causal convolution through the channel-major kernel #18571
- Correct VBWS width tracking and terminal-step finalization #19407
- Drain guided-decoder host functions before forward calls that hold the GIL #19455
- Prevent out-of-bounds indices in vanilla top-p sampling #19496
- Reuse pending MNNVL split communicators across workspace construction retries #19101
-
Documentation
- Document GVR V2 self-sampling and multi-threshold exact top-k methods #19244
-
Test & Infra
- Upgrade public PyTorch to 2.14, Triton to 3.8, and C++ to 20 #19477, #19249
- Add Helix zero-KV, disaggregated closure, and speculative acceptance regression coverage #18995, #19478, #19176
- Expand Qwen3.8 coverage across GPU configurations #19530
- Disable context overlap for the DeepSeek V4 Pro disaggregated test configuration #19424
- Fix GPT-OSS Eagle3 disaggregated fixtures and DeepSeek V4 assertions #19531, #19554
- Correct DFlash budget, MPI FlashInfer, and MNNVL memory test fixtures #19582, #19254, #19509
- Set bounded retries for VisualGen ffmpeg installation in tests #19433
- Correct Wan 2.2 LPIPS thresholds and restore its accuracy test #19438
- Preserve CI logs while detecting GCP hosts and identify SLURM job origins #19345, #19327
- Key the legacy lint baseline on POSIX paths #17745
- Update telemetry code ownership for configuration review #19599
- Regenerate the LLM-args telemetry manifest for sparse-attention configuration #19493
- Remove deprecated GPT-OSS Hopper tests and obsolete TensorRT leftovers #19470, #19428
What's Changed
- [TRTLLM-16564][feat] MLA-backboned standalone DSpark drafter (Inferact/Kimi-K3-DSpark) by @dc3671 in #19040
- [None][fix] Responses API: treat explicit null as unset for optional fields by @brnguyen2 in #19454
- [None][perf] defer dynamic NVFP4 scale finalization for fused GELU MLPs by @daichu-nv in #18906
- [None][fix] Regenerate LLM-args golden manifest for sparse_attention_config.uses_spcompress by @brnguyen2 in #19493
- [None][feat] agent-flow: backend-independent tools and required-tool policy by @kaiyux in #19236
- [None][feat] Kimi K3: unlock the CUTEDSL MoE backend for NVFP4 SiTU by @xguannv in #19003
- [None][fix] Share speculative capture buffers across CUDA graph buckets by @brnguyen2 in #19350
- [None][fix] Validate media data-URIs and gate private-URL fetches behind an opt-in by @brnguyen2 in #19458
- [None][perf] Avoid prefix probes when the KV cache v2 context chunk budget is insufficient by @xwang233 in #19366
- [https://nvbugs/6708299][fix] Support group_size=32 for W4A16_NVFP4 and validate NVFP4 group sizes by @moraxu in #19091
- [https://nvbugs/6649739][fix] Skip deprecated GPT-OSS Hopper tests and remove stale waivers by @dongfengy in #19470
- [None][feat] Helix speculative verify groups: fp8 + fp4 MLA and DSpark by @reasonsolo in #19273
- [None][test] Disable ctx overlap scheduler for deepseek-v4-pro con8 disagg config by @chenfeiz0326 in #19424
- [None][feat] NVFP4 MLA residual switch for DeepSeek-V4 by @Tracin in #19305
- [TRTLLM-16398] [test] Add disaggregated speculative decoding acceptance-length baselines by @allisonlim-nv in #19176
- [TRTLLM-15820][feat] Expose Nemotron-H VL LoRA configuration hook by @eopXD in #19151
- [TRTLLM-15899][feat] Add new KDA decode non-mtp kernel and dispatch logic by @pengbowang-nv in #18938
- [https://nvbugs/6607482][logging] Add quiet PyExecutor hang diagnostics by @2ez4bz in #19045
- [https://nvbugs/6705034][fix] Route zero-token causal-conv input to the channel-major kernel by @nv-guomingz in #18571
- [TRTLLMINF-440][infra] Upgrade C++ from 17 to 20 by @EmmaQiaoCh in #19249
- [TRTLLM-16394][refactor] Consolidate and group the cpp MoE torch ops under thop/moe/ by @lori-ren in #19216
- [https://nvbugs/6758990][fix] Fix MPI flashinfer test failure by @tongyuantongyu in #19254
- [None][doc] GVR V2: Self-Sampling and Multi-Thresholding for Faster Exact Top-K by @longcheng-nv in #19244
- [#19337][fix] Correct VBWS width tracking and terminal-step finalization by @yifanQi98 in #19407
- [None][fix] make disagg load balancing router global by @reasonsolo in #19399
- [None][fix] fix _CacheReuseAdapterV1 by @bo-nv in #19480
- [TRTLLMINF-420][fix] Add Jenkins instance name to SLURM job by @lyxxn0414 in #19327
- [https://nvbugs/6777522][test] Guard VisualGen ffmpeg apt install with timeout and retry by @zhenhuaw-me in #19433
- [https://nvbugs/6813629][fix] Fix GPT-OSS Eagle3 disaggregated test configuration by @zhaoyangwang-nvidia in #19531
- [None][perf] PDL on the missing decode edges by @brb-nv in #19446
- [None][fix] size FP4 block-scale MoE permute buffers from local_num_experts by @reasonsolo in #19251
- [#19487][fix] Fix CUDA graph spec dec RNG corner case by @mikeiovine in #19505
- [TRTLLM-15718][test] Add disagg closure tests, demote one E2E to post-merge and replace another with CPU/1-GPU tests by @nv-xtf in #19478
- [None][perf] MLA context: fuse the Q FP8 quantization into the absorb bmm epilogue (CuTe-DSL BF16->FP8 bmm) + RoPE kOutputFp8Q by @Tabrizian in #19082
- [None][feat] OpenAI protocol: cap tool enum values and accept parallel_tool_calls by @brnguyen2 in #19459
- [None][fix] Give test_mnnvl_memory an expert-parallel mapping and assert its checks by @brnguyen2 in #19509
- [None][chore] Add per-rank routing/scheduling observability to the iteration log by @lancelly in #16323
- [https://nvbugs/6812347][fix] Do not pass expert-count args to the Rubin MoE finalize op by @zhangcl in #19512
- [None][fix] Drain guided-decoder host functions before GIL-holding forward calls by @brnguyen2 in #19455
- [https://nvbugs/6535765][fix] Raise Wan 2.2 LPIPS threshold to 0.25 and unwaive by @o-stoner in #19438
- [None][perf] Publish MoE all-to-all counters before the dispatch data issue by @zhangcl in #19425
- [#19485][bugfix] Prevent OOB index in vanilla top-p sampling by @DhineshPonnarasan in #19496
- [None][build] remove obsolete cpp_only build option by @Funatiq in #19421
- [fix] Reuse the pending MNNVL all-reduce communicator after workspace construction retries by @trtllm-agent in #19101
- [None][fix] DFlash buffer sizing/position fixes and an accurate FlashInfer speculation gate by @brnguyen2 in #19466
- [None][perf] Enable flashinfer gdn prefill on sm120/121 by @amukkara in #19376
- [None][feat] Add opt-in per-request cached KV token logging by @qiaoxj07 in #19375
- [None][feat] Rubin fp4 mla refactor(P1) by @Tracin in #18478
- [None][fix] suppress GCP host detection xtrace in CI images by @hanjingtian in #19345
- [None][fix] Guard model_engine access in PyExecutor._handle_responses by @brnguyen2 in #19559
- [TRTLLMINF-391][infra] Upgrade public PyTorch to 2.14 and integrate Triton 3.8 by @EmmaQiaoCh in #19477
- [https://nvbugs/6771102][fix] Support Qwen3.5 global FP8 checkpoints by @2ez4bz in #19519
- [None][perf] Triton top-k combine for Marlin NVFP4 MoE by @rmeghwal-nv in #18999
- [None][fix] Synchronize disaggregated KV receive destinations by @xwang233 in #19444
- [None][test] Add Helix zero-KV multi-GPU regression coverage by @xguannv in #18995
- [TRTLLM-16614][chore] Remove more TRT leftovers by @Funatiq in #19428
- [None][fix] Refuse paged-context attention when no fused kernel exists by @brnguyen2 in #19353
- [https://nvbugs/6812588][fix] Stop unbindable IPv6 addresses from reaching disaggregated listeners by @reasonsolo in #19544
- [None][fix] Fix DFlash budget test broken by CachedModelLoader signature change by @weiminwang-nv in #19582
- [None][perf] Extend MiniMax-M3 piecewise CUDA graphs coverage by @peihu-nv in #19423
- [None][feat] Add support for mnnvl allreduce under ray orchestrator by @shuyixiong in #18231
- [https://nvbugs/6737124][fix] Get rid of incorrect assertion for DSv4 E2E test by @2ez4bz in #19554
- [None][fix] Keep one-sided MoE combine workspace offset stable by @roborluo in #19312
- [None][test] expand Qwen3.8 test GPU coverage by @crazydemo in #19530
- [TRTLLM-16022][fix] Authenticate disaggregated subagent affinity by @xwang233 in #19175
- [https://nvbugs/6771023][fix] Trace startup and warmup, allow 900s and unwaive Blackwell PP4 tests by @zhaoyangwang-nvidia in #19402
- [None][fix] Add SM107 (Rubin) to SM100-family support statements; gate W4A8 NVFP4 FP8 MoE to SM100/103 by @farazkh80 in #19361
- [None][models] Avoid GEMM in Gemma4 vision RoPE by @2ez4bz in #19571
- [None][fix] Update Scaffolding MCP transport and KV cache truncation by @wu1du2 in #19055
- [None][chore] BREAKING: Remove AutoDeploy from TensorRT-LLM by @bmarimuthu-nv in #19028
- [None][fix] Preserve KV transfer outcomes independently of polling by @chienchunhung in #19377
- [https://nvbugs/6813351][fix] Revert "optimize qwen3.5 perf by removing gather and scatter op (#17365)" by @nv-guomingz in #19578
- [None][chore] update telemetry codeowners by @tburt-nv in #19599
- [#17743][fix] Key the legacy lint baseline on POSIX paths by @edenfunf in #17745
- [None][fix] Compute the Cosmos3 rotary table with a broadcast multiply, not a K=1 matmul by @ishovkun in #19532
New Contributors
- @daichu-nv made their first contribution in #18906
- @yifanQi98 made their first contribution in #19407
- @lyxxn0414 made their first contribution in #19327
Full Changelog: v1.3.0rc28...v1.3.0rc29