What's Changed
🚀 Features
- [Feat]: Support output input logprobs by @RunningLeon in #4793
- Expand SM90 quantized GEMMs and unify TurboMind linear execution by @lzhangzz in #4943
- [ascend] support glm52 by @wanfengcxz in #4928
- Support dflash for qwen3.5 by @RunningLeon in #4789
- Move shared expert into the MoE FFN layer; shard dense FFN node-locally in MoE models by @lzhangzz in #4973
💥 Improvements
- fix: shard MLA attention under hybrid DP and TP by @CUHKSZzxy in #4908
- refactor(serving): share chat serving runner by @lvhan028 in #4876
- feat(pytorch): prototype TurboMind W4A16 backend for AWQ linear by @qescccczmr in #4897
- refactor(pytorch): redesign CacheEngine around plans and allocations by @grimoire in #4862
- fix: reject multimodal requests for text models by @CUHKSZzxy in #4926
- refactor(pytorch): replace backend operator builders with typed build specs by @grimoire in #4890
- Migrate TurboMind to C++20 by @lzhangzz in #4946
- feat(pytorch): add piecewise CUDA graph prefill by @grimoire in #4895
- feat: add XTuner TileLang sparse MLA backend by @CUHKSZzxy in #4904
- feat: add request-only cache usage metric by @grimoire in #4940
- refactor: stream tool parameters incrementally by @lvhan028 in #4802
- feat: support checkpoint-engine weight updates by @lvhan028 in #4913
- feat: add symmetric-memory LM-head all-gather by @qescccczmr in #4915
- Refactor PyTorch scheduler ownership and API boundaries by @grimoire in #4921
- feat(gpt-oss): combine xgrammar structural_tag grammar with Harmony parser by @windreamer in #4907
- Decouple GEMM workspace from Linear handle by @lzhangzz in #4978
- feat(kv_connector): mooncake store support mtp by @caikun-pjlab in #4948
- fix(pytorch): avoid health RPC scheduling delay by @lvhan028 in #4966
- refactor: remove legacy OpenAI API client by @lvhan028 in #4993
🐞 Bug fixes
- fix: fall back to model generation config by @CUHKSZzxy in #4929
- fix: initialize MLA KV-B after online FP8 quantization by @CUHKSZzxy in #4925
- fix tool_choice='required' by @lvhan028 in #4909
- fix: exclude glm-5.2 mtp projection from online fp8 by @CUHKSZzxy in #4945
- Fix TurboMind vision workspace overflows by @irexyc in #4924
- fix(proxy): interpolate node_url in the terminate_node failure response by @Anai-Guo in #4941
- fix(disagg): make zmq_disconnect synchronous so p2p_drop_connect actually closes the sockets by @Anai-Guo in #4916
- Remove cuBLAS grouped GEMM and give cuBLAS dense priority over native BF16/FP16 kernels by @lzhangzz in #4975
- fix(internvit): allocate pinned input buffers on demand by @irexyc in #4982
- fix(serve): tolerate messages without a content key in multimodal preprocessing by @matrix72c in #4994
- fix: enforce weights-only loading for MemDecode router checkpoints by @lvhan028 in #4988
- fix(dlinfer): remove CUDA hardcode in NTK rotary embedding for non-CUDA accelerators by @li-lizhe in #4986
- fix: build NCCL stub with public device API header by @lvhan028 in #4998
- fix: correct tool choice handling for Intern-S and GPT-OSS by @lvhan028 in #4997
- fix: restore getenv after environment parsing errors by @adenzhou1350 in #5000
🌐 Other
- [Docs] Fix A100 FP16 benchmark labels by @BingH225 in #4919
- fix(deps): pin xgrammar below 0.2.5.post1 to unbreak unit-test CI by @windreamer in #4931
- Assert the logprobs count in the restful return checks by @David-Wu1119 in #4934
- docs: align docstring parameter names with signatures by @simpleqt in #4938
- [Docs] Fix syntax in the Chinese output logits example by @BingH225 in #4959
- docs: fix dead heading anchors in CONTRIBUTING and get_started pages by @simpleqt in #4937
- improve(autotest): replace APIClient with OpenAI SDK and strict param checks by @littlegy in #4887
- ci: add swap for CUDA release builds by @lvhan028 in #4974
- [ci] refactor(autotest): consolidate layout-based pytest cases by @zhulinJulia24 in #4881
- Test: update sleep/wakeup and abort scenarios by @littlegy in #4528
- docs: sync ja README quickstart python version with en/zh (3.10 -> 3.12) by @simpleqt in #4939
- bump version to v0.18.0 by @lvhan028 in #4990
New Contributors
- @BingH225 made their first contribution in #4919
- @David-Wu1119 made their first contribution in #4934
- @simpleqt made their first contribution in #4938
- @li-lizhe made their first contribution in #4986
- @adenzhou1350 made their first contribution in #5000
Full Changelog: v0.17.0...v0.18.0