github InternLM/lmdeploy v0.18.0

3 hours ago

What's Changed

🚀 Features

💥 Improvements

  • fix: shard MLA attention under hybrid DP and TP by @CUHKSZzxy in #4908
  • refactor(serving): share chat serving runner by @lvhan028 in #4876
  • feat(pytorch): prototype TurboMind W4A16 backend for AWQ linear by @qescccczmr in #4897
  • refactor(pytorch): redesign CacheEngine around plans and allocations by @grimoire in #4862
  • fix: reject multimodal requests for text models by @CUHKSZzxy in #4926
  • refactor(pytorch): replace backend operator builders with typed build specs by @grimoire in #4890
  • Migrate TurboMind to C++20 by @lzhangzz in #4946
  • feat(pytorch): add piecewise CUDA graph prefill by @grimoire in #4895
  • feat: add XTuner TileLang sparse MLA backend by @CUHKSZzxy in #4904
  • feat: add request-only cache usage metric by @grimoire in #4940
  • refactor: stream tool parameters incrementally by @lvhan028 in #4802
  • feat: support checkpoint-engine weight updates by @lvhan028 in #4913
  • feat: add symmetric-memory LM-head all-gather by @qescccczmr in #4915
  • Refactor PyTorch scheduler ownership and API boundaries by @grimoire in #4921
  • feat(gpt-oss): combine xgrammar structural_tag grammar with Harmony parser by @windreamer in #4907
  • Decouple GEMM workspace from Linear handle by @lzhangzz in #4978
  • feat(kv_connector): mooncake store support mtp by @caikun-pjlab in #4948
  • fix(pytorch): avoid health RPC scheduling delay by @lvhan028 in #4966
  • refactor: remove legacy OpenAI API client by @lvhan028 in #4993

🐞 Bug fixes

  • fix: fall back to model generation config by @CUHKSZzxy in #4929
  • fix: initialize MLA KV-B after online FP8 quantization by @CUHKSZzxy in #4925
  • fix tool_choice='required' by @lvhan028 in #4909
  • fix: exclude glm-5.2 mtp projection from online fp8 by @CUHKSZzxy in #4945
  • Fix TurboMind vision workspace overflows by @irexyc in #4924
  • fix(proxy): interpolate node_url in the terminate_node failure response by @Anai-Guo in #4941
  • fix(disagg): make zmq_disconnect synchronous so p2p_drop_connect actually closes the sockets by @Anai-Guo in #4916
  • Remove cuBLAS grouped GEMM and give cuBLAS dense priority over native BF16/FP16 kernels by @lzhangzz in #4975
  • fix(internvit): allocate pinned input buffers on demand by @irexyc in #4982
  • fix(serve): tolerate messages without a content key in multimodal preprocessing by @matrix72c in #4994
  • fix: enforce weights-only loading for MemDecode router checkpoints by @lvhan028 in #4988
  • fix(dlinfer): remove CUDA hardcode in NTK rotary embedding for non-CUDA accelerators by @li-lizhe in #4986
  • fix: build NCCL stub with public device API header by @lvhan028 in #4998
  • fix: correct tool choice handling for Intern-S and GPT-OSS by @lvhan028 in #4997
  • fix: restore getenv after environment parsing errors by @adenzhou1350 in #5000

🌐 Other

  • [Docs] Fix A100 FP16 benchmark labels by @BingH225 in #4919
  • fix(deps): pin xgrammar below 0.2.5.post1 to unbreak unit-test CI by @windreamer in #4931
  • Assert the logprobs count in the restful return checks by @David-Wu1119 in #4934
  • docs: align docstring parameter names with signatures by @simpleqt in #4938
  • [Docs] Fix syntax in the Chinese output logits example by @BingH225 in #4959
  • docs: fix dead heading anchors in CONTRIBUTING and get_started pages by @simpleqt in #4937
  • improve(autotest): replace APIClient with OpenAI SDK and strict param checks by @littlegy in #4887
  • ci: add swap for CUDA release builds by @lvhan028 in #4974
  • [ci] refactor(autotest): consolidate layout-based pytest cases by @zhulinJulia24 in #4881
  • Test: update sleep/wakeup and abort scenarios by @littlegy in #4528
  • docs: sync ja README quickstart python version with en/zh (3.10 -> 3.12) by @simpleqt in #4939
  • bump version to v0.18.0 by @lvhan028 in #4990

New Contributors

Full Changelog: v0.17.0...v0.18.0

Don't miss a new lmdeploy release

NewReleases is sending notifications on new releases.