github vllm-project/vllm-metal v0.30.0.dev20260930202258

pre-release2 hours ago

What's Changed

  • Add paged MLA/KDA serving for Ling-3.0 (Bailing V3) by @FENP in #607
  • [Bugfix] Refuse a prebuilt extension whose MLX differs from the pinned version by @abhijithneilabraham in #925
  • [Bugfix] Keep the full-row logits profile for forward-ready multimodal adapters by @abhijithneilabraham in #923
  • [Test] Poison one threadgroup's masked tile per case and compare from its first row by @abhijithneilabraham in #920
  • [Bugfix] Scale the GDN q/k normalization epsilon by the head dimension by @abhijithneilabraham in #919
  • [Bugfix] Skip sampling for intermediate prefill chunks by @StevenWang-CY in #917
  • [Cleanup] Dedupe TurboQuant prefill formulas and share admission checks by @begcdn in #916
  • [Perf] Cache decode batch indices per forward and chunk over the token cap by @begcdn in #913
  • [Spec Decode] Bucket DFlash context for compiled drafting by @StevenWang-CY in #915

Full Changelog: v0.30.0.dev20260930163440...v0.30.0.dev20260930202258

Don't miss a new vllm-metal release

NewReleases is sending notifications on new releases.