github vllm-project/vllm-metal v0.30.0.dev20260926202403

pre-release4 hours ago

What's Changed

  • [Bugfix] Synchronize before reading the MLX cache in profile_run by @abhijithneilabraham in #836
  • [Spec Decode] Extract shared target feature capture by @StevenWang-CY in #834
  • [Perf] Materialize MLA prefill K/V for chunks with cached context by @begcdn in #833
  • [Attention] Bidirectional attention inside Gemma 4 image blocks by @Lavmee in #830
  • [Multimodal] Keep only the most recent backbone mode decisions by @Lavmee in #829

Full Changelog: v0.30.0.dev20260926140952...v0.30.0.dev20260926202403

Don't miss a new vllm-metal release

NewReleases is sending notifications on new releases.