github vllm-project/vllm-metal v0.30.0.dev20260929101557

pre-release3 hours ago

What's Changed

  • [Tools] Reject failed server startup in the STT smoke test by @HanningLin in #883
  • [Cleanup] Narrow the capacity-path runtime and drop an unused ModelAdapter member by @Lavmee in #872
  • [Tools] Match temperatures in the sampling microbenchmark by @HanningLin in #886
  • [Bugfix] Honor --logprobs-mode for sample logprobs by @tak-bro in #881
  • [Attention] Log the MLA kernel and spec-verify window paths at startup by @Lavmee in #874
  • [Kernel] Keep window-masked scans neutral in the online softmax by @begcdn in #884
  • [Perf] Route bounded TurboQuant prefill through NAX/tiled attention by @suntp in #853
  • [Docs] Say which Gemma 4 conversions ship vision weights by @luke-edward in #879

New Contributors

Full Changelog: v0.30.0.dev20260928205037...v0.30.0.dev20260929101557

Don't miss a new vllm-metal release

NewReleases is sending notifications on new releases.