github vllm-project/vllm-metal v0.30.0.dev20261004205607

pre-release3 hours ago

What's Changed

  • [Perf] Project only the sampled logits rows on multimodal steps by @luke-edward in #1000
  • [Model] Serve DiffusionGemma block diffusion on the paged Metal path by @marcomd in #994
  • [KV Offload] KV cache offloading for the Metal backend by @RobbieJ in #998
  • [Perf] Calibrate hd128 TurboQuant prefill admission by @suntp in #990
  • GGUF: share the Q4 nibble split and accept remote :Q5_0/:Q5_1 tags by @begcdn in #997
  • [Benchmarks] Compare DSpark HTTP serving with matched baselines by @StevenWang-CY in #999
  • Add a KV offloading serving benchmark by @RobbieJ in #737

New Contributors

Full Changelog: v0.30.0.dev20261004100413...v0.30.0.dev20261004205607

Don't miss a new vllm-metal release

NewReleases is sending notifications on new releases.