github vllm-project/vllm-metal v0.30.0.dev20261001204624

pre-release2 hours ago

What's Changed

  • [Docs] Point the Gemma 4 paged example at a live checkpoint by @Luan-Fuzi in #937
  • [Test] Reuse run_in_spawn_process in the chunked-prefill e2e by @begcdn in #935
  • [Test] Pin MLA_KERNEL_BLOCK_SIZES to the mla.metal instantiations by @begcdn in #933
  • [Cleanup] Deduplicate the boolean environment switch parsing by @begcdn in #938
  • [Perf] Cast attention sinks to float32 once, not per layer per forward by @begcdn in #939
  • Keep selective logits when multimodal input is disabled by @LxYuan0420 in #931
  • [Spec Decode] Support batch-size-based DFlash drafting by @StevenWang-CY in #932
  • [Test] Check the lazy GGUF skeleton keeps load peak off the eager path by @begcdn in #934
  • [Cleanup] Fold the DFlashProposer isinstance checks into MetalProposer by @begcdn in #936

Full Changelog: v0.30.0.dev20261001081948...v0.30.0.dev20261001204624

Don't miss a new vllm-metal release

NewReleases is sending notifications on new releases.