What's Changed
- [Multimodal] Serve Gemma 4 E2B/E4B images through the vision sidecar by @luke-edward in https://github.com/vllm-project/vllm-metal/pull/992
- [Spec Decode] Add experimental greedy DSpark serving by @StevenWang-CY in https://github.com/vllm-project/vllm-metal/pull/986
- Load standalone draft models before memory profiling by @LxYuan0420 in https://github.com/vllm-project/vllm-metal/pull/993
- [Cleanup] Point missing capability bindings at a rebuild, drop dead probe by @begcdn in https://github.com/vllm-project/vllm-metal/pull/984
- [Tests] Import compare() from attention_bench_utils, merge its cases by @begcdn in https://github.com/vllm-project/vllm-metal/pull/985
- [Tests] Add a truth-table test for _mm_forward_forced by @begcdn in https://github.com/vllm-project/vllm-metal/pull/983
- [Docs] Add a macOS serving benchmark guide by @suntp in https://github.com/vllm-project/vllm-metal/pull/989
Full Changelog: https://github.com/vllm-project/vllm-metal/compare/v0.30.0.dev20261003204550...v0.30.0.dev20261004083402