github vllm-project/vllm-metal v0.30.0.dev20261002095739

pre-release5 hours ago

What's Changed

  • [Bugfix] Sanitize speech checkpoint keys before matching them for quantization by @abhijithneilabraham in #946
  • [Test] Let run_in_spawn_process label its child for failure messages by @begcdn in #949
  • [Test] Assert profile_run invokes the drafter's profile_warmup by @begcdn in #950
  • [Perf] Warm the float32 sink cache at patch time by @begcdn in #951
  • [Test] Drift-check the full MLA instantiation space against the C++ gate by @begcdn in #952
  • [Perf] Add P256/P512 GQA decode kernels and tests by @suntp in #715

Full Changelog: v0.30.0.dev20261002084958...v0.30.0.dev20261002095739

Don't miss a new vllm-metal release

NewReleases is sending notifications on new releases.