What's Changed
- [Docs] Point the Gemma 4 paged example at a live checkpoint by @Luan-Fuzi in #937
- [Test] Reuse run_in_spawn_process in the chunked-prefill e2e by @begcdn in #935
- [Test] Pin MLA_KERNEL_BLOCK_SIZES to the mla.metal instantiations by @begcdn in #933
- [Cleanup] Deduplicate the boolean environment switch parsing by @begcdn in #938
- [Perf] Cast attention sinks to float32 once, not per layer per forward by @begcdn in #939
- Keep selective logits when multimodal input is disabled by @LxYuan0420 in #931
- [Spec Decode] Support batch-size-based DFlash drafting by @StevenWang-CY in #932
- [Test] Check the lazy GGUF skeleton keeps load peak off the eager path by @begcdn in #934
- [Cleanup] Fold the DFlashProposer isinstance checks into MetalProposer by @begcdn in #936
Full Changelog: v0.30.0.dev20261001081948...v0.30.0.dev20261001204624