What's Changed
- Add and improve MLX-VLM agent skills by @Lazarus-931 in #1747
- Skip full-sequence LM head during Gemma 4 prefill by @Lazarus-931 in #1750
- Support current LFM2 feed-forward configs by @Blaizzy in #1786
- Load Kimi K3 processor and tokenizer without remote code by @Blaizzy in #1787
- Bump version to 0.6.10 by @Blaizzy in #1789
- Convert Kimi K3 tiktoken tokenizers at runtime by @Blaizzy in #1790
- Expose per-request speculative decoding acceptance stats in response timings by @Lazarus-931 in #1788
Full Changelog: v0.6.9...v0.6.10