What's Changed
- Enable Qwen3-Omni multimodal batching by @Lazarus-931 in #2349
- Review/extraction test harness by @Lazarus-931 in #2394
- Fix broken CI: Qwen-compatible tokenizer in the processor regression test by @Lazarus-931 in #2401
- Add Decider and the shared typed decision API by @Lazarus-931 in #2396
- Add Laya on the shared decision API by @Lazarus-931 in #2397
- Keep the most likely MoG component for a tiny top_p in Nemotron VoiceChat by @pierre427 in #2405
- Stop gradients on MoE router selection indices (MLX >= 0.32.1) by @Lazarus-931 in #2409
- removing duplicate rotate_half and check_array_shape funcs by @Lazarus-931 in #2414
- Feat/decision cli by @Lazarus-931 in #2402
- Add conversation compaction to the Responses API by @Blaizzy in #2408
- Fix prequantized Qwen4 PLE loading by @Lazarus-931 in #2417
- Fix PhiMoE routing and plain RoPE configs by @Lazarus-931 in #2419
- Compute top-p and typical-p probability mass in float32 by @pierre427 in #2404
- qwen3_vl: take video timestamps from the frames that were actually sampled by @Marian2110 in #2413
- Bump version to 0.7.5 by @lucasnewman in #2432
New Contributors
- @Marian2110 made their first contribution in #2413
Full Changelog: v0.7.4...v0.7.5