- Add KimiLinearForCausalLM
- Add DFlash2
- TP support for Qwen3.8-Flash-Next, GLM5.3, DeepseekV3/V4
- MoE optimizations
- Fractional-trellis quantization mode
- Self-calibration pipeline improvements (including highly experimental YAQA support)
- Fix some memory leaks
- Fix EXL3_MOE_PINNED_ARENA on Windows and Conda
- Other optimizations and bugfixes
Full Changelog: v1.5.0...v1.5.1