- Reduced (flat) VRAM overhead for GDN/KDA/Mamba drafting
- Support recurrent draft models, switch DFlash to sliding-attention (reduced VRAM overhead)
- New hybrid drafting mode (draft model + long-ngram)
- Workaround for Triton issue breaking kernel autotune on mixed-architecture setups
- Lots of ROCm fixes and optimizations
- Turing (sm_75) optimizations
- AVX2 optimizations
- Many bugfixes
Full Changelog: v1.6.0...v1.6.1