github turboderp-org/exllamav3 v1.6.1
1.6.1

4 hours ago
  • Reduced (flat) VRAM overhead for GDN/KDA/Mamba drafting
  • Support recurrent draft models, switch DFlash to sliding-attention (reduced VRAM overhead)
  • New hybrid drafting mode (draft model + long-ngram)
  • Workaround for Triton issue breaking kernel autotune on mixed-architecture setups
  • Lots of ROCm fixes and optimizations
  • Turing (sm_75) optimizations
  • AVX2 optimizations
  • Many bugfixes

Full Changelog: v1.6.0...v1.6.1

Don't miss a new exllamav3 release

NewReleases is sending notifications on new releases.