github turboderp-org/exllamav3 v0.0.32
0.0.32

latest releases: v1.5.1, v1.5.0, v1.4.9...
4 months ago
  • Fix TP regression in v0.0.31
  • Fix Gemma4 vision model (wrong norm)
  • GEMM kernel autotuning with disk cache
  • Specialized GDN kernel instances
  • Allow constrained generation with draft model
  • Many kernel optimizations and fused ops

Speed improvement (TG) relative to v0.0.31:

. 3090¹ 4090¹ 5090¹ 6000 Pro¹ 5090² 6000 Pro²
Qwen3.5-35B-A3B 4.00bpw 5.3% 5.8% 8.6% 10.3% 21.0% 23.5%
Qwen3.5-27B 4.00bpw 0.0% 1.9% 8.1% 11.7% 13.1% 15.0%
Trinity-Nano 4.15bpw 29.5% 48.6% 52.3% 52.9% 70.5% 72.4%
Gemma4-26B-A4B 4.10bpw 3.1% 2.9% 7.8% 9.6% 16.4% 19.2%
Gemma4-31B 4.00bpw 4.0% 4.9% 10.0% 8.0% 16.0% 12.0%

¹ CUDA 12.8
² CUDA 13.2 (improves performance further for Blackwell, no change for Ampere/Ada)

Full Changelog: v0.0.31...v0.0.32

Don't miss a new exllamav3 release

NewReleases is sending notifications on new releases.