- Fix TP regression in v0.0.31
- Fix Gemma4 vision model (wrong norm)
- GEMM kernel autotuning with disk cache
- Specialized GDN kernel instances
- Allow constrained generation with draft model
- Many kernel optimizations and fused ops
Speed improvement (TG) relative to v0.0.31:
| . | 3090¹ | 4090¹ | 5090¹ | 6000 Pro¹ | 5090² | 6000 Pro² |
|---|---|---|---|---|---|---|
| Qwen3.5-35B-A3B 4.00bpw | 5.3% | 5.8% | 8.6% | 10.3% | 21.0% | 23.5% |
| Qwen3.5-27B 4.00bpw | 0.0% | 1.9% | 8.1% | 11.7% | 13.1% | 15.0% |
| Trinity-Nano 4.15bpw | 29.5% | 48.6% | 52.3% | 52.9% | 70.5% | 72.4% |
| Gemma4-26B-A4B 4.10bpw | 3.1% | 2.9% | 7.8% | 9.6% | 16.4% | 19.2% |
| Gemma4-31B 4.00bpw | 4.0% | 4.9% | 10.0% | 8.0% | 16.0% | 12.0% |
¹ CUDA 12.8
² CUDA 13.2 (improves performance further for Blackwell, no change for Ampere/Ada)
Full Changelog: v0.0.31...v0.0.32