github turboderp-org/exllamav3 v1.4.3
1.4.3

latest releases: v1.6.0, v1.5.4, v1.5.3...
one month ago
  • Support GlmMoeDsaForCauslLM (GLM 5.2, still probably somewhat WIP)
  • Partial CPU layer expert offloading option with dynamic placement (supersedes expert cache)
  • Improved dynamic draft sizing with auto-calibrated confidence thresholds
  • New (experimental) quant-optimizer pipeline
  • Support for mid-stream text injection in generator (enables reasoning token budget)
  • More precise autosplit allocation
  • Reduced (and now stable) VRAM footprint for DSA prefill
  • CPU cache offloading supported in TP mode
  • Bugfixes and QoL improvements
  • Remove incomplete Nanochat implementation

Full Changelog: v1.4.2...v1.4.3

Don't miss a new exllamav3 release

NewReleases is sending notifications on new releases.