FastFlowLM v0.9.23 introduces a 35% speed boost for image understanding in vision-enabled Gemma models.
β‘ Vision Prefill Optimization
Models Improved:
gemma3:4bmedgemma:4b
Whatβs new:
- Reduced Time to First Token (TTFT) from ~4.5s β ~3.4s for single-image cases (35% speedup).
- Noticeably faster responses in visual chat and medical imaging scenarios.
π οΈ FLM Runtime Improvements
- Model downloads now print both the total download size and per-file sizes β thanks to @jeremyfowers for the suggestion!
π€ Flm-Companion (Independent Project)
Flm-Companion, created by @julienM77, is a modern GUI designed to complement FastFlowLM.
It provides an intuitive way to run local models, monitor the server, and manage configurations.
- Latest version: v0.4.0
- Changelog: https://github.com/julienM77/Flm-Companion/releases
This release continues our push to accelerate on-device multimodal inference, especially for latency-sensitive workflows.