FastFlowLM v0.9.20 introduces substantial performance improvements across multiple model families, with special focus on decoding efficiency.
β‘ Performance Improvements
πΈ 1. GPT-OSS Models
- Decoding speed of
gpt-oss:20bandgpt-oss-safeguard:20bare reaching ~19 tokens/sec and are over 60% faster at 1K context length.
πΈ 2. Gemma3 Models
gemma3:4breaching ~19 tokens/sec and enjoys over ~20% decoding speed boost at 1K context length.gemma3:1b(reaching ~43 tokens/sec)gemma3:270m(reaching ~79 tokens/sec; Note that this model is experimental)
This release is focused on raw speedβmaking FastFlowLM even more efficient for both high-capacity and portable deployments.