github ROCm/FastFlowLM v0.9.20
πŸš€ FastFlowLM v0.9.20 β€” Massive Decoding Speed Boosts for GPT-OSS & Gemma3

latest releases: v1.0.4, v1.0.3, v1.0.2...
9 months ago

FastFlowLM v0.9.20 introduces substantial performance improvements across multiple model families, with special focus on decoding efficiency.


⚑ Performance Improvements

πŸ”Έ 1. GPT-OSS Models

  • Decoding speed of gpt-oss:20b and gpt-oss-safeguard:20b are reaching ~19 tokens/sec and are over 60% faster at 1K context length.

πŸ”Έ 2. Gemma3 Models

  • gemma3:4b reaching ~19 tokens/sec and enjoys over ~20% decoding speed boost at 1K context length.
  • gemma3:1b (reaching ~43 tokens/sec)
  • gemma3:270m (reaching ~79 tokens/sec; Note that this model is experimental)

This release is focused on raw speedβ€”making FastFlowLM even more efficient for both high-capacity and portable deployments.

Don't miss a new FastFlowLM release

NewReleases is sending notifications on new releases.