π¦ Whatβs New
β‘ Prefill Speed Upgrade (Qwen & Gemma Families)
A new attention engine dramatically accelerates prefill, with larger gains at longer context lengths (especially 16K+).
- π Up to 3.8Γ faster prefill
(qwen3:0.6b with a 32K-token prompt)
π Prefill Speed @ 32K Prompt (tokens/sec)
| Model | Before β After | Speedup |
|---|---|---|
| gemma3:1b | 1596 β 1755 | 1.1Γ |
| gemma3:4b | 673 β 926 | 1.4Γ |
| medgemma:4b | 673 β 926 | 1.4Γ |
| qwen3:0.6b | 236 β 1496 | 3.8Γ |
| qwen3:1.7b | 225 β 768 | 3.4Γ |
| qwen3:4b | 164 β 303 | 1.9Γ |
| qwen3-it:4b | 164 β 303 | 1.9Γ |
| qwen3-tk:4b | 164 β 303 | 1.9Γ |
| qwen3vl-it:4b | 164 β 303 | 1.9Γ |
| qwen3:8b | 150 β 260 | 1.7Γ |
| deepseek-r1-0528:8b | 150 β 260 | 1.7Γ |
πΌοΈ Note: For qwen3vl-it:4b, image understanding is also faster in this release.
π Benchmark Results
- π Gemma3 performance: https://fastflowlm.com/docs/benchmarks/gemma3_results/
- π Qwen3 performance: https://fastflowlm.com/docs/benchmarks/qwen3_results/
π οΈ Vision Tool Calling (New)
- β Tool calling is now supported on qwen3vl-it
- π Enables vision tool calling workflows
- π₯ Demo: https://youtu.be/Rf6r0Fm1UVs?si=u45hBgFXyDeEKXxh
π§ Client Compatibility Improvements
- Non-stream mode logic adjusted
- Improves compatibility with client applications