Happy Thanksgiving!
Today weβre dropping one of our biggest speed upgrades ever for LLaMA and DeepSeek models (our first batch of models) β just in time for the holiday break. Fire up your Ryzenβ’ AI NPU and enjoy some seriously boosted performance. π₯
π 1. Quantization Upgrade
- All models migrated from AWQ to Q4_1
- Better LLM accuracy and quality.
β‘ 2. Massive Decoding Speedup
llama3.2:1b: ~50% faster decoding, reaching 66 tpsllama3.2:3b: ~40% faster decoding, reaching 28 tpsllama3.1:8b: ~40% faster decoding, reaching 13 tpsdeepseek-r1:8b: ~40% faster decoding, reaching 13 tps
π 3. Prefill Phase Optimized
- Slight improvements to prefill speed of all above, especially impactful for large context initializations.
ποΈ 4. Standalone Whisper ASR Server
You can now serve Whisper (OpenAIβs ASR model) as a standalone model for speech transcription β or pair it with GPU LLMs in a hybrid pipeline.
Use either:
flm serve -a 1or
flm serve --asr 1This release wraps up a bundle of performance gifts for LLaMA models on FastFlowLM.
Thank you for being part of the FastFlowLM journey β and happy Thanksgiving! π¦π₯