FastFlowLM v0.9.24 is here with a powerful new LiquidAI model, smarter caching, safer downloads, and better runtime control β making your NPU experience faster, smoother, and more reliable than ever.
π Expanded LiquidAI Support
| Feature | Details |
|---|---|
| New Model | LFM2:2.6B added β delivers 31+ tokens/sec (tps) in decoding and 1000+ tps in prefill.
|
| Faster Decoding | LFM2:1.2B now reaches 63+ tokens/sec
|
π Redownload is required for the new
LFM2:1.2B.
π Whatβs New & Why It Matters
| Feature | Benefit |
|---|---|
| Prompt Cache (Server Mode) | Reuses recent inputs to cut latency and speed up multi-turn conversations. Inspired by llamacpp. |
| Download Integrity Checks | Every model download is now checksum-verified to ensure reliability and correctness. Huge thanks to @ramkrishna2910 for reporting and @jeremyfowers for guidance! |
| Interrupt During Decoding | Stop generation mid-stream in serve mode for tighter control over long or unwanted outputs β another great suggestion from @jeremyfowers. |
π A huge, heartfelt thank you to our community and early adopters β your testing, feedback, bug reports, patience, and enthusiasm are what make FastFlowLM better every single day. We truly couldnβt do this without you.
β¨ππ π
Happy Holidays! Wishing you warmth, joy, and inspiration β and we canβt wait to keep building amazing AI together in the new year.