🧠 1. New Model: LiquidAI/LFM2‑2.6B‑Transcript
Tag: lfm2‑trans:2.6b
Summarize conference notes like a pro.
A single‑turn model designed to cleanly condense long transcripts into insights — so you can spend more time sipping ☕ and less time scrolling 📜.
🎬 See it in action: https://youtu.be/hpt0EhR1_vE?si=v9OCKa7VKAzuZ-02
🛠️ 2. Tool Calling — Preview Release
Tool calling is now officially out of preview!
Verified to work with:
qwen3:4bqwen3:8bqwen3-it:4bqwen3-tk-4b
📹 Watch the demo: https://youtu.be/H-i4dztSdVk?si=5keyfkHt3ii8Wlu0
📘 Setup instructions (local):
👉 https://fastflowlm.com/docs/instructions/server/tool_calling/
💽 3. Installer Upgrade — xclbins Inside!
All xclbins are now bundled in the installer, which means:
- 🆙 Faster updates, No re‑downloading models unless weights change
- 🤯 Fewer user headaches
- 🚀 We are able to keep pushing the performance and efficiency.
More performance tuners coming soon… 🔧⚡
🔁 4. Runtime Restructure for Fine‑Tuned Models
We’ve overhauled the FastFlowLM runtime to let YOU plug in fine‑tuned models from supported families.
This is made possible by the upcoming gguf → q4nx conversion tool —
it’s almost ready and the docs are currently baking 🍳.
Stay tuned — this one will unlock a lot of flexibility.
🙌 Acknowledgements
- Huge thanks to @ItzCrazyKns from Perplexica for schooling us in the basics of tool calling and all the help along the way!
- Huge thanks to @jeremyfowers for highlighting and helping us resolving the ambiguity in the JSON-formatted reasoning output!