This release adds qwen3.8-mtp:27b β the largest dense model FastFlowLM has shipped to date, and the first to run speculative decoding on the NPU through its built-in MTP draft head. It also renames hy-mt2:1.8b to hy-mt2-flash:1.8b to match the flash naming convention, and picks up four community fixes to the OpenAI usage contract, CLI exit codes, flm validate, and a stray file in the release packages.
π¦ New Model Support
π§ Qwen3.8-27B
FastFlowLM now supports qwen3.8-mtp:27b, a 27B reasoning model with tool calling β the largest dense model FastFlowLM has shipped, and the first to run speculative decoding on the NPU through its built-in MTP draft head.
- Tag:
qwen3.8-mtp:27b
Run in CLI mode:
flm run qwen3.8-mtp:27bRun in server mode:
flm serve qwen3.8-mtp:27bπΎ This is by far the largest dense model FastFlowLM has shipped. Check that you have the disk space for the pull and enough system memory to hold it resident before running it.
β‘ What MTP Speculative Decoding Means
Qwen3.8 ships with a multi-token prediction (MTP) head β a small draft model trained alongside the main network. FastFlowLM now puts it to work:
- The MTP head drafts several candidate tokens in one shot.
- The full model verifies them in a single pass.
- Accepted drafts are kept; the first rejected one is replaced by the base model's own choice, and drafting restarts from there.
The practical consequences:
- Faster decoding, same output. Verification is an exact compare against what the base model would have emitted, so every token you receive is a token the base model would have produced on its own. Speculation changes throughput, not quality.
- Gains are prompt-dependent. Predictable, low-entropy stretches β code, structured output, long reasoning chains β see the highest draft acceptance. Highly creative or surprising text accepts fewer drafts and converges toward ordinary decode speed.
- Nothing to configure. Speculation is driven by the engine; there is no flag to set and no API change.
Every other model is unaffected β speculation is opt-in at the engine level, so the rest of the lineup decodes exactly as it did in v1.0.6.
For model details, context limits, and measured speedups, see the model card and benchmark results.
π hy-mt2:1.8b β hy-mt2-flash:1.8b
The multilingual translation model introduced in v1.0.5 became single-turn in v1.0.6. Its name now says so: it is hy-mt2-flash:1.8b.
This is a rename only β same weights, same 1k context, same single-turn behavior, same translation quality. It simply brings the tag in line with gemma4e-flash and qwen3vl-flash, so that "flash" consistently signals the same contract: single-turn, fixed context, optimized kernels.
Action required: update any script, config, or client that pins the old tag.
- flm run hy-mt2:1.8b
+ flm run hy-mt2-flash:1.8b- "model": "hy-mt2:1.8b"
+ "model": "hy-mt2-flash:1.8b"The recommended prompt shape from the v1.0.5 guide is unchanged:
{"role": "system", "content": "ε°δ»₯δΈζζ¬ηΏ»θ―δΈΊθ±θ―οΌζ³¨ζεͺιθ¦θΎεΊηΏ»θ―εηη»ζοΌδΈθ¦ι’ε€θ§£ιγθΎεΊεΏ
ι‘»ε
¨ι¨δ½Ώη¨θ±θ―οΌδΈθ¦θΎεΊζΊθ―θ¨ζεζ"},
{"role": "user", "content": "{TEXT}"}For more details, see the model card and benchmark results.
π Fixes from the Community
Four fixes in this release came from outside contributors. Thank you β these are exactly the kind of sharp, well-scoped reports that make the project better. π
π OpenAI usage Now Reports the Full Prompt Length
PR #729 β thanks to @Javinator9889 (Javier Alonso)
The prompt cache strips the matching prefix before prefill, so usage.prompt_tokens was reporting only the newly evaluated suffix rather than the whole input. The first turn of a conversation looked correct; every turn after that reused the history as a cached prefix and reported just the new message.
That broke agentic front-ends badly. Tools like OpenCode size their context from usage.prompt_tokens, so the context appeared to reset on every request β built-in compaction never triggered, and the client eventually hit a context overflow, re-fed the conversation, and overflowed again in a loop.
What changed:
usage.prompt_tokensnow counts the entire input, cached prefix included β matching the OpenAI specification.usage.prompt_tokens_details.cached_tokensis now exposed on the OpenAI endpoints, so clients can see how much of that input was served from cache.prefill_speed_tpsnow divides by the tokens actually evaluated, so a cache hit no longer inflates the reported prefill speed.- Ollama-compatible
prompt_eval_countis unchanged β it continues to report evaluated tokens only, matching upstream Ollama.
If you drive FLM from an agent framework that manages its own context window, this is the fix to upgrade for.
β©οΈ flm help, flm version, and flm port Now Exit 0
PR #723 β thanks to @jtuyls (Jorn Tuyls)
These three commands succeeded and then exited with status 1, because they stopped argument parsing the same way a usage error does. Any script, CI job, or shell with set -e that ran flm version treated a perfectly good invocation as a failure.
They are now handled as successful commands β and handled before the model list is loaded, so flm help and flm version no longer pay that startup cost. Genuine usage errors still exit 1.
β
flm validate No Longer Hard-Checks the Kernel Version
PR #738 β thanks to @superm1 (Mario Limonciello)
flm validate required kernel 6.17 or newer. Several distributions backport the NPU driver to considerably older kernels, so working setups were being reported as invalid.
The check is now removed. Validation already checks the firmware version, which implicitly requires a recent enough driver β making the kernel version test both redundant and wrong for backported kernels.
π§Ή Stray Backup File Removed from the Release Packages
PR #747 β thanks to @yorickvP
src/lib/xrt/libq4_npu_eXpress.so.bak-20260826 β a backup copy of a shared library β was committed by accident and had been shipping inside the release packages, despite being used by neither the build nor the executable. It has been deleted, so the Linux artifacts are a little smaller.
π Acknowledgements
| Contributor | Contribution |
|---|---|
| @Javinator9889 | #729 β full prompt length in OpenAI usage
|
| @jtuyls | #723 β exit 0 for help, version, and port
|
| @superm1 | #738 β drop the kernel check in flm validate
|
| @yorickvP | #747 β remove the stray .so.bak file
|
π Summary
| Highlight | |
|---|---|
| π¦ | New model: Qwen3.8-27B (qwen3.8-mtp:27b) β reasoning and tool calling, the largest dense model FastFlowLM has shipped
|
| β‘ | First model with MTP speculative decoding: a draft head proposes, the full model verifies β faster decode, identical output |
| π | hy-mt2:1.8b renamed to hy-mt2-flash:1.8b β rename only; update pinned tags
|
| π | usage.prompt_tokens reports the full input, cached_tokens is now exposed, and prefill_speed_tps is no longer inflated by cache hits
|
| β©οΈ | flm help, flm version, and flm port exit 0 instead of 1
|
| β | flm validate no longer rejects backported kernels older than 6.17
|
| π§Ή | Stray .so.bak backup file removed from the release packages
|
Thanks for your support β see you in the next one! π