FreeLLMAPI v0.13.2 keeps Claude Code working on models that cap their output below what Claude Code asks for.
Fixes
- Models with a smaller output limit keep working with Claude Code. Claude Code asks for up to 128,000 output tokens. Some free models stop at 65,536 and turned the request down: Groq's gpt-oss-20b and Ollama Cloud's Nemotron 3 Super and Ultra. The router used to set those models aside and wrongly marked them as unable to use tools. Now it learns each model's output limit from the first refusal, retries straight away within that limit, and keeps within it from then on. This covers every endpoint: Claude/Anthropic, OpenAI chat, Responses and Gemini/Ollama. (#1364)
Updating
Download the installer for your platform below, or use ghcr.io/tashfeenahmed/freellmapi:v0.13.2 for Docker. The latest Docker tag also points to this release. Desktop users on v0.13.1 will see this update in the Updates panel.
Existing keys and profiles are retained. This release has no database migrations.
⭐ Like the free router? Go Premium, the live signed catalog, $19/yr, cancel anytime.