- [Model Routing & Deprecated Model Seamless Redirect] Fix Stale gemini-3.1-flash-lite Redirect, Restore Layer-3 Background Summary, and Route Retired 2.5 Family to gemini-3.6-flash-medium (Fixes #3577, Thanks to @Xyloz3n):
- Correct gemini-3.1-flash-lite Direct Passthrough: Permanently removed the hardcoded redirect that sent healthy
gemini-3.1-flash-literequests to the retiredgemini-2.5-flash-lite. Restored 1:1 passthrough to upstream's active 1M-contextMODEL_PLACEHOLDER_M50, eliminating upstream 429/503 errors and serving requests as 200 OK. - Revive internal-background-task: Retargeted the internal virtual model
internal-background-taskand background task constants in OpenAI/Gemini handlers from dead 2.5 models to the fast, healthygemini-3.1-flash-lite(1M context), fully restoring Layer-3 conversation history compression and background summarization. - Purge Dead Models from Advertised Catalog: Removed the entirely retired 2.5 family (
gemini-2.5-pro,gemini-2.5-flash,gemini-2.5-flash-thinking,gemini-2.5-flash-lite) andgemini-3.5-flash-litefromget_supported_models(),is_model_compliant_with_baseline, and frontend menus to eliminate dead options, while addinggemini-3.1-flash-liteto the compliant catalog. - Seamless In-Flight Redirects for Deprecated Names: Retained internal backwards-compatible routing rules for legacy clients and scripts:
gemini-2.5-flash,gemini-2.5-flash-thinking, andgpt-3.5-turbosmoothly redirect togemini-3.6-flash-medium;gemini-2.5-flash-liteandgemini-3.5-flash-literedirect togemini-3.1-flash-lite; andgemini-2.5-proredirects togemini-pro-agent.
- Correct gemini-3.1-flash-lite Direct Passthrough: Permanently removed the hardcoded redirect that sent healthy
- [Downstream SSE Heartbeat Tightening] Reduce SSE Keep-Alive Interval to 3s During Deep Thinking (PR #3578, Thanks to @EricZhou05):
- 3s Stream Heartbeat: Tightened downstream OpenAI-compatible SSE heartbeat interval to 3 seconds, preventing client-side read timeouts and connection drops during extended reasoning and deep-thinking phases.
- [Documentation & Metadata Alignment] Fix README Typos and Synchronize Bilingual Content (PR #3579, Thanks to @EricZhou05):
- Bilingual Sync: Fixed typos in documentation, aligned Chinese and English README content, and refreshed metadata.
- [模型路由治理与淘汰模型平滑重定向] 修复 3.1-flash-lite 误重定向、拯救 Layer-3 后台摘要与压缩并将已退役 2.5 系列平滑重定向至 3.6-flash-medium (Fixes #3577, Thanks to @Xyloz3n):
- 纠正 gemini-3.1-flash-lite 错误降级与健康直传: 彻底移除核心映射表中将健康存活的
gemini-3.1-flash-lite错误重定向至已故gemini-2.5-flash-lite的硬编码。恢复为其自身标准直传(出站 1:1 透传上游具备 1M 上下文的MODEL_PLACEHOLDER_M50),彻底消除由此引发的 429 与 503 报错,上游实测 200 OK。 - 拯救 internal-background-task 后台任务: 将内部虚拟模型
internal-background-task以及 OpenAI/Gemini 适配器中的后台任务模型常量从已下线的 2.5 系列重定向至健康高效且具备 1M 上下文的gemini-3.1-flash-lite,彻底恢复因 2.5 系列退役而瘫痪的 Layer-3 对话历史摘要与长上下文压缩功能。 - 宣传目录与合规基准线净化: 从内置模型列表
get_supported_models()、基准线过滤器is_model_compliant_with_baseline及前端模型菜单中彻底剔除上游已全线 503 的 2.5 全系列(gemini-2.5-pro、gemini-2.5-flash、gemini-2.5-flash-thinking、gemini-2.5-flash-lite)以及gemini-3.5-flash-lite,杜绝死菜单项与无效暴露;同时将健康的gemini-3.1-flash-lite纳入支持列表与合规基准线。 - 已淘汰模型内部平滑重定向: 内部完整保留对已淘汰旧模型的重定向兜底,保障历史客户端与自动化脚本平稳运行:
gemini-2.5-flash、gemini-2.5-flash-thinking与gpt-3.5-turbo重定向至gemini-3.6-flash-medium;gemini-2.5-flash-lite与gemini-3.5-flash-lite重定向至gemini-3.1-flash-lite;gemini-2.5-pro重定向至gemini-pro-agent;前端路由预设同步对齐。
- 纠正 gemini-3.1-flash-lite 错误降级与健康直传: 彻底移除核心映射表中将健康存活的
- [下游 SSE 心跳保活收紧] 缩短流式思考心跳间隔至 3 秒防止长推理连接中断 (PR #3578, Thanks to @EricZhou05):
- 3 秒流式心跳注入: 将下游 OpenAI 协议流式输出中的 SSE 心跳保活间隔缩短至 3 秒,防止客户端在长思考/深度推理期间因长久无响应而提前断开连接。
- [文档与元数据排版校准] 修复 README 错别字并对齐中英双语内容 (PR #3579, Thanks to @EricZhou05):
- 双语文档对齐: 修复 README 文档排版错别字,对齐中英双语内容并更新相关元数据。
What's Changed
- docs: fix README typos, align bilingual content and update metadata by @EricZhou05 in #3579
- fix(proxy): reduce downstream SSE heartbeat interval to 3s to prevent client disconnect during deep thinking by @EricZhou05 in #3578
Full Changelog: v4.9.0...v4.9.1