- [Dynamic Tier Router & Claude 5.5 Spec Architecture Upgrade] Data-Driven DynamicTierRouter, Automated Tier Scanning, and Full Claude 5.5 Support (Inspired by PR #3594, Thanks to @CarlitoDon):
- Pure Data-Driven Dynamic Tier Router (DynamicTierRouter): Fully decoupled from model-specific heuristics; implemented
collect_tiers_for_baseinOfficialModelCatalogto dynamically discover all available tier suffixes for any base model from the official upstream catalog. Whether for Gemini 3.x Flash, Claude 5.5, or future models, automatically strips suffixes to derive bare models and routes requests dynamically based on tier weight hierarchies (lite < low < default < medium < high < xhigh < max) and client effort hints. - Claude 5.5 Model Specs & Factory: Fully transitioned to dynamic
RealModelSpecconstruction, dynamically fetchingmaxOutputTokens(128k) and thinking budgets fromOfficialModelCatalogwithout static constant maintenance. - Permission Gate & Soft-Sorting: Enforced semver-based subscription gating in
TokenManager(Claude >= 5.0) to filter out Free accounts, giving scheduling priority to PRO/ULTRA accounts that have synchronized the model catalog. - Thinking Level Normalization: Added support for snake_case levels such as
extra_low,x_high, andflash_lite, seamlessly normalizing them with kebab-case counterparts. - Unified Thinking Budget & Default Alignment: Unified Gemini and Claude thinking configuration, aligning the default budget to -1 (representing native adaptive defaults) and simplifying settings panels.
- Pure Data-Driven Dynamic Tier Router (DynamicTierRouter): Fully decoupled from model-specific heuristics; implemented
- [Long-Generation & Deep-Thinking Stream Interruption Fix] Relax HTTP/2 Keep-Alive Timeout to 90s and Fix Claude Peek Error Detection (Fixes #3593, Thanks to @Hubitski):
- HTTP/2 Keep-Alive Timeout Relaxed: Extended
keep_alive_timeoutinUpstreamClientto 90s (aligned withpool_idle_timeout) and adjusted probe interval to 5s, completely resolving 10s client connection drops during deep thinking or long generation pauses (Stream interrupted before completion/error reading a body from connection). - Claude Streaming Peek Error Detection: Completed error event detection (
claude_stream_chunk_has_error_event) during the peek pre-read phase inclaude.rs, preventing stream head errors from being mistaken for valid responses and bypassing account pool rotation.
- HTTP/2 Keep-Alive Timeout Relaxed: Extended
- [通用分档路由器与 Claude 5.5 动态规格架构升级] 纯数据驱动的通用 DynamicTierRouter、可用档位自动扫描与 Claude 5.5 全链路接入 (PR #3594 启发推进, Thanks to @CarlitoDon):
- 纯数据驱动通用分档路由器 (DynamicTierRouter): 从硬编码特判中彻底解耦,在
OfficialModelCatalog中实现collect_tiers_for_base,支持从官方目录动态探测任意模型的所有可用分档后缀(available_tiers 列表)。无论是 Gemini 3.x Flash 还是 Claude 5.5 或未来的新模型,自动剥离后缀派生裸模型并根据权重梯队(lite < low < default < medium < high < xhigh < max)与客户端 effort 动态智能路由。 - Claude 5.5 模型规格与动态工厂: 全面对接
RealModelSpec动态构建,优先从OfficialModelCatalog动态挂载maxOutputTokens(128k)与思考预算,彻底告别静态常量维护。 - 权限门禁与优先调度: 在
TokenManager中基于语义版本(Claude >= 5.0)自动拦截 Free 账号,优先调度已同步该模型目录的 PRO/ULTRA 账号。 - 思考等级归一化扩展: 支持
extra_low、x_high、flash_lite等下划线蛇形命名,与连字符形态无缝兼容归一化。 - 思考预算统一精简与默认值对齐: 统合 Gemini 与 Claude 思考配置,默认预算统一为 -1(语义为遵循官方默认自适应档位预算),精简冗余分区。
- 纯数据驱动通用分档路由器 (DynamicTierRouter): 从硬编码特判中彻底解耦,在
- [长生成与深度思考断流根治及流式错误识别] 放宽 HTTP/2 PING 超时至 90s 并修复 Claude Peek 错误检测 (Fixes #3593, Thanks to @Hubitski):
- HTTP/2 保持长连接与超时放宽: 将
UpstreamClient中底层 HTTP/2keep_alive_timeout放宽至 90s(与pool_idle_timeout对齐),探测间隔优化为 5s,彻底根除模型在深度思考或超长生成静默期时底层连接在 10s 被主动掐断断流的报错(Stream interrupted before completion/error reading a body from connection)。 - Claude 流式 Peek 错误检测补齐: 补齐
claude.rs的 Peek 预读阶段错误事件识别逻辑(claude_stream_chunk_has_error_event),防止流式首包错误被误判为合法正文而绕过账号池轮换重试机制。
- HTTP/2 保持长连接与超时放宽: 将
Full Changelog: v4.9.2-beta.4...v4.9.2-beta.5