- [Thinking Budget Root-Cause Fix & Model Suffix Absolute Priority] Eliminate 32768 Overwrite, Enforce Named Suffix Priority & Migrate Legacy Configs (PR #3611, Fixes #3610):
- Enforce Named Model Suffix Absolute Priority: Established absolute priority for explicit suffixes (
-low,-medium,-high) inresolve_custom_budget, restricting clienteffortmapping strictly to bare models (is_bare), preventing named models (such asgemini-3.8-flash-medium4000,gemini-3.8-flash-low1000) from being overridden to inflated 32768 budgets by clientclient_effort: high. (Thanks to @cubelikeplayDaniel) - Restore Default Mode Fallback Branch: Added missing
flash_mode == ThinkingBudgetMode::Defaultinterception for named non-tiered models, allowing gateway default mode to seamlessly fall back to official catalog preset values (Medium 4000, Low 1000, High -1). (Thanks to @cubelikeplayDaniel) - Decouple Claude Adapter (Pipeline First): Removed
.or_else(|| tb_config.effort.as_ref())global fallback from the Claude protocol adapter, ensuring adapters purely pass through client parameters and leaving arbitration entirely to the pipeline. (Thanks to @cubelikeplayDaniel) - Converge Inbound Capacity Expansion: Eliminated two premature and redundant
min_overhead(+8192) expansions from early branching, converging all headroom coordination into the single-point terminal invariant state machine. (Thanks to @cubelikeplayDaniel) - Smooth Legacy 32k Config Migration: Introduced
thinking_budget_32k_legacy_migratedstartup migration to reset legacy default values (32768/16384) back to official adaptive-1without affecting custom configurations. (Thanks to @cubelikeplayDaniel)
- Enforce Named Model Suffix Absolute Priority: Established absolute priority for explicit suffixes (
- [网关思考预算根治与模型后缀绝对优先级保障] 根治思考预算无差别覆盖 32768、确立具名后缀最高优先级并平滑迁移旧版配置 (PR #3611, Fixes #3610):
- 确立模型后缀绝对最高优先级: 在
resolve_custom_budget中建立显式后缀(-low、-medium、-high)绝对优先级判定,仅当纯净无后缀裸模型(is_bare)时才允许客户端effort映射,彻底阻断具名中低档位模型(如gemini-3.8-flash-medium4000、gemini-3.8-flash-low1000)被客户端client_effort: high越权覆盖为 32768 满血预算。 (Thanks to @cubelikeplayDaniel) - 补齐 Default 模式回退分支: 为具名非 Tiered 模型补齐
flash_mode == ThinkingBudgetMode::Default拦截分支,使网关默认模式能无损回退到官方模型结构体权威默认值(如 Medium 4000、Low 1000、High -1)。 (Thanks to @cubelikeplayDaniel) - 落实 Pipeline First 协议适配层解耦: 彻底移除 Claude 协议适配器中
.or_else(|| tb_config.effort.as_ref())全局配置脑补行为,适配器仅纯粹透传客户端原始参数,统一由流水线集中仲裁。 (Thanks to @cubelikeplayDaniel) - 收敛 Inbound 流水线上限扩充逻辑: 移除模式分流中两处超前且重复的
min_overhead(+8192) 强制扩充,所有预算与容量协商统一收敛至 Inbound 尾部的「终审上限保护与不变量协调状态机」单点裁决。 (Thanks to @cubelikeplayDaniel) - 旧版出厂脏配置平滑迁移: 新增
thinking_budget_32k_legacy_migrated启动迁移逻辑,在应用启动时自动将 Default 模式下残留的历史 32768/16384 脏数据重置为官方自适应值-1,彻底消除历史遗留锁定。 (Thanks to @cubelikeplayDaniel)
- 确立模型后缀绝对最高优先级: 在
What's Changed
- fix(pipeline): 彻底根治思考预算统一覆盖为32768根因并保证后缀与默认模式优先级 (#3610) by @cubelikeplayDaniel in #3611
Full Changelog: v4.9.5-beta.3...v4.9.5-beta.4