Bifrost HTTP Transport Release v2.2.6
✨ Features
- OSS Management API Setup Lock - While dashboard auth is not active (no admin account, or auth disabled), every non-public
/apicall on OSS Bifrost now needs the setup token, sent as theX-Bifrost-Setup-Tokenheader or as thebifrost_setup_sessioncookie the dashboard gets fromPOST /api/session/setup. A missing token returns401. A wrong token, or no token set on the server, returns403./health,/api/version, the session login routes,/.well-known/*and whitelisted routes stay public. The lock lifts as soon as an enabled admin is saved (#8010)
Migration: setsetup_tokenin config.json (orBIFROST_SETUP_TOKEN) and restart. Then either enable dashboard auth, or sendX-Bifrost-Setup-Tokenfrom scripts and API clients that call/apiwith auth off. Enterprise is not affected by the lock.
- Inference Auth On by Default -
enforce_auth_on_inferencenow defaults totruefor fresh deployments when config.json leaves it out, file-only deployments included. Creating the first enabled admin also turns inference auth on unless the request sets it explicitly. Inference without a credential then returns401. A stored database value always wins, and an explicitfalseis always kept. This also applies to Bifrost Enterprise (#8010, #7864)
Migration: to keep unauthenticated inference, set"enforce_auth_on_inference": falsein config.json, or send it explicitly when creating the first admin.
- OAuth Discovery Requires issuer_url - When
mcp_server_auth_modeisoauthorboth,oauth2_server_config.issuer_urlmust be set to a non-empty value. Config load fails andPUT /api/configrejects the save otherwise. The issuer is never derived from the requestHostheader any more, and discovery responses carryCache-Control: no-store(#7863)
Migration: setoauth2_server_config.issuer_url(env syntaxenv.MY_VARworks) before upgrading any deployment that has MCP OAuth discovery enabled, or the server will not start.
- Provider Dial Target and Proxy Changes Need Real Auth - With dashboard auth off, these changes now return
403unless the request carries a genuine admin credential (a session, or the OSS setup token): providerbase_url,allow_private_network, key endpoint URLs (Ollama, SGL, vLLM, Azure, etc.), provider proxy, custom CA certs, skipped TLS verification, absoluterequest_path_overrides, and enabling the global proxy with a URL (#7865, #7867) - Private Framework Config URLs Rejected -
pricing_url,model_parameters_urlandmcp_library_urlare now checked at save time and at dial time against private, link-local and CGNAT addresses, and redirects are not followed. Air-gapped setups should usefile://URLs (#7861) - Semantic Cache Scoped per Virtual Key - Cache buckets are now partitioned by virtual key, so a shared
cache_keyordefault_cache_keyno longer serves one virtual key's cached response to another. Entries written before the upgrade under a virtual key are not reused, so expect a cold cache. Unscoped cache keys that start withvk:are moved toraw:vk:. A per-request threshold override can only raise the configured threshold, and is capped at 1.0 (#7862) - zstd Decoder Window Cap - zstd-compressed bodies whose frame header asks for a window above 100 MiB are rejected before any allocation (#7859)
- Dashboard Setup Session - The login page has a new setup screen that trades the setup token for a 12-hour
HttpOnly; SameSite=Strictcookie. The cookie is signed with a key derived from the token, so the browser never stores the token itself. A sidebar card flags missing dashboard auth until an admin exists.GET /api/session/is-auth-enablednow reportssetup_requiredandsetup_token_configured, andPUT /api/configaccepts a setup-token request as first-admin proof (#8010) - Compat: Clamp Over-Limit Output Tokens - With
should_convert_paramson (UI: Convert Unsupported Param Values, orx-bf-compat: ["should_convert_params"]), amax_output_tokens,max_completion_tokensormax_tokensabove the model'smax_output_tokensin the model catalog is lowered to that limit instead of being rejected by the provider. This works for every provider whose model has a limit in the catalog. A thinking budget at or above the lowered cap is moved just below it. Values are never raised, and models with no catalog limit are left alone. Each change is logged as a warning on the request. The setting did nothing before this release. A config.json whoseclient_confighas nocompatblock turns it on by default - Datasheet Control for Per-Message Effort - Per-message effort support can now be set per provider and model with the datasheet field
supports_mid_conversation_output_config, so a surface that ships the feature can be enabled without a release. With no datasheet value the current behaviour applies (Anthropic direct on Fable 5.1, Opus 5+ and Sonnet 5.5). On a provider other than Anthropic, also allow the beta header withbeta_header_overrides: {"mid-conversation-output-config-": true}in that provider's network config - Per-Message Effort Override - A per-turn effort override sent as an effort-only system message (
{"role":"system","content":[],"output_config":{"effort":"low"}}) now reaches Anthropic instead of being dropped, with themid-conversation-output-config-2026-07-01beta added. Models without per-turn effort, and OpenAI-shaped providers, drop it instead of returning an error (#7714) - Claude Code Per-Message Effort on Vertex and Other Cloud Surfaces - Claude Code requests to Opus 5.5, Fable 5.1 and Sonnet 5.5 on Vertex no longer fail with
messages.1.output_config: Extra inputs are not permitted. The per-messageoutput_configis now removed for every provider and model without per-message effort (Vertex, Bedrock, Bedrock Mantle, Azure, DeepSeek, Fireworks, vLLM, SGL). The system message text and the top-level effort are kept
🐞 Fixed
- Claude Tool-Call Argument Streaming on Bedrock and Vertex - Claude tool arguments now stream incrementally, so a long Write call no longer arrives as one burst after minutes of silence and Claude Code no longer aborts with "Stream idle timeout".
eager_input_streamingdefaults on for custom tools that leave it unset (every Claude model on Vertex; Sonnet 4.6, Sonnet 5+, Opus 4.7+ and Fable on Bedrock). An explicitfalseis kept. Converse carriesfine-grained-tool-streaming-2025-05-14inadditionalModelRequestFields(#8009) - Tool-Result Cache Markers for gpt-5.6+ - Anthropic
cache_controlmarkers on tool results (as Claude Code sends them) now becomeprompt_cache_breakpointon both Chat Completions and Responses for gpt-5.6+ on OpenAI, Azure, Bedrock and Bedrock Mantle, so caching keeps advancing past the first tool turn. Responses also marksinput_imageandinput_fileparts. At most four breakpoints are kept, the latest ones (#8012) - Handler Panic Recovery - A panic in a request handler now returns a
500and logs a server-side stack trace instead of crashing the process (#7866) - Secret Redaction in Config Responses - Env and vault-resolved values in provider alias configs (region, project ID, Azure endpoint, Vertex project, Bedrock inference profile ARN) and in Bedrock endpoint overrides are now masked in management API responses (#7858)
- Admin Password Autofill - The Security page's admin fields carry
autoComplete="username"and"new-password", so password managers no longer fill a saved host password into the new admin's password field (#8010)
🗄️ Database Migrations
- No new database migrations in this release.
Installation
Docker
docker run -p 8080:8080 maximhq/bifrost:v2.2.6Binary Download
npx @maximhq/bifrost --transport-version v2.2.6Docker Images
maximhq/bifrost:v2.2.6- This specific versionmaximhq/bifrost:latest- Latest version (updated with this release)
This release was automatically created with dependencies: core v1.11.2, framework v1.8.1. All plugins have been validated and updated.