Bifrost HTTP Transport Release v2.2.0
✨ Features
- Claude Desktop and Cowork Marketplace - Skills stored in Bifrost can now be registered as a marketplace in Claude Desktop and Cowork, which reject the direct JSON URL and require a cloneable Git repository URL. A new
/api/skills/serve/claude-code.gitendpoint implements the two Git smart-HTTP requests used during a clone and serves a repository containing.claude-plugin/marketplace.json; the existing Claude Code flow is unchanged (#7152) - Time-of-Day Peak and Off-Peak Pricing - Model pricing accepts
off_peak_cost_multiplierand apeak_hoursweekly schedule, so providers like DeepSeek that bill the same model at two rates are costed correctly. Base rates are treated as peak; the multiplier scales usage-based charges outside the declared windows. Flat per-request fees, per-search-query fees and guardrail/MCPAdditionalCostare never discounted. Windows use IANA timezones, weekday numbers and half-openHH:MMintervals that may wrap past midnight, and both fields are editable from the custom pricing override sheet (#6574, #6575, #6576, #6577, #6578, #6579, #7054) - GA Realtime Transcription - OpenAI and Azure GA transcription sessions are served over both WebSocket and WebRTC with normal Bifrost authentication, routing, governance, guardrails, logging and transcription-aware pricing. These sessions carry only
intent=transcriptionon the connection and deliver the routing model later insession.update(or in the initial multipart/v1/realtime/callsrequest for WebRTC), so Bifrost now routes on the nested transcription model while preserving realtime connection and turn semantics (#7089) - Regex Model Allow and Block Lists -
allowed_models/blacklisted_modelson virtual keys andmodels/blacklisted_modelson provider keys acceptregex:<pattern>entries next to exact names. Patterns are compiled once as case-insensitive full matches, a pattern that is empty,*or does not compile is refused with 400, and list-models never surfaces a pattern as a model. Provider-key create and update now validatemodelsthe same way asblacklisted_models. This supersedes the separate*_patternsfields, which were added and then withdrawn before release (#6987, #6988, #6989, #6990, #7031, #7133, #7134) - Governance Entity Names on MCP Tool Logs -
mcp_tool_logsgains the same attribution shape thelogstable has:user_name,team_name,customer_nameandbusiness_unit_namebecome real columns instead of transients, the multi-valuedteam_ids/team_names,customer_ids/customer_namesandbusiness_unit_ids/business_unit_namessets are stored as index-aligned JSON arrays, andbudget_idsandrate_limit_idsare recorded. Names are written from the request context at ingestion, with nothing resolved on read, so the dashboard stops rendering raw UUIDs (#7154) - Endpoint-Attributed MCP Inspections - Inspected MCP tool calls are logged with bounded identity sourced from the gateway rather than payload-supplied headers. A new
MCPObservationcarries device, app key, server label, tool name and decision onto both the pending and final log entry, and the MCP logs view falls back toapp_keywhenappis absent so endpoint-attributed rows show the right app icon and name (#6959) - Virtual MCP References by Name - Access profiles and governance projects reference Virtual MCPs through
virtual_mcp_name, making config files portable across environments; names resolve on startup and a name matching nothing is refused.mcp_configs({ mcp_client_id, tools_to_execute }) replaces themcp_servers/mcp_tool_overridesinclude-exclude model with a single allowlist, where["*"]grants all tools including future ones and[]grants none.virtual_mcp_idand the old keys are deprecated, still accepted, and folded into the new shape at load time (#7181) - Normalized
error_typeMetric Label -bifrost_error_requests_totalgains anerror_typelabel with a closed, prefix-structured vocabulary (caller_*,policy_*,provider_*,bifrost_*,_OTHER) so a 429 from a governance rate limit is distinguishable from a 429 from an upstream, and a 503 Bifrost shed under queue pressure from an upstream overload. Classification resolves a declaredExtraFields.ErrorTypefirst, then Bifrost's own markers, then the status code; it deliberately ignores providererror.typestrings, which disagree across providers for the same condition (#7141) - Bedrock OpenAI-Compatible Endpoint Routing - A
use_openai_endpointsflag on Bedrock keys and aliases routes chat completions and responses through Bedrock's/openai/v1surface instead of Converse, for models that support it. It is opt-in by design: Converse carries Bedrock Guardrails,performanceConfigandrequestMetadatathat the OpenAI-compatible surface silently ignores, so diverting automatically could stop a guardrail from being enforced with no visible error. The alias value wins over the key, matchinguse_anthropic_endpointsprecedence (#7071, #7073) - Anthropic Tool Search on Bedrock Claude - Anthropic tool search (
tool_search_tool_*,defer_loading) is served onbedrock/Claude models by routing those requests to InvokeModel / InvokeModelWithResponseStream, the only Bedrock API AWS allows it on; CountTokens counts such requests with the same InvokeModel body. Server-side tool search also survives the Bedrock-native invoke ingress end to end: the tool is carried as an ingress-only marker so the egress predicate can see it, results are returned as aserver_tool_useplustool_search_tool_resultpair rather than a clienttool_use(which the API rejects when echoed back), the streaming path emits the same pair, replayed blocks round-trip unchanged, and the Anthropic-native response path carries them too (#6900, #6908, #7162, #7163, #7164, #7165, #7166) - Namespace Tool Support Across Providers - Responses
namespacetools are flattened in core for every provider whose wire lacks the type, with nested functions renamed to<namespace>__<function>so two namespaces sharing a function name no longer collide into an upstreamTool names must be unique400. Returnedfunction_callitems map back to the bare name plus namespace, prior-turn calls andtool_choicenames are re-aliased, and a still-duplicate name is rejected with a clear 400 before reaching the provider. Flattened names honour each wire's documented tool-name limit, overridable per model throughtool_name_max_length. The names a provider reserves for its own server tools come from the datasheet rowreserved_tool_namespaces, and Codex's literalfunctionsnamespace is unwrapped to top-level tools for every provider (#7039, #7082, #7084, #7161) - Trusted Networks for the SSRF Guard - A
trusted_networkslist of IP/CIDR entries the SSRF guard consults before outbound discovery calls, so a self-hosted IdP on an internal network can be reached by the generic provider's discover-endpoints and discover-claims flows. Declaring the key inconfig.jsonmakes it own the whole list, an explicit empty array clears dashboard-added entries, and omitting it leaves the stored allowlist untouched. Hostnames are refused, since DNS would then decide which requests bypass SSRF protection (#7081) - Prompt Cache Reload Through the Server -
ReloadPromptCachemoves ontoServerCallbacksso enterprise can gossip it across nodes. The prompts plugin's in-memory index was previously rebuilt only in the process that served the write, so on a multi-node deployment a prompt published on node A left node B resolvingx-bf-prompt-id/x-bf-prompt-versionagainst a stale index until restart: an unknown version errored andlatestserved the old content. OSS behaviour is unchanged (#7061) - Guardrail Tool-Call Argument Redaction - Guardrail redaction covers LLM tool-call arguments (Chat function arguments, Responses function arguments and custom-tool input) across the Anthropic streaming and non-streaming paths, reading and writing
delta.partial_jsononinput_json_deltaevents and collecting string-valued paths inside atool_useblock'sinputwithout touching tool names, IDs or definitions. A separate identity-based transformer path lets provider-managed rewrites (Model Armor, Bedrock) land in the correct native JSON field even when the same text appears in several fields, verifyingOriginalbefore patching and the written value after (#6977, #7049) - Regions and Service URLs in Plaintext - Regions and service URLs (Azure endpoint, Vertex/Bedrock/Bedrock Mantle region, vLLM/Ollama/SGL/Databricks URL, MCP connection string) are public identifiers, not credentials, and were being unconditionally redacted into unreadable values in the UI. A new
SecretVar.RedactedIfSecret()returns a plain clone for a literal value and still masks anything sourced from an env var or vault reference (#7085) wait_for_usagefor Custom Providers - Await_for_usageflag oncustom_provider_configtells Bifrost the upstream sends a trailing usage-only frame, so the read loop holds open pastfinish_reasonuntil it arrives instead of synthesizing a zero-usage terminal chunk. Termination stays bounded by the usage chunk, two consecutive post-finish heartbeat comments, EOF, orstream_idle_timeout_in_seconds(#7187)- Pinnable Log Search Mode - The logs search box gains a mode dropdown (Auto, Content, Request ID). Auto-detection treated UUID-shaped input as an ID lookup and everything else as a content scan, which breaks for request IDs that are not UUID-shaped and for UUID-shaped strings that should be searched as content. A pinned mode bypasses all sniffing and re-runs the current input immediately (#7149)
- MCP Usage Guide Auth Methods - The MCP usage guide generates client configs for Virtual key, OAuth and Identity provider authentication instead of requiring a virtual key for every config. Credential resolution is centralized in
buildMCPHeaders(), so OAuth emits no headers, identity provider emits aBearerplaceholder, and virtual key keepsx-bf-vk(#7111) - Chart Color System - Dashboard charts, status badges and components read a structured set of CSS custom properties instead of hard-coded hex values, so colors adapt correctly between light and dark themes. Tokens are grouped as semantic (hues 0 to 70 reserved so no category can look like an error), sequential, ordinal for percentile series, and categorical at matched chroma assigned by rank (#7113)
🐞 Fixed
- Client Disconnect Not Cancelling Requests - A client that closes its socket while Bifrost is still waiting on core (silent upstream header wait, retry backoff) now cancels the request.
ConvertToBifrostContextstarts a socket watcher that peeks the client connection every 500 ms withMSG_PEEKand cancels the context on FIN or RST, so upstream retries stop as soon as nobody is listening; previously fasthttp offered no per-requestDoneand a disconnect was only noticed when an SSE write failed. No-op on non-unix platforms and on in-memory test connections (#7035, #7106) - Silent Upstream Never Timed Out -
default_request_timeout_in_secondsnow bounds the wait for response headers on streaming requests, and cancelling a request closes the upstream socket. Every fasthttp client runs through a Bifrost-ownedRoundTripperthat applies the client's read/write timeouts and the request context to the request write and the header wait, then lifts the socket deadline once headers are parsed sostream_idle_timeout_in_secondsremains the only bound on the body. An upstream that accepts the connection and never answers now fails with 504RequestTimedOutand its fallbacks are used, instead of pinning the provider worker. Unary large-response downloads bound every body read the same way, gzip-encoded bodies are classified for large-response mode by decompressed size, and a mid-body connection drop is again reported as the retryable 502 completion-marker error instead of a generic unexpected EOF (#7034, #7104) - Retry Storm After Client Disconnect - fasthttp-level stale-connection retries no longer multiply
max_retries.contextTransport.RoundTripreports a pre-header failure on a freshly dialed socket withretry=false, soStaleConnectionRetryIfErronly walks past pooled keep-alive sockets the upstream closed while idle; an upstream that closes a fresh connection without answering now costs exactly one attempt instead of up to four. Retry backoff also ends as soon as the request context is cancelled, freeing the worker immediately instead of after up toretry_backoff_max(#7035, #7105) - Abandoned Request Billing Coin Flip - Non-streaming requests whose caller had already disconnected were billed and logged only about half the time. The worker's delivery
selecthad two simultaneously ready cases, a send into a cap-1 channel andctx.Done(), and Go picks uniformly among ready cases, so terminal post-hooks were skipped roughly 50% of the time. The worker now checksreq.Context.Err()before the select and callsbillAbandonedTerminaldeterministically, keeping the 5-second timer guard for a caller that leaves between the check and the send (#6972, #7116) - Stream Never Terminated Without
[DONE]- An OpenAI-compatible upstream that omits[DONE]and then goes silent afterfinish_reasonno longer fails the stream whenstream_idle_timeout_in_secondsfires. The chat and text completion read loops treat an idle timeout after a terminal signal as a parked upstream, abandon the connection rather than drain it, and synthesize the final chunk with the bufferedfinish_reason; a stall beforefinish_reasonstill surfaces as the idle-timeout error (#7108, #7115) - Dropped SSE Frames in
raw_response- Role-only, finish-only and usage-only frames never entered the chunk-forwarding branch, so their bytes were discarded from the reconstructedraw_response, leaving the captured audit trail irreconcilable against a provider invoice since the usage frame carries the token counts Bifrost bills from. ApendingRawFramesbuffer drains onto the next forwarded chunk or the synthetic terminal chunk, the Responses-over-Chat fallback no longer stamps one upstream frame onto every derived event, anddelta.refusalanddelta.annotationsare forwarded instead of dropped entirely (#7144, #7184) - Bedrock Mantle Trailing Usage - Bedrock Mantle chat streaming no longer drops the usage-only chunk that arrives after
finish_reason;ProviderSendsDoneMarkernow treatsbedrock_mantleand the legacy Mantle route under thebedrockkey as sending[DONE], so streamed usage and cost are recorded (#7065, #7076) - Fallbacks Re-Ran the Primary for Image and Video Edits - Fallbacks for image edit, image variation and video edit requests now reach the configured fallback provider and model.
prepareFallbackRequesthad no arm for those three types, so the shallow copy kept the primary's sub-request pointer and the attempt was routed back to the primary while routing info, headers, the log row and thefallback_indexmetric label all reported it as a fallback. The helper now verifies the prepared request targets the fallback and skips it with a warning otherwise, so a future request type added without an arm fails loudly (#6966, #7118) - Provider Response Headers Leaked Across Fallbacks -
clearCtxForFallbackmissedBifrostContextKeyProviderResponseHeaders. Providers set that key before the status check so error paths can forward it, and when a fallback failed pre-flight (key selection failed, a plugin short-circuited, the queue was retiring) nothing overwrote it, so the client received a response attributed to provider B carrying provider A'sRetry-Afterand rate-limit headers (#6973, #7021) (thanks @Huang-404-Q!) - Credential-Bearing Response Headers Forwarded - The provider-response extractors filtered only against a fixed map of 28 exact names, so any credential-named header outside it was re-served to the inference caller in both the HTTP response and
extra_fields.provider_response_headers. The path now also consultsschemas.IsSensitiveHeader, which recognizes credential names by substring and suffix and already knew aboutcf-access-*andx-amzn-oidc-*; a fixed list cannot enumerate the space whennetwork_config.extra_headersexists to carry custom auth headers and some upstreams echo request headers back (#7120, #7121) (thanks @Atharva-Kanherkar!) - Nil Dereference on Incomplete Fallback Errors - A plugin returning a
BifrostErrorwhose nestedErrorfield is nil crashed the request worker. Fallback processing now nil-checks the error and guards access toError.Type, using the nil-safeGetErrorString()helper, and continues to the next fallback when allowed (#6967, #7110) (thanks @Constantine3!) - Bedrock Duplicate Document Names - Every untitled document block was given the literal default name
document, and Converse rejects duplicate document names, so any request with two or more untitled documents failed unconditionally with aValidationException. A per-request document namer now disambiguates with numeric suffixes (document,document-2, and so on) and suffixes titled documents only on an actual collision, on both the Converse and Responses replay paths (#7003, #7027) (thanks @Huang-404-Q!) - Bedrock Text Document Source - Converse rejects document blocks that use a text-only
DocumentSourceunless citations are explicitly enabled, so plain text formats (text/plain,text/markdown,text/csv,text/html) failed withmust set one of the following keys: bytes, s3Location. All document content, including data URLs, percent-encoded payloads andfile_data, now ships base64-encoded throughsource.bytes(#7072, #7079) - Bedrock Tool Result Images - Some Bedrock-hosted models (OpenAI and Grok families) reject image blocks placed directly inside a
toolResultin Converse, even though they accept images in tool output via Responses. AhoistToolResultImagespass moves images out oftoolResultblocks and re-inserts them after the last tool result in the same message, leaving a placeholder text block so the now image-free result is not rejected for being empty. A datasheet rowsupports_converse_tool_result_imagesoverrides the family-level default (#7150) - gpt-oss Message Mistagging on Mantle - Claude Code replaying a prior assistant message through
POST /anthropic/v1/messagesat a Bedrock Mantle model was rejected with hundreds of validation errors: the Bedrock-grouped ingress converter tagged user and system input text asoutput_text(onlyinput_textis valid on input messages) and omitted the requiredstatuson replayed assistant messages. Bedrock requests with no explicitmax_tokensnow also populate it from the model's known capacity instead of truncating silently (#7074, #7075) - Azure Foundry Output Token Cap - Azure Foundry deployments of Fireworks-hosted models were silently capped at 4096 output tokens on
/openai/v1/responsesbecause Microsoft routes those models through chat completions internally. The Azure provider now checks the model's datasheetsupported_endpointsand transparently serves bothResponsesandResponsesStreamthrough/openai/v1/chat/completionswhen/v1/responsesis absent. Separately, a turn truncated by the output-token cap on any OpenAI-shaped Responses provider now reportsstop_reason: max_tokenson the Anthropic egress instead of hiding the truncation asend_turn(#6782, #7142) - Gemini Inline Image and Audio Dropped - Gemini image-generation output (
inlineData) was silently dropped on/v1/chat/completions, both unary and streaming (#7032, #7033) (thanks @Atharva-Kanherkar!) - Gemini Image Edit Misclassified -
isImageEditRequestonly checkedcontents[0].parts[0], so a request with the prompt text before theinlineDatapart, which is the ordering in Google's own REST edit sample, was misclassified as image generation: the image never reached Vertex and the model invented a picture from the prompt alone, returning HTTP 200. Detection now scans all parts across all contents.imageConfig.aspectRatiois also preserved as a typed parameter instead of being folded into aWxHsize string that collapsed any unsupported ratio to1:1(#7173) - Gemini Per-Part Media Resolution Dropped -
Partwas missingmediaResolution, which overridesgenerationConfig.mediaResolutionfor a single part, and becausePart.UnmarshalJSONdecodes into a closed alias the key was discarded before any conversion ran. Per-part image and PDF tokenization fell back to the model default, so anULTRA_HIGHimage billed about 21k prompt tokens through/genaiinstead of about 22.1k direct. The field now round-trips both spellings end to end and is stripped on the OpenAI wire path where it is unknown (#7156) - Gemini
generationConfigLost Across Retries -convertParamsToGenerationConfigResponsesdeletedtop_k,frequency_penalty,presence_penalty,stop_sequencesandmedia_resolutionfromExtraParamswhile mapping them intogenerationConfig, and that conversion runs once per attempt on the same request, so every retry or fallback after the first was sent without them. On Vertex the only visible signal waspromptTokenCounthalving on retried requests, making the downgrade silent (#7138) (thanks @VictorRequenaMaisa!) - MCP Tool Schema Property Order - MCP tool schemas keep one property and
$defsorder across tool syncs, so prompt caching is not invalidated by reordering alone (#7169, #7170) (thanks @dougcalobrisi!) - vLLM Alias Resolution During Key Selection - The generic allow and block lists correctly evaluated the user-facing alias, but vLLM key selection compared each key's physical
model_nameagainst the unresolved alias and rejected valid keys. The alias is now resolved per key before that comparison, so the same public alias can map to different physical model IDs across vLLM instances, while allow and block checks keep matching the original alias. Allowed Models, Blocked Models and Deployments/Aliases are now exposed on the vLLM key form (#6956) (thanks @Constantine3!) - opencode-zen Responses Routing - opencode-zen Responses calls are routed through
/v1/chat/completions, which its upstream serves, instead of/v1/responses, which it does not (#6778, #6819) (thanks @miguelchico!) - DeepSeek
max_completion_tokensIgnored - DeepSeek's chat-completions endpoint only recognizes the legacymax_tokensfield and silently ignoresmax_completion_tokens, so the limit had no effect. The field is now remapped on the wire, matching the behaviour already in place for Opencode and Ollama (#7131) - Anthropic Server Tools on Third-Party Endpoints - Fireworks' Anthropic-compatible endpoint returns 400 when a request includes Anthropic server tools such as
web_search_20250305, because those run on Anthropic-operated infrastructure that does not exist on third-party hosts; clients whose built-in web search is always on hit this on every request. Unsupported server tools are now dropped before the request leaves Bifrost, the caller's function tools are kept, and the drops are reported onDroppedUnsupportedToolsinstead of failing the call. The same applies to vLLM and SGLang (#7090) - Bedrock Guardrail Headers - Bedrock's OpenAI-compatible endpoints apply guardrails through request headers rather than the
guardrailConfigbody field Converse uses, so a configured guardrail was silently ignored there.guardrailIdentifier,guardrailVersionandtraceare now mapped to theX-Amzn-Bedrock-Guardrail*headers on both chat completions and responses, streaming and non-streaming, and the key is consumed so it is not also emitted into the body. A half-formed config with only one of identifier or version is left untouched rather than sent (#7095) - Responses SSE
item: null- Responses stream events that carry no item payload no longer serialize"item": null, which strict OpenAI Responses clients reject as an invalid frame, breaking streamed/v1/responsesusage (#6395) (thanks @ReStranger!) - Responses
actionString Decode -image_generation_callitems where OpenAI emitsactionas a bare JSON string failed to decode, becauseUnmarshalJSONimmediately peeked at a.typefield that cannot be read from a string. That silently dropped theresponse.output_item.doneandresponse.completedevents carrying the image, leaving the stream without a terminal event and surfacing as a bogus "provider closed the stream" truncation error. The completed item also keeps the generation settings OpenAI echoes back (#7060) - Mid-Conversation System Messages Broke Prompt Caching - Mid-conversation
role: "system"messages are inlined in place as<system-reminder>user turns on every converter with a top-level system field: Bedrock Converse (Responses and Chat Completions), Gemini Chat Completions, and the Anthropic wire shape used by DeepSeek, Fireworks and SGL. Previously only Claude on Bedrock and Anthropic inlined; everything else hoisted each reminder into the top-level system block. Claude Code appends a trailing<total_tokens>reminder after every turn, so the hoisted block grew the front of the prompt each turn and prefix-based caches reported a full cache write and zero cache reads on every turn (#7145) - Unsupported
reasoning.contextRejected the Request - Areasoning.contextvalue the target model does not accept is dropped on the OpenAI and Azure Responses path, soall_turnson the original gpt-5 family including gpt-5-pro, gpt-5.1 to gpt-5.3 and the o-series runs under the model's owncurrent_turndefault instead of failing withUnsupported value. gpt-5.4, gpt-5.5 and gpt-5.6 keep it. Accepted values come from the datasheet rowsupported_reasoning_contexts(#7140) - ClickHouse Retention Filled Replica Disks - ClickHouse log store deletes no longer run as heavyweight
ALTER TABLE ... DELETEmutations. The retention cleaner issued one per 100 rows, each rewriting the whole current-month part, and the once-a-minute stale-processingsweeps issued one per table unconditionally. Every delete is now a single lightweightDELETE FROM ... WHEREper run, skipped when nothing matches. The table TTL derived fromlogs_store.retention_daysis reconciled on every startup with a metadata-onlyMODIFY TTL, so changing the value reaches existing tables;0leaves an existing TTL untouched (#7098, #7103) - Governance Cleanup Dump Race -
UsageTracker.Cleanup()took its final budget and rate-limit snapshots before stopping the periodic reset worker, sotrackerCancel()could cancel an in-flight dump and fail withcontext canceled. The worker is now cancelled and awaited before the final dumps, a queued ticker event cannot start another reset cycle during shutdown, andcontext.Canceledis treated as expected only when the tracker context was actually cancelled (#7099, #7100) (thanks @Constantine3!) - OAuth Refresh Failed for Public Clients - Public OAuth2 clients registered against servers that only support
token_endpoint_auth_method: nonehave no client secret, and unconditionally settingclient_secret=in the refresh POST body sent an emptyclient_secret_postattempt. Strict authorization servers answeredinvalid_client, flipping the token row toneeds_reautheven though the refresh token was valid. The parameter is now omitted when the secret is empty, matching the PKCE code-exchange path (#7042) - Complexity Router Skipped Continuation Turns - When a turn is a continuation, such as a tool result following a prior user message, the complexity router discarded the extracted input and skipped classification entirely if no active session was found, leaving new or recovered sessions without a tier. Continuation turns now keep the populated
ComplexityInputand fall back to classifying the recoveredLastUserText; the skip path applies only when that is also empty (#7122) - Runtime Responses-Compat Routing - Bedrock runtime models that serve the Responses API are routed to it through a dedicated surface resolver rather than falling back to Converse (#7071)
- Bedrock Mantle Base Path - Bedrock Mantle serves each model on exactly one of two base paths (
v1oropenai/v1) and returns a 400 on the other. Hard-coded string matching in two packages covered only generations up to GPT-5 and Gemma 4, so GPT-6 and any future closed-generation model silently fell through to the wrong path. Resolution is centralized inResolveBedrockMantleBasePath, backed by abedrock_mantle_base_pathdatasheet field with family-name detection as the fallback, so a new generation needs a datasheet row rather than a code change (#7077) - Virtual Key
allowed_models: ["*"]Handling Reverted - The wildcard handling for governance virtual keys added in #6767 is reverted. Configurations relying onallowed_models: ["*"]with an empty synced catalog return to the previous behaviour (#7053) - Helm
perUserHeaderKeysNot Rendered -mcp.clientConfigs[].perUserHeaderKeysis mapped into the renderedconfig.json(#6033, #6034) (thanks @CallumWayve!) - Helm Plural Access Profiles - Governance roles accept
access_profilesas an array in the Helm and config schemas. Bifrost supports multiple access profiles but the schemas accepted only the deprecated singularaccess_profile. The singular form keeps rendering unchanged, the plural wins when both are present, and an explicitly empty plural list clears existing grants (#7044) (thanks @CarlosLanderas!) - Sidebar Title Overflow - Long sidebar item titles are truncated instead of overflowing (#7069)
- MCP Logs App Icon - App icons in the MCP logs table render at 20x20 and no longer shrink when the column is narrow (#7167)
🗄️ Database Migrations
- add_use_openai_endpoints_column - Adds the
use_openai_endpointscolumn to the provider keys table for Bedrock OpenAI-compatible endpoint routing. Reversible: the rollback drops the added column. Additive and nullable, so it is safe to run during a rolling upgrade. - add_time_of_day_pricing_columns - Adds
off_peak_cost_multiplierandpeak_hourstogovernance_model_pricingfor time-of-day pricing. Reversible: the rollback drops both added columns. Additive and nullable, so it is safe to run during a rolling upgrade. - mcp_tool_logs_add_governance_snapshots - Adds twelve governance attribution columns to
mcp_tool_logs:user_name,team_name,customer_name,business_unit_name, theteam_ids/team_names,customer_ids/customer_namesandbusiness_unit_ids/business_unit_namespairs, plusbudget_idsandrate_limit_ids. Reversible: the rollback drops all twelve in reverse order. Additive and nullable, so it is safe to run during a rolling upgrade. The twelveALTER TABLEs run under a bounded DDL lock wait, so startup does not stall behind a long-running log transaction holdingACCESS EXCLUSIVEon a continuously written table.
🐙 Closed GitHub Issues
- #6033 - Helm chart: mcp.clientConfigs[].perUserHeaderKeys not rendered into config.json
- #6778 - opencode-zen Anthropic endpoint fails, zen upstream doesn't support /v1/responses
- #6782 - Azure Fireworks/Foundry models capped at 4096 output tokens on Responses + Anthropic ingress (chat completions is not); truncation reported as end_turn
- #6825 - Bedrock provider silently drops Anthropic compaction (compact_20260112), capability matrix says supported, but Claude egress is Converse-only
- #6966 - fallbacks silently re-target the primary provider for image edit / variation requests
- #6967 - shouldContinueWithFallbacks nil-derefs BifrostError.Error, crashing the process on a plugin-returned error
- #6972 - abandoned-request billing is a ~50% coin flip when a client disconnects mid-request
- #6973 - provider response headers leak across fallback boundaries (clearCtxForFallback misses ProviderResponseHeaders)
- #7003 - Bedrock Converse assigns duplicate default name "document" to untitled document blocks, ValidationException
- #7032 - Gemini image-generation output (inlineData) silently dropped on /v1/chat/completions, both unary and streaming
- #7034 -
default_request_timeout_in_secondsandstream_idle_timeout_in_secondsdo not fire while waiting for response headers, a silent upstream blocks the request until the upstream closes, and the fallback is never used - #7035 - upstream retries continue after the client has disconnected, and go past
max_retries, one abandoned request produced 10 upstream attempts over ~20 minutes - #7048 - Compat namespace flattening creates duplicate tool names for DeepSeek 4.1 Flash
- #7065 - Bedrock Mantle chat streaming drops trailing usage after finish_reason
- #7072 - Bedrock Converse drops text-format document bytes, DocumentSource "must set one of the following keys: bytes, s3Location" (v1 to v2 regression)
- #7074 - openai.gpt-oss-120b via Bedrock Responses API mistags replayed history as output_text instead of input_text, breaks multi-turn Claude Code sessions
- #7098 - ClickHouse logs store: retention cleaner runs one
ALTER TABLE ... DELETEmutation per 100 rows and fills replica disks - #7099 - UsageTracker cleanup races the periodic rate-limit dump during shutdown
- #7108 - Custom-provider streaming never terminates when the upstream omits [DONE] (heartbeats mask stream_idle_timeout_in_seconds)
- #7120 - Provider response-header filter ignores IsSensitiveHeader, forwarding credential-named headers to inference callers
- #7143 - does_not_send_done_marker drops trailing Chat Completions usage and records zero cost
- #7144 - Chat Completions streaming raw_response omits usage-only and finish-only SSE frames
- #7155 - Bedrock-native invoke ingress silently drops Anthropic tool search (
tool_search_tool_*/defer_loading), served eagerly over Converse - #7169 - MCP tool schema property order changes between tool syncs, breaking prompt caching
Installation
Docker
docker run -p 8080:8080 maximhq/bifrost:v2.2.0Binary Download
npx @maximhq/bifrost --transport-version v2.2.0Docker Images
maximhq/bifrost:v2.2.0- This specific versionmaximhq/bifrost:latest- Latest version (updated with this release)
This release was automatically created with dependencies: core v1.9.0, framework v1.7.0. All plugins have been validated and updated.