What's Changed
⚠️ Compatibility Notes
- Require Python 3.11 or newer in every package; Python 3.10 installs resolve to 2.54.0 and earlier by @Kludex in #9526
- Record the tool requests of
ExaSearch,YouSearch,LocalStackand other harness capabilities under durable execution by @DouweM in #9705 - Bound retained realtime audio like retained images with
retain_audio_max_secondsby @DouweM in #9702 - Prefix
LogfireMCPtool names withlogfire_by @mpfaffenberger in #9868 - Serialize
FileUrl.media_typeasnullfor extensionless URLs instead of raising by @pydanty in #8405 - Take realtime history, usage, and reply waits from the session core on OpenAI, Azure OpenAI, and xAI by @DouweM in #9612
- Keep default
run_idandconversation_idstable when a durable run is re-executed by @DouweM in #9701 - Record each failed
FallbackModelattempt with its model, timing, error and usage, and count rejected responses inRunUsageby @DouweM in #8745
🚀 Features
- Add an
OpenAIDecisionsModelbackend with image input by @DenysMoskalenko in #9634 - Halve
import pydantic_aitime by loadingpydantic_ai.mcponly when anMCPcapability needs it by @DouweM in #10046 - Add
Conversationto carry and store a conversation between runs, accepted asconversation=by every entry point by @DouweM in #8337 - Record prompt cache diagnostics in
provider_details: on by default for OpenAI Responses,anthropic_cache_diagnosticsopt-in by @DouweM in #9406 - Report prompt-cache health per conversation via
pydantic_ai.cache.*span attributes and collapse events by @DouweM in #6534 - Default Azure AI Voice Live to API version
2026-07-15and applyparallel_tool_calls=Falsethere by @DouweM in #9409 - Raise Azure AI Voice Live
warningevents as Python warnings instead of dropping them by @DouweM in #9420 - Add
openai_cache_instructionsto place a prompt cache breakpoint on the instructions by @adtyavrdhn in #7559 - Add
hang_up()to end a WebRTC call from the server on gpt-realtime and GPT-Live by @DouweM in #9342 - Add
PostgresStepStoreandPostgresMediaStore, the hosted Postgres backends for messages and media by @rscholz98 in #9899 - Add unified
cachesetting andCachingcapability for cross-provider prompt caching by @adtyavrdhn in #7560 - Add support for
claude-haiku-5-5by @dsfaccini in #9998 - Seed
message_historyinto a WebRTC call throughanswer_webrtc_offerby @DouweM in #9338 - Configure Azure AI Voice Live semantic VAD, end-of-utterance detection, noise suppression and echo cancellation by @DouweM in #9414
- Cache the instructions and tool definitions too when
cache=Trueuses Anthropic's automatic caching by @DouweM in #10047 - Turn on prompt caching by default in harness
Coderby @DouweM in #10040 - Add
openai_live_idle_audioso text-only GPT-Live sessions work without a microphone by @DouweM in #9295
🐛 Bug Fixes
- Read Codex CLI
auth.jsonas UTF-8 regardless of platform locale by @pydanty in #8254 - Avoid false
WarnOnCacheBustswarnings after native tool calls by @aweis89 in #9746 - Run a capability
durable_operationinline when it is called from inside a durable unit by @DouweM in #9700 - End an xAI realtime session on a
max_durationerror instead of reconnecting into the same limit by @DouweM in #9415 - Attach a leading
CachePointto the previous message on OpenAI instead of raising by @DouweM in #9582 - Keep the run state recoverable when Ctrl-C interrupts
run_stream_sync()mid-output by @DouweM in #9709 - Raise
BedrockEmbeddingModelerrors asModelHTTPError/ModelAPIErrorinstead of anExceptionGroup, and map botocore transport errors by @DouweM in #9314 - Record GPT-Live's final usage on close by ending the session with
session.closeby @DouweM in #9403 - Stop a numeric request timeout from lengthening the connect and pool timeouts of Pydantic AI-created HTTP clients by @DouweM in #9599
- Don't record an OpenAI realtime
idle_timeout_msfollow-up as an empty user turn by @DouweM in #9396 - Close out interrupted tool calls when resuming a sub-agent with
delegate_task(resume=...)by @DouweM in #9783 - Publish
AgentRunResult's public shape in its validation JSON schema by @DouweM in #9791 - Record
Planning's plan store calls under durable execution, so recovery doesn't repeat its writes by @DouweM in #9789 - Report malformed GPT-Live events as recoverable errors instead of dropping them by @DouweM in #9393
- Record each
Memorytool call under durable execution, so a forked DBOS workflow doesn't repeat its writes by @DouweM in #9832 - Replay a spoken realtime turn without a transcript as a marker on reconnect instead of dropping it by @DouweM in #9782
- Don't let
openai_prompt_cache_retentionwiden the cache-health window on GPT-5.6+ by @DouweM in #9851 - Hide MCP Apps tools whose
_meta.ui.visibilityleaves out"model"from the model inMCPToolsetby @adtyavrdhn in #9862 - Drop the empty
moderationstub fromSnowflakeModelresponses by @specterbuilds in #9764 - Handle JSON-pointer
$refs and draft-07definitionsin MCP TypeScript SDK tool schemas by @dsfaccini in #9643 - Populate
span_treeand metrics forrun_on_errorsevaluators when a task raises by @Diwak4r in #7070 - Send a generic question when a
SystemOneModelquestion would have noinstructionsby @DouweM in #9848 - Move
MCPReadOnlyNoToolsWarningto_warnso harness submodule imports stay light by @pydanty in #9922 - Accrue
SpendLimitscontinuation chains per segment boundary so a durable retry charges newly billed segments by @mpfaffenberger in #10002 - Keep OpenAI's implicit prompt cache when
cache={'messages': False}can't place an instruction breakpoint by @DouweM in #10007 - Cancel the streamed model request when a parallel
InputGuardrailblocks by @mpfaffenberger in #10003 - Ignore media-type parameters in
is_text_like_media_typeby @pydanty in #9948 - Decode percent-encoded Base64 payloads in
BinaryContent.from_data_uriby @pydanty in #9950 - Add
gpt-6.1-solsupport on Bedrock Converse by @pydanty in #9817 - Fix history repair for a
tool_call_idrepeated within one response by @pydanty in #9819 - Preserve the current request when retained user turns exhaust the
SummarizingCompactiontail budget by @pydanty in #9715 - Resume deferred Code Mode calls in dispatch order in global sequential modes by @pydanty in #9972
- Fix
Researcherconstruction underTemporalDurabilityandPrefectDurabilityby @pydanty in #9564 - Fix
FallbackModelignoring handler functions infallback_ontuples by @pydanty in #10031 - Reject whitespace-only
SummarizingCompactionsummaries instead of replacing history with an empty summary by @aweis89 in #9748 - Catch overflowing
Retry-AfterHTTP dates by @ktz03 in #10023 - Keep OpenAI Responses tool-call arguments that arrive only in
function_call_arguments.doneoroutput_item.doneby @DouweM in #10039 - Return a 422 instead of a 500 from
Agent.to_web()for invalid chat requests by @DouweM in #10042 - Stop
EvaluationReport.render()andprint()from parsing case, task and evaluator text as Rich markup by @lets-order-some-fries in #9544 - Rewrite
$refs underpropertyNamesandpatternPropertieswhen merging schema defs by @jayzuccarelli in #8805 - Keep the current request when a
SlidingWindowCompactionreceipt would take the only kept slot by @aweis89 in #9719 - Reconnect xAI realtime sessions by replaying local history, since xAI conversation resumption drops typed user turns and tool calls by @DouweM in #9421
- Keep completed tool returns under
parallel_ordered_eventswhen a sibling raises or the run is cancelled by @DouweM in #10049 - Fix retried Anthropic requests sending empty base64 image and PDF bytes by @pydanty in #10014
- Keep the live workspace on a settled
StreamedRunResult.resultby @DouweM in #10048 - Fix ID-only tool call deltas emitting a
nullargs fragment in AG-UI and Vercel AI adapters by @pydanty in #9724 - Keep streamed OpenAI Responses tool-call arguments when the
donesnapshot disagrees, instead of starting the part again by @DouweM in #10044 - Stop dynamic system-prompt re-evaluation from mutating caller-owned message history by @pydanty in #9627
- Count a file in a tool return as its
See filereference, not itsrepr, in harness compaction by @pepijn-m in #9632 - List every module Monty can import in the
CodeModerun_codedescription by @mpfaffenberger in #9315 - Fix
merge_json_schema_defscorrupting branch schemas on transitive def-name collisions by @pydanty in #9735 - Charge
SpendLimitsbudgets for responses aFallbackModelrejected by @DouweM in #9792 - Fix
TextOutputJSON schema to infer the wrapped function's return annotation by @pydanty in #9652 - Charge
FallbackModelattempts once inSpendLimitswhen a failed continuation chain is retried by @DouweM in #10059 - Raise
ModelAPIErrorfor transport errors that escaped Google, Mistral, Cohere, Hugging Face, Anthropic, Groq and Bedrock by @DouweM in #9336 - Raise earlier Anthropic cache breakpoints to a later, longer TTL so requests aren't rejected by @DouweM in #10057
📦 Dependencies
- Raise the
openaifloor to 3.26.0 for Decisions support in theopenai,openrouter,bedrock-mantle,zai,snowflake,crusoeandcerebrasextras by @dsfaccini in #9968
New Contributors
- @atharva-mashalkar made their first contribution in #9567
- @pankajkoti made their first contribution in #8785
- @specterbuilds made their first contribution in #9764
- @ktz03 made their first contribution in #10023
- @lets-order-some-fries made their first contribution in #9544
- @jayzuccarelli made their first contribution in #8805
- @pepijn-m made their first contribution in #9632
Full Changelog: v2.54.0...v2.55.0