pypi pydantic-ai 2.55.0
v2.55.0 (2026-10-09)

3 hours ago

What's Changed

⚠️ Compatibility Notes

  • Require Python 3.11 or newer in every package; Python 3.10 installs resolve to 2.54.0 and earlier by @Kludex in #9526
  • Record the tool requests of ExaSearch, YouSearch, LocalStack and other harness capabilities under durable execution by @DouweM in #9705
  • Bound retained realtime audio like retained images with retain_audio_max_seconds by @DouweM in #9702
  • Prefix LogfireMCP tool names with logfire_ by @mpfaffenberger in #9868
  • Serialize FileUrl.media_type as null for extensionless URLs instead of raising by @pydanty in #8405
  • Take realtime history, usage, and reply waits from the session core on OpenAI, Azure OpenAI, and xAI by @DouweM in #9612
  • Keep default run_id and conversation_id stable when a durable run is re-executed by @DouweM in #9701
  • Record each failed FallbackModel attempt with its model, timing, error and usage, and count rejected responses in RunUsage by @DouweM in #8745

🚀 Features

  • Add an OpenAIDecisionsModel backend with image input by @DenysMoskalenko in #9634
  • Halve import pydantic_ai time by loading pydantic_ai.mcp only when an MCP capability needs it by @DouweM in #10046
  • Add Conversation to carry and store a conversation between runs, accepted as conversation= by every entry point by @DouweM in #8337
  • Record prompt cache diagnostics in provider_details: on by default for OpenAI Responses, anthropic_cache_diagnostics opt-in by @DouweM in #9406
  • Report prompt-cache health per conversation via pydantic_ai.cache.* span attributes and collapse events by @DouweM in #6534
  • Default Azure AI Voice Live to API version 2026-07-15 and apply parallel_tool_calls=False there by @DouweM in #9409
  • Raise Azure AI Voice Live warning events as Python warnings instead of dropping them by @DouweM in #9420
  • Add openai_cache_instructions to place a prompt cache breakpoint on the instructions by @adtyavrdhn in #7559
  • Add hang_up() to end a WebRTC call from the server on gpt-realtime and GPT-Live by @DouweM in #9342
  • Add PostgresStepStore and PostgresMediaStore, the hosted Postgres backends for messages and media by @rscholz98 in #9899
  • Add unified cache setting and Caching capability for cross-provider prompt caching by @adtyavrdhn in #7560
  • Add support for claude-haiku-5-5 by @dsfaccini in #9998
  • Seed message_history into a WebRTC call through answer_webrtc_offer by @DouweM in #9338
  • Configure Azure AI Voice Live semantic VAD, end-of-utterance detection, noise suppression and echo cancellation by @DouweM in #9414
  • Cache the instructions and tool definitions too when cache=True uses Anthropic's automatic caching by @DouweM in #10047
  • Turn on prompt caching by default in harness Coder by @DouweM in #10040
  • Add openai_live_idle_audio so text-only GPT-Live sessions work without a microphone by @DouweM in #9295

🐛 Bug Fixes

  • Read Codex CLI auth.json as UTF-8 regardless of platform locale by @pydanty in #8254
  • Avoid false WarnOnCacheBusts warnings after native tool calls by @aweis89 in #9746
  • Run a capability durable_operation inline when it is called from inside a durable unit by @DouweM in #9700
  • End an xAI realtime session on a max_duration error instead of reconnecting into the same limit by @DouweM in #9415
  • Attach a leading CachePoint to the previous message on OpenAI instead of raising by @DouweM in #9582
  • Keep the run state recoverable when Ctrl-C interrupts run_stream_sync() mid-output by @DouweM in #9709
  • Raise BedrockEmbeddingModel errors as ModelHTTPError/ModelAPIError instead of an ExceptionGroup, and map botocore transport errors by @DouweM in #9314
  • Record GPT-Live's final usage on close by ending the session with session.close by @DouweM in #9403
  • Stop a numeric request timeout from lengthening the connect and pool timeouts of Pydantic AI-created HTTP clients by @DouweM in #9599
  • Don't record an OpenAI realtime idle_timeout_ms follow-up as an empty user turn by @DouweM in #9396
  • Close out interrupted tool calls when resuming a sub-agent with delegate_task(resume=...) by @DouweM in #9783
  • Publish AgentRunResult's public shape in its validation JSON schema by @DouweM in #9791
  • Record Planning's plan store calls under durable execution, so recovery doesn't repeat its writes by @DouweM in #9789
  • Report malformed GPT-Live events as recoverable errors instead of dropping them by @DouweM in #9393
  • Record each Memory tool call under durable execution, so a forked DBOS workflow doesn't repeat its writes by @DouweM in #9832
  • Replay a spoken realtime turn without a transcript as a marker on reconnect instead of dropping it by @DouweM in #9782
  • Don't let openai_prompt_cache_retention widen the cache-health window on GPT-5.6+ by @DouweM in #9851
  • Hide MCP Apps tools whose _meta.ui.visibility leaves out "model" from the model in MCPToolset by @adtyavrdhn in #9862
  • Drop the empty moderation stub from SnowflakeModel responses by @specterbuilds in #9764
  • Handle JSON-pointer $refs and draft-07 definitions in MCP TypeScript SDK tool schemas by @dsfaccini in #9643
  • Populate span_tree and metrics for run_on_errors evaluators when a task raises by @Diwak4r in #7070
  • Send a generic question when a SystemOneModel question would have no instructions by @DouweM in #9848
  • Move MCPReadOnlyNoToolsWarning to _warn so harness submodule imports stay light by @pydanty in #9922
  • Accrue SpendLimits continuation chains per segment boundary so a durable retry charges newly billed segments by @mpfaffenberger in #10002
  • Keep OpenAI's implicit prompt cache when cache={'messages': False} can't place an instruction breakpoint by @DouweM in #10007
  • Cancel the streamed model request when a parallel InputGuardrail blocks by @mpfaffenberger in #10003
  • Ignore media-type parameters in is_text_like_media_type by @pydanty in #9948
  • Decode percent-encoded Base64 payloads in BinaryContent.from_data_uri by @pydanty in #9950
  • Add gpt-6.1-sol support on Bedrock Converse by @pydanty in #9817
  • Fix history repair for a tool_call_id repeated within one response by @pydanty in #9819
  • Preserve the current request when retained user turns exhaust the SummarizingCompaction tail budget by @pydanty in #9715
  • Resume deferred Code Mode calls in dispatch order in global sequential modes by @pydanty in #9972
  • Fix Researcher construction under TemporalDurability and PrefectDurability by @pydanty in #9564
  • Fix FallbackModel ignoring handler functions in fallback_on tuples by @pydanty in #10031
  • Reject whitespace-only SummarizingCompaction summaries instead of replacing history with an empty summary by @aweis89 in #9748
  • Catch overflowing Retry-After HTTP dates by @ktz03 in #10023
  • Keep OpenAI Responses tool-call arguments that arrive only in function_call_arguments.done or output_item.done by @DouweM in #10039
  • Return a 422 instead of a 500 from Agent.to_web() for invalid chat requests by @DouweM in #10042
  • Stop EvaluationReport.render() and print() from parsing case, task and evaluator text as Rich markup by @lets-order-some-fries in #9544
  • Rewrite $refs under propertyNames and patternProperties when merging schema defs by @jayzuccarelli in #8805
  • Keep the current request when a SlidingWindowCompaction receipt would take the only kept slot by @aweis89 in #9719
  • Reconnect xAI realtime sessions by replaying local history, since xAI conversation resumption drops typed user turns and tool calls by @DouweM in #9421
  • Keep completed tool returns under parallel_ordered_events when a sibling raises or the run is cancelled by @DouweM in #10049
  • Fix retried Anthropic requests sending empty base64 image and PDF bytes by @pydanty in #10014
  • Keep the live workspace on a settled StreamedRunResult.result by @DouweM in #10048
  • Fix ID-only tool call deltas emitting a null args fragment in AG-UI and Vercel AI adapters by @pydanty in #9724
  • Keep streamed OpenAI Responses tool-call arguments when the done snapshot disagrees, instead of starting the part again by @DouweM in #10044
  • Stop dynamic system-prompt re-evaluation from mutating caller-owned message history by @pydanty in #9627
  • Count a file in a tool return as its See file reference, not its repr, in harness compaction by @pepijn-m in #9632
  • List every module Monty can import in the CodeMode run_code description by @mpfaffenberger in #9315
  • Fix merge_json_schema_defs corrupting branch schemas on transitive def-name collisions by @pydanty in #9735
  • Charge SpendLimits budgets for responses a FallbackModel rejected by @DouweM in #9792
  • Fix TextOutput JSON schema to infer the wrapped function's return annotation by @pydanty in #9652
  • Charge FallbackModel attempts once in SpendLimits when a failed continuation chain is retried by @DouweM in #10059
  • Raise ModelAPIError for transport errors that escaped Google, Mistral, Cohere, Hugging Face, Anthropic, Groq and Bedrock by @DouweM in #9336
  • Raise earlier Anthropic cache breakpoints to a later, longer TTL so requests aren't rejected by @DouweM in #10057

📦 Dependencies

  • Raise the openai floor to 3.26.0 for Decisions support in the openai, openrouter, bedrock-mantle, zai, snowflake, crusoe and cerebras extras by @dsfaccini in #9968

New Contributors

Full Changelog: v2.54.0...v2.55.0

Don't miss a new pydantic-ai release

NewReleases is sending notifications on new releases.