Highlights
- Fix Gemini thinking configuration passthrough, batch tool placement, and typed configuration serialization (#553).
- Fix Vertex AI batch jobs failing at import when
google_searchis configured without options (#561). - Surface blocked or failed Gemini batch items instead of returning empty output (#537).
- Fix legacy
format_typenormalization in extraction (#544). - Release completed prompts before building the next batch, reducing retained memory (#543).
- Clarify suppressed parsing errors (#521).
- Harden fork live tests and dispatch them from
main(CI only; #552).
Compatibility notes
- In batch mode with the default
ignore_item_errors=False, blocked items, items with a Vertex AI error status, and items that return no text with a finish reason other thanSTOP(for example,MAX_TOKENS) now raiseInferenceRuntimeErrorinstead of returning empty output, and results from that call are not cached. Setignore_item_errors=Trueto continue past them (#537). - A
thinking_configpassed to the Gemini provider is now sent to the API; v1.7.0 ignored it. Callers that already set it may see different output, latency, or cost, or an API error if the model rejects the setting (#553). - Batch cache keys now include tools. Cache entries written by v1.7.0 will not be reused, even for requests without tools, and repeated requests may incur new API charges (#553).
- In strict mode, empty cached batch results are rechecked, because the cache cannot distinguish a genuine empty completion from a previously ignored failure. This may repeat requests and incur API charges. Nonempty cache hits and explicit ignore mode are unchanged (#537).
- In Vertex AI batch mode, any tool other than
google_searchwith an empty configuration, such ascode_executionorurl_context, now fails before upload with anInferenceRuntimeErrorcaused byInferenceConfigError. Set an option on the tool if it has one, or disable batch mode (#561).
Full Changelog: v1.7.0...v1.7.1