pypi transformers 5.19.0
Release v5.19.0

4 hours ago

Release v5.19.0

New Model additions

EmbeddingGemma2

image

EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture. It encodes text, images, audio, and video, individually or combined in one input, into a shared 768-dimensional vector space for cross-modal retrieval, semantic similarity, clustering, and classification. It uses Matryoshka Representation Learning, so embeddings can be truncated to 512, 256, or 128 dimensions. It also offers configurable visual and video token budgets, and unused vision or audio towers can be disabled at load time to save memory.

Links: Documentation

Breaking changes

All MoE models whose routers compute logits now return them when output_router_logits=True, following the Qwen3-MoE pattern (a router_logits recorder on the base model, MoeModelOutputWithPast from the backbone, and a MoE causal LM output from the head), so code that relied on the previous outputs or their absence should read the router logits from these output classes.

  • 🚨 Return router logits from every MoE model that computes them (#48920) by @qgallouedec

Owlv2ForObjectDetection.embed_image_query now selects the query box with the highest objectness score, as in the original OWLv2 notebook, instead of the OWL-ViT heuristic, so image-guided query embeddings and detections may differ from earlier releases.

The "paged|" prefix for SDPA and flash attention implementations is deprecated, so users should set the regular attention implementation (e.g. sdpa or flash_attention_2) for continuous batching instead of paged|sdpa or paged|flash_attention_2.

  • 🚨 Attention 🚨 Deprecate "paged|" prefix for SDPA and flash (#49112) by @remi-or

The regular flash and SDPA attention functions (flash_attention.py, sdpa_attention.py) now support continuous batching directly, and "paged|..." implementations for these are redirected to them, while eager still requires the "paged|eager" prefix.

  • 🚨 Attention 🚨 Make regular attention support CB (#49101) by @remi-or

In continuous batching, the cache update for the index-based and block-table paths is now fused into a single call, which slightly changes the cache update function's behavior and affects any custom code that calls the separate update paths.

  • 🚨 [CB] 🚨 Fuse update for index and block table path (#49088) by @remi-or

Continuous batching internals changed in preparation for removing "paged": `max

  • 🚨 [CB] 🚨 Little fixes before removing "paged" (#49069) by @remi-or

Parallelization

Expert parallelism gains a token-dispatch implementation, selected via the new ep_dispatch_experts plan rule and now the default for Qwen3 MoE and Mellum, which removes the requirement that EP size equal TP size. The Trainer was also adapted to work with expert parallelism, and the docs now note that PEFT adapters support tensor parallelism. A CI-related fix for pipeline-parallel chart2table inference was also included.

Cache

This release fixes quantized cache handling: generate no longer mutates the user's cache_config, and QuantizedLayer.reorder_cache is repaired. It also adds per-layer cache configuration, so DynamicCache and StaticCache initialize each layer from its own config (sliding window, attention chunk size, conv states, and attention head counts) to better support heterogeneous models. Separately, the deprecation cycle on mask and cache

Bugfixes and improvements

Significant community contributions

The following contributors have made significant changes to the library over the last release:

  • @vasqu
  • @remi-or
    • [Fix] Remove old test file with two ancient tests (#49280)
    • [CB] Make CB more device agnostic and add support for XPU (#49156)
    • 🚨 Attention 🚨 Deprecate "paged|" prefix for SDPA and flash (#49112)
    • 🚨 Attention 🚨 Make regular attention support CB (#49101)
    • 🚨 [CB] 🚨 Fuse update for index and block table path (#49088)
    • [Refactor] Make some flash-attention utils more readable (#49071)
    • 🚨 [CB] 🚨 Little fixes before removing "paged" (#49069)

Don't miss a new transformers release

NewReleases is sending notifications on new releases.