github deepset-ai/haystack v3.3.0-rc1

pre-release2 hours ago

⭐️ Highlights

⬆️ Upgrade Notes

  • InMemoryDocumentStore.bm25_retrieval and InMemoryBM25Retriever with BM25L (the default) or BM25Plus now return only documents that contain at least one query term, with or without scale_score. Previously, documents without any query term were returned with a positive score and filled up top_k. As a result, retrieval can now return fewer than top_k documents, or none at all.

    The delta lower bound now applies only to query terms that occur in the document, as in the original BM25L/BM25+ definitions. So scores of matching documents are lower than before whenever the query has terms that a document does not contain. If you filter results with a fixed score threshold, re-check the threshold. BM25Okapi is not affected.

⚡️ Enhancement Notes

  • TextCleaner.run now raises a clear TypeError when texts is not a list or any element is not a str, instead of failing later or producing unexpected results.
  • SentenceWindowRetriever now queries the Document Store once per run or run_async call instead of once per retrieved document. It combines the windows of all retrieved documents into a single OR filter and assigns the returned documents to each window in memory. This reduces the load on Document Stores such as OpenSearch or Elasticsearch when many documents are retrieved. The context_windows and context_documents outputs are unchanged.

🔒 Security Notes

  • Haystack now requires anyio>=4.14.2. anyio is installed transitively through httpx and openai; versions before 4.14.2 are affected by CVE-2026-63374 (GHSA-82r6-8w77-94w6).

🐛 Bug Fixes

  • Fixed BM25 tokenization in InMemoryDocumentStore for Chinese, Japanese and Korean text. The default bm25_tokenization_regex now splits CJK characters into one token each, so bare-term queries can match words inside longer unspaced runs. Text is also NFC-normalized before tokenization, so composed and decomposed spellings produce the same tokens.

  • Fix Pipeline.connect() raising TypeError: object of type 'ellipsis' has no len() when one of the sockets is annotated with Callable[..., T]. A Callable[..., T] now matches callables with any parameters, in both directions, as long as the return types are compatible.

  • Fixed ChatPromptBuilder dropping content parts when the template is a list of ChatMessage objects. Only the first TextContent part was rendered, and every other part - additional texts, images, files -was removed from the rendered prompt without a warning. Template variables used in those dropped parts were not detected either, so they were not exposed as inputs. Now all TextContent parts are rendered and the remaining content parts are passed through unchanged.

  • Fixed CompactionHook.close() and close_async() to release the token counter's resources as well as the compactor's. Previously, resources such as OpenAITokenCounter's HTTP client remained open after the hook or Agent was closed.

  • ConfirmationHook no longer drops the messages that come after the last user or tool message when it rewrites the conversation, such as a system message added by another hook or an assistant answer followed by an on_exit reminder. Before, those messages were removed from the history the Agent sends to the model.

  • DocumentToImageContent no longer raises a KeyError for the whole batch when a document points to a PDF page that cannot be converted, because the page is out of range or the PDF cannot be read. That document now gets None in image_contents, like other invalid documents, and the rest of the batch is still converted. The logged warnings now include the path of the PDF file.

  • Fixed DOCXToDocument breaking a Markdown table when a cell contains a pipe or spans several paragraphs. A pipe was emitted as a column separator, and a cell's line break ended the row in the middle of it. Pipes are now escaped and line breaks inside a cell are collapsed to a space. The csv table format is unchanged: it already kept such cells intact by quoting them.

  • Fixed DOCXToDocument dropping hyperlink addresses inside tables. With link_format set to markdown or plain, links in table cells were written as their display text only, while links in body paragraphs kept their address. Links in table cells are now formatted the same way, in both the markdown and csv table formats. The default link_format="none" output is unchanged.

  • Fixed ConditionalRouter.from_dict mutating the caller's routes data in place: serialized output_type strings were deserialized directly inside the caller's dictionaries, so reusing the same serialized pipeline dict afterwards yielded already-deserialized type objects. from_dict now works on a copy and the caller's data is left untouched. BranchJoiner.from_dict received the same treatment.

  • LinkContentFetcher no longer adds its timeout and follow_redirects defaults to the client_kwargs dictionary passed by the caller. The dictionary is now copied before the defaults are applied, so reusing one HTTP client configuration across components no longer leaks these defaults.

  • Fix MetaFieldRanker to return no documents when missing_meta="drop" and all documents lack the ranking field or have a None value.

  • MetaFieldRanker now treats a Document whose meta_field value is None the same as a Document that is missing the field, applying the missing_meta setting to it. Previously a single None value made sorting fail, so the ranker logged a warning and returned all Documents in their original order, ignoring missing_meta="drop" as well.

  • MockTextEmbedder and MockDocumentEmbedder now accept non-positive dimension values when embedding or embedding_fn is provided, matching the documented behavior. The default deterministic embedding mode still requires a positive dimension.

  • ChatMessage.from_openai_dict_format now accepts tool-call arguments that are already a dictionary instead of raising a TypeError. Some OpenAI-compatible servers send a parsed object rather than a JSON string.

  • MetaFieldGroupingRanker now treats a group_by or subgroup_by value of None as missing, the same way it already treats sort_docs_by. Documents with None go to the end with the other ungrouped documents, instead of forming a group named "None" that also absorbed documents whose value was the string "None".

  • Fixed JSONConverter failing with ValueError when content_key contains numeric or boolean scalar values. These values are now converted to strings before creating the Document, while null values remain unchanged.

  • Fixed SentenceSplitter losing the whitespace between two sentences when the first one ends with a closing quote, as in He said "Hi." Bye.. Those characters were missing from the chunk text and shifted the split_idx_start offset of every following chunk, so chunks could no longer be mapped back onto the original text. This affects all components that split on sentences, such as DocumentSplitter, RecursiveDocumentSplitter, MarkdownHeaderSplitter and EmbeddingBasedDocumentSplitter. Chunk boundaries change for text that contains quoted sentences, so re-indexing an existing corpus produces different chunks than before.

  • Fixed LLMMetadataExtractor incorrectly treating a chat generator output that carries its own error field as a failed LLM call, which sent every document to failed_documents with metadata_extraction_error set to None. Such documents are now processed normally.

  • Retrievers now handle top_k consistently. A negative top_k passed at runtime to InMemoryBM25Retriever, InMemoryEmbeddingRetriever or MultiRetriever (top_k and top_k_per_retriever) now raises a ValueError. Previously it was applied as a negative slice, silently dropping the last documents. A runtime top_k=0 returns no documents.

    MultiRetriever now also validates top_k and top_k_per_retriever at initialization, raising a ValueError when they are set and not greater than 0, matching the in-memory retrievers. None still means no limit.

  • Fixed Pipeline.run(), Pipeline.run_async(), and Pipeline.stream() treating a flat input with a dictionary value as a component name. For example, {"payload": {"x": 1}} now reaches components with a payload input without requiring the component name in the input data.

  • Pipeline.run_async now lets a BreakpointException or PipelineRuntimeError raised by a nested component propagate unchanged, instead of wrapping it in another PipelineRuntimeError. This matches the synchronous Pipeline.run and preserves the original error context (for example an agent snapshot) when a component internally runs a pipeline.

  • XLSXToDocument now escapes pipe characters and replaces in-cell line breaks with spaces in the default Markdown pipe table output. This keeps cell content from being interpreted as additional table columns or rows.

💙 Big thank you to everyone who contributed to this release!

@alanhuangyoo, @anakin87, @bilgeyucel, @carey-bk, @Cha-Imaa, @chrikrah, @dakjdakd, @gauravch-code, @Harsh23Kashyap, @Jayanth-reflex, @JbravoI, @julian-risch, @L4XB, @Lesereingrape, @lets-order-some-fries, @mnm-matin, @MohammadHijjawi97, @nanhanq1, @pcbeingused333, @PeterSmith0127-lcm, @Rainmemery, @sclfcz, @serhiizghama, @shivsin25, @ShousenZHANG, @simpleqt, @winklemad

Don't miss a new haystack release

NewReleases is sending notifications on new releases.