⭐️ Highlights
⬆️ Upgrade Notes
-
InMemoryDocumentStore.bm25_retrievalandInMemoryBM25RetrieverwithBM25L(the default) orBM25Plusnow return only documents that contain at least one query term, with or withoutscale_score. Previously, documents without any query term were returned with a positive score and filled uptop_k. As a result, retrieval can now return fewer thantop_kdocuments, or none at all.The
deltalower bound now applies only to query terms that occur in the document, as in the original BM25L/BM25+ definitions. So scores of matching documents are lower than before whenever the query has terms that a document does not contain. If you filter results with a fixed score threshold, re-check the threshold.BM25Okapiis not affected.
⚡️ Enhancement Notes
TextCleaner.runnow raises a clearTypeErrorwhentextsis not a list or any element is not astr, instead of failing later or producing unexpected results.SentenceWindowRetrievernow queries the Document Store once perrunorrun_asynccall instead of once per retrieved document. It combines the windows of all retrieved documents into a singleORfilter and assigns the returned documents to each window in memory. This reduces the load on Document Stores such as OpenSearch or Elasticsearch when many documents are retrieved. Thecontext_windowsandcontext_documentsoutputs are unchanged.
🔒 Security Notes
- Haystack now requires
anyio>=4.14.2.anyiois installed transitively throughhttpxandopenai; versions before 4.14.2 are affected by CVE-2026-63374 (GHSA-82r6-8w77-94w6).
🐛 Bug Fixes
-
Fixed BM25 tokenization in
InMemoryDocumentStorefor Chinese, Japanese and Korean text. The defaultbm25_tokenization_regexnow splits CJK characters into one token each, so bare-term queries can match words inside longer unspaced runs. Text is also NFC-normalized before tokenization, so composed and decomposed spellings produce the same tokens. -
Fix
Pipeline.connect()raisingTypeError: object of type 'ellipsis' has no len()when one of the sockets is annotated withCallable[..., T]. ACallable[..., T]now matches callables with any parameters, in both directions, as long as the return types are compatible. -
Fixed
ChatPromptBuilderdropping content parts when the template is a list ofChatMessageobjects. Only the firstTextContentpart was rendered, and every other part - additional texts, images, files -was removed from the rendered prompt without a warning. Template variables used in those dropped parts were not detected either, so they were not exposed as inputs. Now allTextContentparts are rendered and the remaining content parts are passed through unchanged. -
Fixed
CompactionHook.close()andclose_async()to release the token counter's resources as well as the compactor's. Previously, resources such asOpenAITokenCounter's HTTP client remained open after the hook or Agent was closed. -
ConfirmationHookno longer drops the messages that come after the last user or tool message when it rewrites the conversation, such as a system message added by another hook or an assistant answer followed by anon_exitreminder. Before, those messages were removed from the history the Agent sends to the model. -
DocumentToImageContentno longer raises aKeyErrorfor the whole batch when a document points to a PDF page that cannot be converted, because the page is out of range or the PDF cannot be read. That document now getsNoneinimage_contents, like other invalid documents, and the rest of the batch is still converted. The logged warnings now include the path of the PDF file. -
Fixed
DOCXToDocumentbreaking a Markdown table when a cell contains a pipe or spans several paragraphs. A pipe was emitted as a column separator, and a cell's line break ended the row in the middle of it. Pipes are now escaped and line breaks inside a cell are collapsed to a space. Thecsvtable format is unchanged: it already kept such cells intact by quoting them. -
Fixed
DOCXToDocumentdropping hyperlink addresses inside tables. Withlink_formatset tomarkdownorplain, links in table cells were written as their display text only, while links in body paragraphs kept their address. Links in table cells are now formatted the same way, in both themarkdownandcsvtable formats. The defaultlink_format="none"output is unchanged. -
Fixed
ConditionalRouter.from_dictmutating the caller'sroutesdata in place: serializedoutput_typestrings were deserialized directly inside the caller's dictionaries, so reusing the same serialized pipeline dict afterwards yielded already-deserialized type objects.from_dictnow works on a copy and the caller's data is left untouched.BranchJoiner.from_dictreceived the same treatment. -
LinkContentFetcherno longer adds itstimeoutandfollow_redirectsdefaults to theclient_kwargsdictionary passed by the caller. The dictionary is now copied before the defaults are applied, so reusing one HTTP client configuration across components no longer leaks these defaults. -
Fix
MetaFieldRankerto return no documents whenmissing_meta="drop"and all documents lack the ranking field or have aNonevalue. -
MetaFieldRankernow treats a Document whosemeta_fieldvalue isNonethe same as a Document that is missing the field, applying themissing_metasetting to it. Previously a singleNonevalue made sorting fail, so the ranker logged a warning and returned all Documents in their original order, ignoringmissing_meta="drop"as well. -
MockTextEmbedderandMockDocumentEmbeddernow accept non-positivedimensionvalues whenembeddingorembedding_fnis provided, matching the documented behavior. The default deterministic embedding mode still requires a positivedimension. -
ChatMessage.from_openai_dict_formatnow accepts tool-callargumentsthat are already a dictionary instead of raising aTypeError. Some OpenAI-compatible servers send a parsed object rather than a JSON string. -
MetaFieldGroupingRankernow treats agroup_byorsubgroup_byvalue ofNoneas missing, the same way it already treatssort_docs_by. Documents withNonego to the end with the other ungrouped documents, instead of forming a group named"None"that also absorbed documents whose value was the string"None". -
Fixed
JSONConverterfailing withValueErrorwhencontent_keycontains numeric or boolean scalar values. These values are now converted to strings before creating theDocument, whilenullvalues remain unchanged. -
Fixed
SentenceSplitterlosing the whitespace between two sentences when the first one ends with a closing quote, as inHe said "Hi." Bye.. Those characters were missing from the chunk text and shifted thesplit_idx_startoffset of every following chunk, so chunks could no longer be mapped back onto the original text. This affects all components that split on sentences, such asDocumentSplitter,RecursiveDocumentSplitter,MarkdownHeaderSplitterandEmbeddingBasedDocumentSplitter. Chunk boundaries change for text that contains quoted sentences, so re-indexing an existing corpus produces different chunks than before. -
Fixed
LLMMetadataExtractorincorrectly treating a chat generator output that carries its ownerrorfield as a failed LLM call, which sent every document tofailed_documentswithmetadata_extraction_errorset toNone. Such documents are now processed normally. -
Retrievers now handle
top_kconsistently. A negativetop_kpassed at runtime toInMemoryBM25Retriever,InMemoryEmbeddingRetrieverorMultiRetriever(top_kandtop_k_per_retriever) now raises aValueError. Previously it was applied as a negative slice, silently dropping the last documents. A runtimetop_k=0returns no documents.MultiRetrievernow also validatestop_kandtop_k_per_retrieverat initialization, raising aValueErrorwhen they are set and not greater than 0, matching the in-memory retrievers.Nonestill means no limit. -
Fixed
Pipeline.run(),Pipeline.run_async(), andPipeline.stream()treating a flat input with a dictionary value as a component name. For example,{"payload": {"x": 1}}now reaches components with apayloadinput without requiring the component name in the input data. -
Pipeline.run_asyncnow lets aBreakpointExceptionorPipelineRuntimeErrorraised by a nested component propagate unchanged, instead of wrapping it in anotherPipelineRuntimeError. This matches the synchronousPipeline.runand preserves the original error context (for example an agent snapshot) when a component internally runs a pipeline. -
XLSXToDocumentnow escapes pipe characters and replaces in-cell line breaks with spaces in the default Markdown pipe table output. This keeps cell content from being interpreted as additional table columns or rows.
💙 Big thank you to everyone who contributed to this release!
@alanhuangyoo, @anakin87, @bilgeyucel, @carey-bk, @Cha-Imaa, @chrikrah, @dakjdakd, @gauravch-code, @Harsh23Kashyap, @Jayanth-reflex, @JbravoI, @julian-risch, @L4XB, @Lesereingrape, @lets-order-some-fries, @mnm-matin, @MohammadHijjawi97, @nanhanq1, @pcbeingused333, @PeterSmith0127-lcm, @Rainmemery, @sclfcz, @serhiizghama, @shivsin25, @ShousenZHANG, @simpleqt, @winklemad