Added
- (redaction): findings from an external content-inspection engine can drive redaction.
RedactionConfig.findingsandfindings_pathaccept inline JSON, JSON Lines, or a file, while the additiveextract_with_external_redactionandredact_externaloperations accept raw JSON or JSON Lines for request-scoped findings. Each value is redacted at every occurrence in every textual field and reported asPiiCategory::Custom(label);redaction.min_scorefilters scored findings below a validated[0.0, 1.0]threshold while retaining unscored findings. Unknown fields are ignored, Presidio, AWS Comprehend, Azure Language, and GCP DLP field names are accepted as aliases, and offsets support UTF-8 bytes, Unicode code points (the default), and UTF-16 code units. An unresolvable finding fails instead of being skipped, configured findings are capped at the lower ofSecurityLimits.max_iterationsand 10,000, andparse_external_findings_boundedenforces caller-supplied limits while parsing. (GH#1941) - (cli):
--redaction-findings <PATH|->passes an external engine's findings toextractandbatch. It takes a JSON array or JSON Lines file, or-to read them from stdin, and turns redaction on if the config has not.-cannot be combined withextract --stdin. (GH#1941) - (captioning): generated image captions can preserve, combine with, or replace existing alt text. Set
captioning.alt_textto"preserve"(the backward-compatible default),"combine", or"replace";ExtractedImage.captioncontinues to hold the generated caption whiledescriptionand rendered DOCX/PPTX image placeholders follow the selected precedence. (GH#1998) - (rendering):
output_format = "docx"returns the document as a Word file.contentholds the.docxpackage base64-encoded andmetadata.output_formatreads"docx"; the CLI's--content-format docxwrites the file itself to stdout. The package is built from the Markdown rendering after every post-processor has run, so redaction removes a term from the Word file exactly as it does from Markdown output. Headings, paragraphs, emphasis, links, nested lists, tables and code blocks carry over; images are not embedded, and chunks describe the Markdown the file was built from. Requires theofficefeature. (GH#1942) - (rendering):
output_format = "pdf"returns the document as a PDF file.contentholds the PDF base64-encoded andmetadata.output_formatreads"pdf"; the CLI's--content-format pdfwrites the file itself to stdout. The PDF is laid out from the Markdown rendering after every post-processor has run, so redaction removes a term from it exactly as it does from Markdown output. Text stays selectable and extractable: it is set in an embedded subset of DejaVu Sans with aToUnicodemap, and code in the standard Courier faces. Headings (also listed in the PDF outline), emphasis, links, nested lists, tables and code blocks carry over; images are not embedded, and characters DejaVu Sans has no glyph for, such as CJK text, draw as boxes but still extract. Requires thepdffeature. (GH#1943) - (tesseract):
ResultIterator::extract_all_words_with_line_startsreports which extracted words begin Tesseract text lines. It returns the existing word extraction outcome plus an aligned vector of flags, moving a failed first word's flag to the next successfully extracted word without changingWordDataconstruction. (GH#1978) - (ocr): typed PaddleOCR settings are available without breaking the legacy config field.
OcrConfig.paddle_ocr_settingsandOcrPipelineStage.paddle_ocr_settingsexposePaddleOcrConfigin every binding, whilepaddle_ocr_configremains raw JSON with its existing Rust, C and binding APIs. Typed settings take precedence when both fields are set. Config files, JSON configs, and REST or MCP request bodies use snake_case keys underpaddle_ocr_settingsand reject unknown keys; binding objects use their language's usual field names. Thepaddle-ocr-typesCargo feature remains accepted as a compatibility alias. (GH#2010)
Changed
- Breaking (Rust and generated bindings): configuration constructors change with the new fields and corrected defaults. Exhaustive Rust struct literals must include
CaptioningConfig.alt_text, typedpaddle_ocr_settingsonOcrConfigandOcrPipelineStage, and the external-finding fields onRedactionConfig, or use..Default::default(). Java record canonical constructors gain the corresponding components. PHP constructors move fields with generated defaults after required parameters, andRedactionConfiggains a requiredfindingsarray; positional PHP calls must be updated or replaced with named arguments. Builder, default-constructor, and JSON surfaces continue to apply defaults where their binding supports omission. These compatibility breaks are intentionally shipped in the1.3.2patch release so the completed fixes remain available together. (GH#1941, GH#1998, GH#2010) - (ocr): Tesseract returns line elements next to its word elements. Each line holds its words' text in reading order, the union of their boxes and their word-count-weighted mean confidence. A caller who requests
min_level: "word"now gets the lines first, then the words, so element positions and the ids frombuild_hierarchyshift. Withlayoutconfigured, image OCR text is now assembled from these lines, so it reads in the source's word order instead of a per-word position sort. (GH#1978)
Deprecated
- (ocr): raw
paddle_ocr_configfields and thepaddle-ocr-typesCargo feature are deprecated. Use typedpaddle_ocr_settingsonOcrConfigandOcrPipelineStage. The raw JSON fields remain lenient toward unknown extension keys for compatibility, while invalid values for known settings still fail validation. The Cargo feature remains accepted as a no-op alias. Both are planned for removal in 2.0. (GH#2010)
Fixed
- (python): nested configuration fields accept the public config classes. Generated Python config objects such as
OcrPipelineStagenow convert publicxberg.*Configdataclasses, dictionaries, and JSON strings instead of requiring callers to construct privatexberg._xbergclasses. (GH#2015) - (redaction): structured table column labels no longer retain PII after redaction. Redaction now rewrites
tables[].columnsandpages[].tables[].columnsalongside table cells and Markdown, so a header copied into the structured column metadata cannot expose a value removed from the document text. (GH#1991) - (cli): a
redactionblock in a CLI config is applied. No CLI build compiled the redaction engine in, so the block parsed, did nothing, and the output came back unredacted. Theredactionfeature is now part of the default andbinstallbuilds, andallthrough the default. (GH#1941) - (redaction): a validation error raised while redacting now fails the extraction instead of returning the document unredacted. The pipeline kept a post-processor's validation error as a processing warning and returned the document as extracted, so a redaction run that failed, for example an
LlmNER backend with noNerConfig.llm, handed back the very text it was asked to redact. The redaction processor now reports these as a plugin error, which fails the extraction. (GH#1941) - (ocr): PaddleOCR's selected-page PDF route keeps embedded-image prose beside recovered tables in Markdown. When a page render was blank, the embedded-image retry recovered both prose and a table, but document restructuring rebuilt the page from an empty paragraph list and retained only the table. The retry's non-table text now feeds that restructuring pass, matching plain output without duplicating table text. (GH#2021)
- (plugins): cached file extractions now reflect the currently registered post-processors and validators. Registering, removing, or replacing a lifecycle hook invalidates entries produced by the older registry state, so the next extraction runs every current hook exactly once and refreshes the cache. (GH#2032)
- (ocr): macOS Apple Silicon release binaries now accelerate Candle OCR with Metal. The downloadable
aarch64-apple-darwinCLI enablescandle-metal, so the bundled TrOCR, PaddleOCR-VL, GLM-OCR, and DeepSeek-OCR backends no longer fall back to impractically slow CPU inference on supported Macs. Other release targets keep their existing feature sets. (GH#2033) - (config): project-local YAML and JSON configuration files are now auto-discovered.
ExtractionConfig::discover()probesxberg.toml,xberg.yaml,xberg.yml, andxberg.jsonin that order in the current directory and each parent before falling back to the user config directory. (GH#2017) - (excel): spreadsheets with valid producer quirks no longer lose content or fail parsing, and malformed XLS files cannot abort on a forged stream length. XLSX worksheets whose dimension names only rows now extract their shared-string cells, legal indentation whitespace between ODS cells is ignored, and legacy XLS/XLA compound streams are checked against the container size and
security_limits.max_archive_sizebefore parsing. (GH#2002, GH#2003, GH#2005) - (odt): text boxes in page-anchored frames are extracted, and encrypted documents report encryption explicitly. Paragraphs under
draw:frame > draw:text-boxno longer produce an empty success when the frame sits outside normal paragraph flow, and recursive frames, lists, and inline wrappers respect the configured nesting limit. An ODT whose manifest markscontent.xmlas encrypted now returns an encryption error instead of reporting its ciphertext as invalid UTF-8, without assuming whether password- or public-key encryption was used. (GH#2006, GH#2004) - (ocr): automatic PaddleOCR fallback no longer enables table detection unless the caller asks for it. The synthesized Tesseract-to-Paddle pipeline keeps PaddleOCR's safe default instead of inferring Paddle table intent from Tesseract. Explicit legacy or typed PaddleOCR settings remain authoritative, and selecting a Paddle model version or tier does not change table-detection intent. (GH#2018, GH#2030, GH#2031)
- (ocr): scanned tables keep sparse amount columns split by digit-width drift. Table post-processing no longer folds an unnamed value track into a label when the matching numeric track is on its right. It now folds row-disjoint values into the numeric neighbour on either side, recognizes multi-row period labels and consecutive plain years as headers, and keeps a title fragment with the title while moving values under the correct year. (GH#1832, GH#2028, GH#2029)
- (ocr): a scanned table is no longer discarded over one unnamed, half-empty amount column. Table post-processing rejected the whole table when a column with no header was empty in more than half of its rows and held few characters, but folded the same column into its neighbour when it was nearly empty. A right-aligned amount column that OCR splits by digit width sits between those two bars, and a correct row merge elsewhere in the grid could move it from "folded" to "rejected", so the page fell back to plain text. An unnamed value column beside a value column is now folded whenever that gate would reject the table on it, and the short amounts return to the column they were split from. (GH#2019)
- (ocr): table post-processing no longer breaks when it folds the first column. The fold looked for a title column left of column 0. Debug builds panicked on that lookup, and release builds dropped the header of a title and value track that an earlier fold had moved to column 0. (GH#2026)
- (ocr): a PDF page that renders blank keeps the text OCR recovers from its embedded images when those images also hold a table. With
force_ocrand table detection on, such a page is read again from its embedded images. The result kept the recovered table but dropped the recovered text, with Tesseract and with PaddleOCR. The page text is now left out only when the OCR backend reports that the detected tables hold all of it, which also keeps a full-page table from printing twice. The recovered text also no longer repeats the recovered table: for an image with tables, the retry keeps only the lines those tables do not hold, so each table appears once. (GH#2008, GH#2014) - (cache): structured extraction results now produce cache hits. Extraction cache entries use named MessagePack fields so tables and document nodes round-trip correctly. Unreadable legacy entries are reported and safely replaced after re-extraction. (GH#1990)
- (mime): large CSS and TOML files with array-like prefixes are no longer mistaken for JSON. Valid JSON arrays and objects larger than the 4 KiB sniffing window remain detected as JSON. (GH#1982)
- (ocr): label-and-value tables no longer disappear when labels contain most of the text. A two-column grid whose first column contains recurring labels and whose second column is mostly numeric values now bypasses the prose-oriented dominant-column rejection. PaddleOCR also reports structurally rejected table candidates in
processing_warningsinstead of leaving only a debug trace. (GH#1970) - (wasm): Tesseract OCR returns word elements when
element_config.include_elementsis enabled. The WebAssembly backend now includes each word's bounding box and recognition confidence, applies the configured level and confidence filters, and leaves element extraction disabled unless requested. (GH#1974) - (ocr): a tall word box no longer splits one scanned-table row into two. Row detection now anchors on representative-height words, lets a tall word seed a row only when it overlaps no existing row band, then assigns every word by overlap with those bands. A word whose OCR box grows above or below its neighbours therefore remains in their row regardless of input order, while a standalone tall heading remains its own row. (GH#1977)
- (transcription): a 30-second window the Whisper decoder cannot process no longer aborts the whole recording. The greedy loop treated the decoder's 448-token context as a budget for generated tokens alone, so once a repetition loop pushed the prompt plus generated tokens past the position-embedding table, the decoder sliced an empty position range and crashed with an ONNX
Reshapeerror (Input shape:{8,0,64}). The prompt length is now reserved from the context budget, and a window that still fails is skipped, reported as aprocessing_warningsentry, and the remaining windows are transcribed instead of the file failing. (GH#1944) - (pdf): numeric footnotes stay separate in structured paragraphs as well as page text. A reference such as
comma 3followed by note5no longer becomescomma 35indocument.nodesor Markdown. Inline-script rejoining now uses the same matching-note evidence as plain-text extraction; scripts without that evidence and adjacent mathematical expressions retain their existing joins. - (ocr): numbered list items detected in standalone images keep their printed numbers. Image layout extraction now separates a leading numeric marker from the item text and records it as the source label, so Markdown rendering no longer strips the only copy of the number. (GH#1979)
- (ocr): a table header label wider than the amounts under it no longer splits its column. On a scan, a right-aligned label that reads as text, such as a period written as letters and digits, starts further left than every amount and formed a header-only column of its own. The column's long amounts then followed the label's left edge, the short amounts stayed in a headerless column, and the sparse-column check dropped the whole table. A header-only track now folds into the value column whose right edge it shares, whatever its label reads, unless that column has a label of its own in the same header row: two labels side by side stay two columns. With
normalize_shaded_rowson, which stays opt-in, a scanned table whose recovered header band carries such period labels keeps its table. (GH#2001) - (ocr): a scanned table is no longer dropped when OCR reads one junk cell in a split amount column, or when its header band spans several rows. On a scan, one right-aligned amount column splits into two adjacent columns by digit width. The fold that rejoins them ran only when every token below the first row read as a value, so one misread letter, one rule mark, or the labels of a second header row kept the column split. The leftover column then failed the sparse-column check and the whole table was lost. The fold now reads the header as the band of rows above the first row with a number, and folds a column whose values outnumber its other tokens. A text column never folds, whatever numbers it holds, and two columns that each hold a number in the same row never fold. (GH#1952)
- (pdf): retain alignment instructions beside split description cells. Cell detection now searches for a closing boundary on both sides instead of stopping at a neighbouring row separator. (GH#1958)
- (pdf): table-cell text keeps the same word spacing as the page text. On a tight face the glyphs on either side of a source space can abut or overlap, so the word merger read the zero-or-negative gap as "one word" and fused the two — turning
corrispettivi superiori al costointocorrispettivisuperiori alcostoandd’impresaintod’im presainside table cells while the page text was correct. The word merger now honours the source whitespace as a word boundary, table words carry that source space even when their boxes touch, and the fragment re-glue band was widened so a font that kerns one word apart by more than a normal space still reassembles. (GH#1948) - (dbf): every row of a dBASE table now keeps its values under the headers they belong to. Each record was read into a name-keyed
HashMapand walked in hash order, which differs from record to record and from run to run, so a table with more than one field came back with every row's cells permuted andmetadata.format.fieldscould pair a field name with another field's type. Records are now read in field-declaration order, which also keeps both values when two fields share a name instead of dropping one. (GH#1968) - (pdf): preserve grouped table header spans in structured output. Native span geometry now reaches document nodes, while dense cells and Markdown retain consistent columns. Mixed row/column spans use the shared placement rules, and redaction covers retained native grid text. (GH#1959)
- (ocr): Tesseract returns OCR elements when
include_elementsis set without amin_level. The default level isline, and Tesseract produced only words, so images and PDFs returned no elements at all. (GH#1978) - (heuristics): the OCR confidence aggregate and the quality-score evidence floor count each recognized word once. When a result carries words and the lines that hold them, as PaddleOCR and Tesseract results do at
min_level: "word", only the finest level is folded. Previously every word counted twice, so ten words reached the 20-word floor of the quality-score cap. (GH#1978) - (ocr): a table that PaddleOCR detects in an image no longer repeats its text in the content. With
enable_table_detectionon, the content listed each recognised line of the table as its own paragraph and then the same text again as the table rows. PaddleOCR now removes the lines that a detected table carries, with the same rule the Tesseract backend uses. A table that lost words from its region keeps those lines, so no text is lost. A scanned PDF page that is only a table no longer shows the table text a second time on theforce_ocrroute. A prose image that table detection turns into a table now shows its text once, as that table. (GH#1973) - (ocr): a table whose rows are spaced widely no longer loses its trailing line items. Rows spaced beyond the vertical region-gap threshold became one-row regions; with fewer than the six-word table minimum they were dropped, so an invoice kept only its header and first item. A region below the minimum now attaches to a column-aligned neighbour instead of being discarded, while genuinely separate tables — each at least the minimum size — still stay apart. (GH#1957)
- (ocr): a scanned table whose amount column splits into two tracks is no longer dropped when OCR reads the page's rules as cells. On a scan, one right-aligned amount column can split into two adjacent columns by digit width. The merge that rejoins them refused whenever a track held a mark read from a rule or a shaded band (
-, a dash run,:,~), so the leftover track failed the sparse-column check and the whole table was lost. A cell with no letter or digit now counts as empty when the merge compares the two tracks, such a mark in the header band is no longer a column label, and the merged cell keeps the real amount. A sign, currency sign or bracket in front of an amount still keeps the two tracks apart, so it is never dropped. (GH#1949) - (ocr): a scanned page whose text layer has no usable character map keeps the scan's segmentation mode. Automatic OCR routing gave Tesseract block mode (PSM 6) to every page whose text layer has no usable character map, including a scan that carries such a layer over its image. On a scanned table, block mode loses the table reconstruction, so the values leave their rows. Block mode now applies only to a page without a scan raster: a page whose images cover a quarter of it or more keeps the mode a scan gets. A page of unmapped vector text still gets block mode, and an explicit caller setting still wins. (GH#1946)