Added
- (pdf): pre-extracted spans can use canonical page reading order.
order_page_spansapplies tagged structure,
article-thread, and geometric ordering without parsing page content a second time, including for layer-filtered text. - (ocr): failed PDF OCR pages are reported as a typed list in every binding.
ExtractedDocumenthas a new
ocr_page_failuresfield. It is absent (None) when no page failed. EachOcrPageFailuregives the 1-basedpage,
the backenderrortext, andrecovered, which istruewhen the page still has content from native text or from
its embedded images. When an automatic OCR run of selected pages fails as a whole, each selected page has a record.
The records were a JSON array undermetadata.additional["ocr_page_failures"]; that key is no longer set. The
warnings do not change. (GH#2064)
Changed
- (benchmarks): cold-process and warm in-process latency are reported as separate regimes. Opt-in steady-state
Xberg runs retain model sessions across iterations, publish a separate cold-start probe, and remain excluded from
cross-framework rankings until competitors expose equivalent persistent-process adapters. - (batch): default in-memory extraction avoids no-op cache work. Engines without an injected cache backend no
longer hash every byte input or serialize completed results for a cache that discards them, reducing CPU time and
transient allocations without changing extractor or OCR cache behavior.
Fixed
-
(node): extraction works with
AsyncLocalStorageand async hooks enabled. Node bindings heap-pin extraction
futures before creating their JavaScript promises, avoiding a synchronous stack overflow on Node 22 and later. -
(pdf, ocr): Arabic and Hebrew extraction preserve logical reading order. Tagged native PDFs use trustworthy
structure order without reversing table cells, and Tesseract output orders mixed-direction lines and tables from
their local text direction and geometry. Sparse whole-image RTL OCR also retries a single-block segmentation mode
when it recovers more strong-script tokens without materially reducing confidence, and mixed Arabic/Latin scans
recover identifiers only when a bounded English crop provides compatible evidence. -
(benchmarks): benchmark provenance is bound to the executable that actually ran. Clean-checkout runs now reject
stale Xberg binaries whose embedded build identifier does not match the repository commit, and machine-readable
xberg versionoutput includes that identifier for auditability. -
(doc): legacy Word documents preserve headers, footers, footnotes, comments, and text boxes. Non-body stories
are now emitted with their document layers and respect the existing header, footer, and footnote filters instead of
being dropped whenever the document also contains body text. (GH#2054) -
(php): Linux packages run on Debian 12 and other glibc 2.36 systems. Release artifacts are built against the
declared ABI floor, reject newer GLIBC, GLIBCXX, or CXXABI requirements, and are load-tested in Debian 12 before
publication. (GH#2060) -
(ocr): columned prose and punctuation fragments are no longer fabricated as tables. Long, sparse newsletter
columns remain ordinary text in both layout and non-layout extraction, while punctuation-only scan regions fall
back to paragraph output without losing recognized content. -
(ocr): layout extraction preserves scanned-PDF recognition settings. Scanned PDFs no longer force Tesseract
into whole-image single-block segmentation when layout assembly has detected regions, while known scans keep their
sparse-text segmentation and source-resolution preprocessing hints. -
(typst): tables wrapped in figures are preserved. A
#table(...)nested inside a#figure(...), including
through analign(...)wrapper, now reaches structured and rendered output instead of being discarded with the
figure body. -
(latex): multiline section titles remain headings. A section command whose braced title spans physical lines
now produces one heading with its label anchor instead of flattened paragraph text and command fragments. -
(python, pdf): Python post-processors preserve formatted output and hide internal PDF heading annotations.
Returning anExtractedDocumentfrom a Python post-processor no longer replaces Rust-only document state, so a
no-op processor keeps rendered headings and chunks. Font-size annotations used for native PDF heading detection are
also excluded from plain text and keyword extraction. (GH#2050) -
(pdf): conflicting JPEG 2000 palette color spaces fail consistently across architectures. Native PDF extraction
now rejects a codestream palette whose component count conflicts with an explicit/ColorSpace, instead of relying
on architecture-dependent decoder rounding that could accept the same malformed image on ARM and reject it on x86. -
(zig): generated bindings compile with Zig 0.17. String duplication now uses the Zig 0.17 sentinel API, and
plugin vtables cross the C ABI through layout-compatible pointers instead of rejected value bitcasts.