github xberg-io/xberg v1.3.7

4 hours ago

Added

  • (pdf): pre-extracted spans can use canonical page reading order. order_page_spans applies tagged structure,
    article-thread, and geometric ordering without parsing page content a second time, including for layer-filtered text.
  • (ocr): failed PDF OCR pages are reported as a typed list in every binding. ExtractedDocument has a new
    ocr_page_failures field. It is absent (None) when no page failed. Each OcrPageFailure gives the 1-based page,
    the backend error text, and recovered, which is true when the page still has content from native text or from
    its embedded images. When an automatic OCR run of selected pages fails as a whole, each selected page has a record.
    The records were a JSON array under metadata.additional["ocr_page_failures"]; that key is no longer set. The
    warnings do not change. (GH#2064)

Changed

  • (benchmarks): cold-process and warm in-process latency are reported as separate regimes. Opt-in steady-state
    Xberg runs retain model sessions across iterations, publish a separate cold-start probe, and remain excluded from
    cross-framework rankings until competitors expose equivalent persistent-process adapters.
  • (batch): default in-memory extraction avoids no-op cache work. Engines without an injected cache backend no
    longer hash every byte input or serialize completed results for a cache that discards them, reducing CPU time and
    transient allocations without changing extractor or OCR cache behavior.

Fixed

  • (node): extraction works with AsyncLocalStorage and async hooks enabled. Node bindings heap-pin extraction
    futures before creating their JavaScript promises, avoiding a synchronous stack overflow on Node 22 and later.

  • (pdf, ocr): Arabic and Hebrew extraction preserve logical reading order. Tagged native PDFs use trustworthy
    structure order without reversing table cells, and Tesseract output orders mixed-direction lines and tables from
    their local text direction and geometry. Sparse whole-image RTL OCR also retries a single-block segmentation mode
    when it recovers more strong-script tokens without materially reducing confidence, and mixed Arabic/Latin scans
    recover identifiers only when a bounded English crop provides compatible evidence.

  • (benchmarks): benchmark provenance is bound to the executable that actually ran. Clean-checkout runs now reject
    stale Xberg binaries whose embedded build identifier does not match the repository commit, and machine-readable
    xberg version output includes that identifier for auditability.

  • (doc): legacy Word documents preserve headers, footers, footnotes, comments, and text boxes. Non-body stories
    are now emitted with their document layers and respect the existing header, footer, and footnote filters instead of
    being dropped whenever the document also contains body text. (GH#2054)

  • (php): Linux packages run on Debian 12 and other glibc 2.36 systems. Release artifacts are built against the
    declared ABI floor, reject newer GLIBC, GLIBCXX, or CXXABI requirements, and are load-tested in Debian 12 before
    publication. (GH#2060)

  • (ocr): columned prose and punctuation fragments are no longer fabricated as tables. Long, sparse newsletter
    columns remain ordinary text in both layout and non-layout extraction, while punctuation-only scan regions fall
    back to paragraph output without losing recognized content.

  • (ocr): layout extraction preserves scanned-PDF recognition settings. Scanned PDFs no longer force Tesseract
    into whole-image single-block segmentation when layout assembly has detected regions, while known scans keep their
    sparse-text segmentation and source-resolution preprocessing hints.

  • (typst): tables wrapped in figures are preserved. A #table(...) nested inside a #figure(...), including
    through an align(...) wrapper, now reaches structured and rendered output instead of being discarded with the
    figure body.

  • (latex): multiline section titles remain headings. A section command whose braced title spans physical lines
    now produces one heading with its label anchor instead of flattened paragraph text and command fragments.

  • (python, pdf): Python post-processors preserve formatted output and hide internal PDF heading annotations.
    Returning an ExtractedDocument from a Python post-processor no longer replaces Rust-only document state, so a
    no-op processor keeps rendered headings and chunks. Font-size annotations used for native PDF heading detection are
    also excluded from plain text and keyword extraction. (GH#2050)

  • (pdf): conflicting JPEG 2000 palette color spaces fail consistently across architectures. Native PDF extraction
    now rejects a codestream palette whose component count conflicts with an explicit /ColorSpace, instead of relying
    on architecture-dependent decoder rounding that could accept the same malformed image on ARM and reject it on x86.

  • (zig): generated bindings compile with Zig 0.17. String duplication now uses the Zig 0.17 sentinel API, and
    plugin vtables cross the C ABI through layout-compatible pointers instead of rejected value bitcasts.

Don't miss a new xberg release

NewReleases is sending notifications on new releases.