github xberg-io/xberg v1.3.3

3 hours ago

What's Changed

  • fix(pdf): render pages in parallel so OCR scales with the thread budget by @tobocop2 in #1689
  • fix(ocr): resolve the language probe's data directory the way a job does by @tobocop2 in #1697
  • fix(ocr): let candle backends ignore the PDF page hints in backend options by @tobocop2 in #1679
  • fix(pdf): keep native tables on pages OCR did not replace by @tobocop2 in #1680
  • fix(pdf): bound the render batch by the content limit, not the thread budget by @tobocop2 in #1684
  • fix(docx): read embedded image bytes when image OCR will run by @tobocop2 in #1686
  • test(office): cover embedded-image OCR for PPT and PPTX by @tobocop2 in #1692
  • test(ocr): guard parallel page rendering by observing thread dispatch by @tobocop2 in #1693
  • fix(ocr): report what the embedded image retry actually did by @tobocop2 in #1698
  • fix(ocr): populate the aggregate confidence on the page route and weight both routes alike by @tobocop2 in #1695
  • fix(ocr): stop candle engine pools from building duplicate engines by @tobocop2 in #1691
  • fix(pdf): repair the provenance lookup and route fabricated mappings to OCR by @tobocop2 in #1678
  • fix(ocr): let low OCR confidence lower the quality score by @tobocop2 in #1682
  • fix(pdf): run the structured build when document structure is requested by @tobocop2 in #1685
  • fix(docx): inherit list numbering from the paragraph style by @tobocop2 in #1688
  • fix(candle-ocr): read the whole page with DeepSeek-OCR by @Goldziher in #1702
  • fix(ocr): drop implausible-script and bare-LaTeX lines from candle OCR by @Goldziher in #1705
  • fix(ocr): drop embedded image bytes read only to feed OCR by @Goldziher in #1706
  • fix(candle-ocr): mask the DeepSeek-OCR prefill and resize its rel-pos table by @Goldziher in #1707
  • fix(pdf): route a text layer that decodes to the wrong letters to OCR by @Goldziher in #1708
  • fix(candle-ocr): stop GLM-OCR penalising repeats per occurrence by @Goldziher in #1704
  • ci(rust): cover xberg-candle-ocr in the workspace test and lint jobs by @tobocop2 in #1729
  • perf(candle-ocr): read DeepSeek-OCR device tensors in one transfer instead of per element by @tobocop2 in #1715
  • fix(ocr): resolve the layout model once per process, not once per page by @tobocop2 in #1720
  • fix(pdf): say when the text-plausibility check could not judge a page by @tobocop2 in #1710
  • fix(pdf): reject a row-padded pixel buffer instead of panicking by @tobocop2 in #1736
  • fix(ocr): report when auto device selection falls back to the CPU by @tobocop2 in #1722
  • refactor(pdf): read the native page texts through one ordered helper by @tobocop2 in #1726
  • fix(layout): charge each detection batch against the budget, not the whole document by @tobocop2 in #1730
  • fix(pdf): extract embedded images in parallel across the thread budget by @tobocop2 in #1734
  • fix(pdf): size the per-page OCR batch by free memory, not the content limit by @tobocop2 in #1724
  • fix(ocr): charge each page against the image limit, not the whole batch by @tobocop2 in #1741
  • perf(candle-ocr): remove two per-element device loops from DeepSeek-OCR by @tobocop2 in #1739
  • feat(ocr): let recognition concurrency follow the thread budget and free memory by @tobocop2 in #1733
  • fix(pdf): make a page's text independent of the order the pages are read by @tobocop2 in #1737
  • perf(pdf): spread scan detection and provenance routing across the thread budget by @tobocop2 in #1743
  • fix(pdf): keep merged cells and zebra rows inside ruled tables by @dergachoff in #1766
  • fix(pdf): inherit past dangling page attributes by @Goldziher in #1777
  • test(pdf): cover single-column bullet list regression by @thisislvca in #1776
  • fix(native-pdf): separate numeric superscript markers by @thisislvca in #1773
  • fix(ocr): compile the backend lookup only with its consumers by @Goldziher in #1831
  • fix(pdf): report parallel render warnings in page order by @Goldziher in #1852
  • fix(layout): restore the wasm32 gate on the custom-model dispatch by @Goldziher in #1862
  • fix(pdf): decode /Indexed JPEG 2000 images through the PDF palette by @tobocop2 in #1889
  • fix(ocr): integrate page-local routing and resilience by @Goldziher in #1936
  • fix(tables): integrate OCR and PDF table reconstruction fixes by @Goldziher in #1937
  • fix(legacy): restore Word 6/7 and OLE spreadsheet extraction by @Goldziher in #1940
  • fix(ocr): keep the scan segmentation mode on scans with unmapped text by @tobocop2 in #1947
  • fix(ocr): merge a split amount column across rule marks by @tobocop2 in #1950
  • feat(rendering): render extraction output as DOCX by @MannXo in #1955
  • fix(ocr): keep widely spaced table rows in their table's region by @Goldziher in #1967
  • fix(ocr): stop a PaddleOCR table repeating its text in content by @tobocop2 in #1975
  • feat(rendering): render extraction output as PDF by @MannXo in #1966
  • fix(build): gate OCR test helpers for the formula-recognition leg by @Goldziher in #1962
  • fix(ocr): return Tesseract line elements at the default element level by @tobocop2 in #1981
  • fix(pdf): preserve grouped header spans through extraction by @thisislvca in #1960
  • fix(dbf): keep record values in field declaration order by @ihibti in #1969
  • fix(ocr): satisfy clippy in WordData tests by @Goldziher in #1983
  • fix(pdf): keep table-cell word spacing with the page text by @Goldziher in #1964
  • fix(pdf): retain cells beside partial row separators by @thisislvca in #1961
  • fix(ocr): fold a split amount column on a scan despite junk tokens by @tobocop2 in #1954
  • fix(ocr): preserve image list source numbers by @Goldziher in #1984
  • fix(pdf): preserve numeric footnotes in structured paragraphs by @thisislvca in #1956
  • fix(transcription): stop a bad 30s window from failing the whole file by @Goldziher in #1963
  • fix(ocr): keep tall word boxes in table rows by @Goldziher in #1985
  • fix(wasm): return OCR word elements by @Goldziher in #1986
  • docs(ocr): document that OCR elements are opt-in, for every binding by @tobocop2 in #1972
  • fix(ocr): retain label and value tables by @Goldziher in #1987
  • fix(mime): reject array-like text as JSON by @Goldziher in #1988
  • fix(codegen): restore Alef generation freshness by @Goldziher in #1989
  • fix(cache): preserve structured extraction entries by @Goldziher in #1992
  • test(benchmark): provision extraction test stack by @Goldziher in #1994
  • ci: verify formatted Alef output by @Goldziher in #1995
  • ci: align Alef Java formatter toolchain by @Goldziher in #1996
  • fix(ci): scope clang-format to C sources by @Goldziher in #1999
  • fix(ci): pin Alef Java formatters by @Goldziher in #2000
  • fix(ocr): fold a header-only label into the amount column it ends with by @tobocop2 in #2007
  • test(ocr): make the line-row twins depend on the Tesseract line again by @tobocop2 in #2013
  • fix(ocr): fold a half-empty unnamed amount column, keep the table by @tobocop2 in #2020
  • fix(ocr): restore shaded tables and fallback intent by @Goldziher in #2022
  • fix(odt): extract text boxes and report encryption by @Goldziher in #2023
  • fix(excel): harden spreadsheet parsing by @Goldziher in #2024
  • fix(config): discover project YAML and JSON by @Goldziher in #2025
  • fix(ocr): stop the table fold subtracting from column 0 by @tobocop2 in #2027
  • fix(ocr): keep embedded-image retry text beside a table that claimed no page text by @tobocop2 in #2009
  • fix: batch OCR, cache, and macOS release repairs by @Goldziher in #2034

New Contributors

Full Changelog: v1.2.4...v1.3.3

Don't miss a new xberg release

NewReleases is sending notifications on new releases.