What's Changed
- fix(pdf): render pages in parallel so OCR scales with the thread budget by @tobocop2 in #1689
- fix(ocr): resolve the language probe's data directory the way a job does by @tobocop2 in #1697
- fix(ocr): let candle backends ignore the PDF page hints in backend options by @tobocop2 in #1679
- fix(pdf): keep native tables on pages OCR did not replace by @tobocop2 in #1680
- fix(pdf): bound the render batch by the content limit, not the thread budget by @tobocop2 in #1684
- fix(docx): read embedded image bytes when image OCR will run by @tobocop2 in #1686
- test(office): cover embedded-image OCR for PPT and PPTX by @tobocop2 in #1692
- test(ocr): guard parallel page rendering by observing thread dispatch by @tobocop2 in #1693
- fix(ocr): report what the embedded image retry actually did by @tobocop2 in #1698
- fix(ocr): populate the aggregate confidence on the page route and weight both routes alike by @tobocop2 in #1695
- fix(ocr): stop candle engine pools from building duplicate engines by @tobocop2 in #1691
- fix(pdf): repair the provenance lookup and route fabricated mappings to OCR by @tobocop2 in #1678
- fix(ocr): let low OCR confidence lower the quality score by @tobocop2 in #1682
- fix(pdf): run the structured build when document structure is requested by @tobocop2 in #1685
- fix(docx): inherit list numbering from the paragraph style by @tobocop2 in #1688
- fix(candle-ocr): read the whole page with DeepSeek-OCR by @Goldziher in #1702
- fix(ocr): drop implausible-script and bare-LaTeX lines from candle OCR by @Goldziher in #1705
- fix(ocr): drop embedded image bytes read only to feed OCR by @Goldziher in #1706
- fix(candle-ocr): mask the DeepSeek-OCR prefill and resize its rel-pos table by @Goldziher in #1707
- fix(pdf): route a text layer that decodes to the wrong letters to OCR by @Goldziher in #1708
- fix(candle-ocr): stop GLM-OCR penalising repeats per occurrence by @Goldziher in #1704
- ci(rust): cover xberg-candle-ocr in the workspace test and lint jobs by @tobocop2 in #1729
- perf(candle-ocr): read DeepSeek-OCR device tensors in one transfer instead of per element by @tobocop2 in #1715
- fix(ocr): resolve the layout model once per process, not once per page by @tobocop2 in #1720
- fix(pdf): say when the text-plausibility check could not judge a page by @tobocop2 in #1710
- fix(pdf): reject a row-padded pixel buffer instead of panicking by @tobocop2 in #1736
- fix(ocr): report when auto device selection falls back to the CPU by @tobocop2 in #1722
- refactor(pdf): read the native page texts through one ordered helper by @tobocop2 in #1726
- fix(layout): charge each detection batch against the budget, not the whole document by @tobocop2 in #1730
- fix(pdf): extract embedded images in parallel across the thread budget by @tobocop2 in #1734
- fix(pdf): size the per-page OCR batch by free memory, not the content limit by @tobocop2 in #1724
- fix(ocr): charge each page against the image limit, not the whole batch by @tobocop2 in #1741
- perf(candle-ocr): remove two per-element device loops from DeepSeek-OCR by @tobocop2 in #1739
- feat(ocr): let recognition concurrency follow the thread budget and free memory by @tobocop2 in #1733
- fix(pdf): make a page's text independent of the order the pages are read by @tobocop2 in #1737
- perf(pdf): spread scan detection and provenance routing across the thread budget by @tobocop2 in #1743
- fix(pdf): keep merged cells and zebra rows inside ruled tables by @dergachoff in #1766
- fix(pdf): inherit past dangling page attributes by @Goldziher in #1777
- test(pdf): cover single-column bullet list regression by @thisislvca in #1776
- fix(native-pdf): separate numeric superscript markers by @thisislvca in #1773
- fix(ocr): compile the backend lookup only with its consumers by @Goldziher in #1831
- fix(pdf): report parallel render warnings in page order by @Goldziher in #1852
- fix(layout): restore the wasm32 gate on the custom-model dispatch by @Goldziher in #1862
- fix(pdf): decode /Indexed JPEG 2000 images through the PDF palette by @tobocop2 in #1889
- fix(ocr): integrate page-local routing and resilience by @Goldziher in #1936
- fix(tables): integrate OCR and PDF table reconstruction fixes by @Goldziher in #1937
- fix(legacy): restore Word 6/7 and OLE spreadsheet extraction by @Goldziher in #1940
- fix(ocr): keep the scan segmentation mode on scans with unmapped text by @tobocop2 in #1947
- fix(ocr): merge a split amount column across rule marks by @tobocop2 in #1950
- feat(rendering): render extraction output as DOCX by @MannXo in #1955
- fix(ocr): keep widely spaced table rows in their table's region by @Goldziher in #1967
- fix(ocr): stop a PaddleOCR table repeating its text in content by @tobocop2 in #1975
- feat(rendering): render extraction output as PDF by @MannXo in #1966
- fix(build): gate OCR test helpers for the formula-recognition leg by @Goldziher in #1962
- fix(ocr): return Tesseract line elements at the default element level by @tobocop2 in #1981
- fix(pdf): preserve grouped header spans through extraction by @thisislvca in #1960
- fix(dbf): keep record values in field declaration order by @ihibti in #1969
- fix(ocr): satisfy clippy in WordData tests by @Goldziher in #1983
- fix(pdf): keep table-cell word spacing with the page text by @Goldziher in #1964
- fix(pdf): retain cells beside partial row separators by @thisislvca in #1961
- fix(ocr): fold a split amount column on a scan despite junk tokens by @tobocop2 in #1954
- fix(ocr): preserve image list source numbers by @Goldziher in #1984
- fix(pdf): preserve numeric footnotes in structured paragraphs by @thisislvca in #1956
- fix(transcription): stop a bad 30s window from failing the whole file by @Goldziher in #1963
- fix(ocr): keep tall word boxes in table rows by @Goldziher in #1985
- fix(wasm): return OCR word elements by @Goldziher in #1986
- docs(ocr): document that OCR elements are opt-in, for every binding by @tobocop2 in #1972
- fix(ocr): retain label and value tables by @Goldziher in #1987
- fix(mime): reject array-like text as JSON by @Goldziher in #1988
- fix(codegen): restore Alef generation freshness by @Goldziher in #1989
- fix(cache): preserve structured extraction entries by @Goldziher in #1992
- test(benchmark): provision extraction test stack by @Goldziher in #1994
- ci: verify formatted Alef output by @Goldziher in #1995
- ci: align Alef Java formatter toolchain by @Goldziher in #1996
- fix(ci): scope clang-format to C sources by @Goldziher in #1999
- fix(ci): pin Alef Java formatters by @Goldziher in #2000
- fix(ocr): fold a header-only label into the amount column it ends with by @tobocop2 in #2007
- test(ocr): make the line-row twins depend on the Tesseract line again by @tobocop2 in #2013
- fix(ocr): fold a half-empty unnamed amount column, keep the table by @tobocop2 in #2020
- fix(ocr): restore shaded tables and fallback intent by @Goldziher in #2022
- fix(odt): extract text boxes and report encryption by @Goldziher in #2023
- fix(excel): harden spreadsheet parsing by @Goldziher in #2024
- fix(config): discover project YAML and JSON by @Goldziher in #2025
- fix(ocr): stop the table fold subtracting from column 0 by @tobocop2 in #2027
- fix(ocr): keep embedded-image retry text beside a table that claimed no page text by @tobocop2 in #2009
- fix: batch OCR, cache, and macOS release repairs by @Goldziher in #2034
New Contributors
- @thisislvca made their first contribution in #1776
- @ihibti made their first contribution in #1969
Full Changelog: v1.2.4...v1.3.3