github xberg-io/xberg v1.3.0

3 hours ago

[1.3.0] - 2026-09-28

Added

  • (ocr): PDF pages with an unusable text-layer character map use Tesseract block segmentation per page. During automatic OCR fallback, only pages routed by the fabricated-character-map detector use PSM 6 so table row labels and values stay together; ordinary scans and explicit force_ocr or force_ocr_pages requests retain their existing mode, and an explicit caller setting still wins. Accepted Tesseract pages record their effective mode in metadata.additional.ocr_page_segmentation_modes. (GH#1896)
  • (python): extract and extract_batch accept an optional on_progress callback for per-page OCR progress. Each ProgressEvent reports the 1-based page number, total pages, completion sequence, OCR backend, and the zero-based input index for batch extraction. Callbacks run off the extraction worker; a callback exception is logged and ignored so extraction can finish. (GH#1887)
  • (bindings): pdf_page_count is available again in every language binding. It returns the number of pages in a PDF without rendering any of them, takes an optional password for an encrypted file, and raises an error on bytes that are not a PDF. The 1.0.0 notes list it as the cheap way to size a render loop, but a later regeneration dropped it from the bindings, so it existed only in Rust. It now ships in Python, TypeScript (Node and WASM), Ruby, PHP, Go, Java, C#, Elixir, Dart, Kotlin, Swift, Zig and the C FFI. (GH#1869)
  • (bindings): the published Python, Node, PHP, Ruby and Elixir packages compile the pdfium PDF backend in. PdfBackend::Pdfium is documented in every one of their generated API references, but selecting it always failed: the Cargo feature that compiles the engine was enabled for no published artifact. These packages still do not bundle libpdfium -- supply it yourself and point PDFIUM_DYNAMIC_LIB_PATH at the directory holding it, as described under "Pdfium PDF backend" in the installation guide. The engine reaches the Linux (x86_64, aarch64) and macOS Apple Silicon artifacts only; the Windows and Intel-macOS packages build from a reduced feature set, and WebAssembly, Android and iOS cannot load a shared library at all. native remains the default backend and nothing changes for a caller that does not select pdfium. (GH#1863)
  • (tesseract): ResultIterator::extract_word_symbols reads the text and box of each symbol in the words that a filter accepts. It walks the page under one lock, like extract_all_words. (GH#1833)
  • (ocr): OcrBackend::supported_languages_for and supports_language_for report a backend's languages under a specific OcrConfig. list_ocr_backend_capabilities_for and ocr_backend_supports_language_for are the corresponding config-aware forms of the existing config-less functions. Use these when the caller sets OcrConfig.tessdata_path: for Tesseract the language list is a property of the resolved tessdata directory, and the config-less forms always answer for the no-override directory. (GH#1857)

Changed

  • (wasm): assigning a class-typed property no longer destroys the handle you assigned, and clearing an optional one now needs clearX(). wasm-bindgen lowers a by-value exported-struct argument through __destroy_into_raw(), so options.field = handle left the caller holding a dead handle: reusing that handle, or assigning it to a second options object, threw null pointer passed to rust. Such a setter now takes a borrow and stores a clone, so the handle stays alive. Because Option<&T> has no OptionFromWasmAbi implementation, an optional field of a generated class type takes the same borrow and wraps it in Some, and a generated clear{Field}() companion is what unsets it -- opts.x = null now throws rather than clearing the field, which is the one breaking change for existing callers. The generated .d.ts accessor pair for such a field is asymmetric (get x(): T | undefined; set x(value: T);) and therefore requires TypeScript 5.1 or newer. Separately, a trait-bridge options field assigned as a property was blanked before the call and so had no effect at all; it is now the fallback when no explicit visitor argument is given. (alef #470)
  • (ocr): TesseractConfig::thresholding_method is an integer selecting the binarization algorithm, where it used to be a bool. The field was sent to Tesseract as the string "true" or "false", which its integer parameter parser cannot read -- and SetVariable reports no error for a value it fails to parse, so the check around that call never fired. The setting therefore did nothing at all, and methods 1 (LeptonicaOtsu) and 2 (Sauvola) were unreachable: on the reporting document Tesseract's own CLI reads 135 of 138 values at method 1 where xberg read 120. Valid values are 0 (Otsu, the default and the previous effective behaviour), 1 and 2, and they are now validated rather than passed through. Configuration files and any caller that supplies a map or dictionary are unaffected -- a legacy false still deserializes to 0 and true to 1, the method its old documentation described. This is a breaking change only for callers that assign the field in typed code, in Rust or in a generated binding, where true becomes 1. (GH#1784)
  • (pdf): metadata.pages is now populated for ocr_near_empty_fallback or ocr_scanned_page_quality_gate set without an ocr block, at parity with what an ocr block alone already produced. The annotation fallback's implicit page-boundary tracking was discarded from metadata.pages based on a narrower, re-derived condition than the one that decides whether boundaries are tracked in the first place, so a caller relying on either GH#1752 setting alone still got metadata.pages == None. (GH#1752)

Fixed

  • (pdf): a borderless table beside a prose column is no longer reconstructed across the page gutter. The heuristic table pass now separates a dominant whitespace corridor before column detection, while rejecting a split that leaves labels on one side and values on the other so financial tables keep their rows. Recurring numeric tracks keep a real table with wrapped row labels from being rejected as flowing prose, and full-width blocks below the columns remain intact. (GH#1769)
  • (ocr): a multi-word table header now stays in the data column it labels. Header words that had already merged into one cell for column detection were split apart again during assignment, so a trailing word could move into the next header. A merged header is now assigned as one cell when its detected column has data rows, while header-only spans keep the existing fragment repair. (GH#1934)
  • (ocr): shaded-row normalization no longer loses dark and mid-grey table rows. Overlapping inversion boxes could invert part of a dark row twice, while normalized band edges could be read as [ or | frame glyphs and hide otherwise recognized rows during table reconstruction. Overlapping boxes are now inverted once, and repeated normalization artifacts at a table's outer edges are removed before cell assignment. (GH#1837)
  • (excel): legacy XLS and XLA compound files no longer undergo ZIP-bomb validation as outer ZIP archives. An embedded ZIP whose central directory could be discovered from the OLE bytes made security accounting seek its local headers against the wrong container and reject the workbook before the XLS parser ran. The bypass requires both a legacy spreadsheet extension and the OLE signature; ZIP-backed spreadsheet formats retain the configured archive limits. (GH#1938)
  • (doc): Word 6/95 binary documents now reach the legacy text extractor. The parser accepts the 0xA5DC signature, recognizes legacy nFib versions before opening a Word 97 table stream, and reads the legacy ccpText field from its actual header offset. Fast-saved Word 6/95 files remain unsupported and now return an explicit remedy instead of being misparsed as Word 97. (GH#1939)
  • (ocr): a shaded table row split by its own baseline gap is no longer struck through by the shaded-row normalisation. A dark-fill row is found as a run of rows that read mostly dark, and the baseline gap of a line of capitals can read light enough to end that run and start a new one, so one printed row is found as two bands. The padding that grows each band onto the fill's anti-aliased edge then grew both halves back across the gap, and the rows they both covered were inverted twice -- returning them to the fill's own dark value as a rule drawn through the middle of the row's text. On the reference scan that rule crossed a totals row, which lost its label and read its 30,474 as 20,474, and it disturbed the plain row above enough that its label came back as a stray fragment plus a misread word. Each row is now inverted exactly once. Measured on that scan, ground-truth values landing in the correct cell rose from 102 to 108 of 138 at the segmentation mode the PDF OCR route selects by default, with its light-fill and plain rows both becoming complete; dark and mid-fill rows are unaffected and remain the subject of GH#1837. (GH#1918)
  • (ocr): the automatic Tesseract-to-Paddle fallback is decided independently for each PDF page. Whole-document OCR previously scored the pages' joined text, so one readable page could hide an unreadable page and prevent Paddle from running; selected-page OCR silently skipped the automatic pipeline and ran Tesseract alone. Both routes now stop at Tesseract for a good page and invoke Paddle only for a page below the configured quality threshold. (GH#1908)
  • (ocr): automatic OCR fallback keeps a valid native page when other pages in the document are scans. When trustworthy page boundaries were available, the quality gate still evaluated the whole-document average first. A document with one native page and nine empty scanned pages therefore sent all ten pages through OCR, replacing the native page and its tables. The gate now decides each bounded page independently and uses whole-document scoring only when page boundaries are absent or untrustworthy. (GH#1927)
  • (ocr): scanned-page settings now survive detached-image OCR pipeline calls. PDF pages selected by force_ocr_pages, scan detection, or per-page fallback now keep the same source DPI, scan preprocessing signal, and default Tesseract segmentation mode as whole-document OCR when layout detection, vlm_fallback, or an explicit pipeline is active. Standalone image pipelines likewise use the image's embedded or inferred DPI and whole-image mode. (GH#1926)
  • (ocr): an inset scanned page with more than 400 unmapped Type 3 glyphs is still recognised as a scan. The inset-raster check counted every extracted glyph as readable native text, including fallback characters fabricated for a font with no usable Unicode mapping, so a broken OCR sidecar could make the page look like body text beside a figure. The check now counts only mapped glyphs toward its native-text limit; a page with more than 400 genuinely mapped glyphs remains a text page. (GH#1921)
  • (ocr): one selected PDF page's OCR failure no longer discards successful OCR from the other pages. Render, upright-rotation, backend, and pipeline failures are isolated to their page with a page-numbered warning. An explicit force_ocr_pages request still returns an error when every requested page fails; automatic per-page OCR keeps the native result. (GH#1925)
  • (ocr): PDF pages rendered for OCR now lower their DPI before exceeding the configured content-size limit. The render chooser uses the same live-byte accounting as the later PNG encode check and reports the page and applied cap, so a supported high DPI or a legal-size scan no longer aborts extraction under the default limit. (GH#1924)
  • (ocr): PDF page rasters are no longer OCR'd a second time as embedded images. Enabling images.include_page_rasters now keeps the raster in the result without appending a duplicate copy of the page's OCR text to document content. (GH#1929)
  • (ocr): right-aligned amount columns of mixed width in a scanned table keep their values apart. When a one-digit amount of one column starts close to where an eight-digit amount of the next column starts, column detection put both into one column. Two amounts of one row then shared a cell, and the table was dropped or showed values in the wrong column. A column that holds two cells of one row is now split by right edge, and a header label no longer stops the value tracks under it from folding into one column. (GH#1909)
  • (tables): an amount column with thousands separators is cleaned up like a column of plain digits. Table cleanup empties a nil dash and joins - 5,000 to -5,000 only in a column it reads as numeric, and it did not read 40,218,965 as a number. Two amount columns of one table could therefore show the same nil dash as empty in one and as - in the other. Comma-grouped amounts such as 1,234,567 and 1,234.50 now count as numbers; 1,2 and 12,34 still do not. (GH#1914)
  • (ocr): numbered row labels no longer make table reconstruction discard the whole table. OCR can split labels such as Item 1 into a word column and a narrow numeric tail column, making the tail look like an independent field and causing an otherwise complete table to be rejected. A short numeric tail is now folded into its adjacent label only when every row resolves to the same repeated label stem; genuine numeric and amount columns remain separate. (GH#1928)
  • (ocr): a dark scanned PDF page now gets the same default OCR preprocessing as a lighter page in the same file. With no tesseract_config set, the OCR step judged every image by its pixels alone -- mean luminance and near-white pixel fraction -- to decide whether to apply the default resample to 300 dpi, binarisation and deskew. A scanned page with shaded table rows or a grey background failed that test and reached Tesseract at render resolution with no binarisation, while a lighter page in the same document was preprocessed and read its table values correctly. The PDF OCR route already knows which pages are whole-page scans from its own scan-detection check; that signal now reaches image preparation directly, so a known scan page always takes the default preprocessing regardless of how dark it reads. A bare image handed to the standalone image extractor, which has no such signal, is unaffected and still judged by its pixels. (GH#1894)
  • (ocr): the OCR retry on a scanned page's embedded image uses the same Tesseract settings as the page. When the OCR of a whole PDF document (force_ocr, or a document that needs OCR throughout) comes back blank or fails on a page, xberg retries OCR on the page's embedded images, and that retry used the configured OCR settings unchanged. The embedded image of a page that is one full-page scan therefore reached Tesseract without the whole-image page segmentation mode and without the scan signal that gives the page the default preprocessing, although the same page gets both when it renders. The retry now takes both on a scan page. A configured psm or preprocessing still wins, and a page that is not a scan and the other OCR backends are unchanged. (GH#1907)
  • (ocr): OCR of selected pages retries a blank or failed page on its embedded images. When the OCR of a page comes back blank or fails, the whole-document OCR route retries OCR on the page's embedded images. The per-page route did not. That route runs for force_ocr_pages, for pages that scan detection marks as scanned, and for the per-page OCR fallback. With the default settings it had no retry, and with a pipeline, a vlm_fallback policy or layout detection the retry found no document to read the images from. So a scanned page that the renderer could not draw came back blank, with no warning. The per-page route now retries in every setting and adds the same warnings as the whole-document route. A failed page that the retry cannot recover still fails the extraction. (GH#1912)
  • (ocr): the psm documentation of TesseractConfig describes the default for a scanned PDF page. It said that a rendered PDF page always gets the engine's automatic-layout mode and suggested setting psm: 11 for table pages. Since GH#1786, a page that is one full-page scan gets the whole-image mode that image OCR uses when no psm is set. Only a page that is not a scan keeps the engine default. The Rust docs, every binding's docs and the API reference now say this. (GH#1911)
  • (pdf): a colour-key /Mask on an /Indexed image is applied again. A colour-key mask states its ranges in the image's raw index space, but extraction expands an Indexed image's indices to RGB for display and always cleared the flag the renderer's colour-key path used to decide whether masking was still possible, so the mask was silently skipped and pixels the file marked transparent were painted with their palette colour instead. The raw index plane is now kept alongside the palette-expanded RGB, and the colour-key path reads it for an Indexed image rather than giving up. (GH#1899)
  • (pdf): an /Indexed image now reaches the separation-plate renderer. The plate classifier had no case for /Indexed, so a palette image resolved to Unknown and was skipped before it was even decoded, whatever ink intent its palette's base colour space carried -- a DeviceCMYK-, Separation- or DeviceN-based palette painted no plate at all. An /Indexed colour space is now classified by its base, and when that base carries ink intent the palette lookup's own base-space bytes -- not the RGB the renderer produces for on-screen display -- route straight to its plates. (GH#1898)
  • (pdf): counts.pages now reports a PDF's real page count instead of the number of pages that happened to produce content. The count came from page-boundary tracking, which only runs when an explicit pages config or one of a handful of OCR settings switches it on, and which in any case only counts pages that yielded text. A scanned PDF extracted with disable_ocr therefore reported 0 pages, and a PDF with one native page beside one scanned page reported 1 of its 2, while the same result's PDF metadata carried the correct count throughout. The format's own page count, read from the page tree independently of any tracking, is now a floor on the reported number. (GH#1888)
  • (pdf): a JPEG 2000 image whose dictionary /Width or /Height disagrees with its codestream now extracts at the codestream's size. A /JPXDecode image's pixel buffer is always sized to the codestream, but the image reported the dictionary's declared dimensions regardless -- a smaller declared size read a scrambled prefix of the buffer at the wrong stride, and a larger one made the buffer come up short for anything that trusted width * height * components, so the image came out silently garbled or dropped, with nothing recorded either way. ISO 32000-1:2008 §7.4.9 treats /Width and /Height as informative for JPEG 2000, so the codestream's own SIZ marker segment is now authoritative for both, and a disagreement between the two now raises a warning. (GH#1900)
  • (ocr): a table on a resampled scanned page is placed inside the page instead of spilling off it. The mixed-OCR route converted a table's bounding box from raster pixels to PDF points using the size of the page render, but OCR preprocessing resamples a sub-300 dpi render up to target_dpi before recognition, so the box comes back in the resampled frame. Dividing by the render's size overstated every table box by exactly that ratio -- a table on a 150 dpi render came out at twice its size, with part of it outside the page. The element boxes on the same route already resolved the processed frame, which is why only tables were affected. (GH#1895)
  • (ocr): layout.table_model = "disabled" now turns off table recognition on the force-OCR PDF route too. The route gated TATR on whether a layout pass had run and nothing else, so the setting was honoured on the native PDF route and for standalone images and silently ignored for every force_ocr extraction. The SLANet values are unchanged on this route: they are wired only on the native route, so treating them as "not TATR" here would remove table recognition outright rather than substitute a model. (GH#1813)
  • (ocr): a table a backend reported itself is now placed using the pixel frame its coordinates are in. The bounding box was rescaled to PDF points using the page render's dimensions while the backend's coordinates are in its own processed-image space, which on a Tesseract page is roughly twice as large in each axis, so the table landed at about half its true position and size. The formula path on the same route already resolved the processed frame; the table path now does the same. (GH#1813)
  • (ocr): a page whose OCR word boxes cannot be mapped into the render raster now yields its text as prose instead of an empty table. The transform answered three untrustworthy metadata shapes -- no processed dimensions reported, an auto-rotation with no usable orientation, and a raster that is not a scaled copy of the render -- by handing back untransformed boxes. Those sit outside the render-space table bounding box, so the cell-overlap filter discarded every one of them and the region published a fully shaped table holding none of the page's values. Such a page now skips table recognition and records a warning, and its words stay in the paragraph stream. A scale pair whose two axes disagree by more than 1% is rejected under the same tolerance the standalone-image route has always applied. (GH#1813)
  • (ocr): numeric_repair now runs on whole-document OCR, not only on per-page OCR. GH#1840 extended the repair across the mixed OCR/native route's outputs, but the whole-document route -- the one force_ocr, the near-empty Auto fallback and the OCR gate's whole-document fallback all take -- had no repair at all, so {numeric_repair: true, force_ocr_pages: [1, 2]} repaired those pages while {numeric_repair: true, force_ocr: true} repaired nothing from the same flag. That route's flat text, structured document and reconstructed tables are now repaired too, under the same opt-in flag and the same markup-renderer exclusion, and ahead of formula recognition so recognised LaTeX is not re-punctuated. The repair's documented limitation is unchanged and now reaches this route as well: a bare four-digit year with no FY prefix is grouped as a thousands value. (GH#1892)
  • (ocr): the Tesseract language probe now asks for the languages the caller configured, and its memo tells apart callers that resolve different directories. GH#1857 threaded OcrConfig.tessdata_path into the probe but left it asking for English, and a candidate tessdata directory is accepted only when it holds every language asked for -- so a tessdata_path holding just the configured language was rejected, resolution fell through to TESSDATA_PREFIX, the xberg cache and the system paths, and supports_language_for/supported_languages_for denied a language every real job with that config loads. The probe now resolves with the config's effective languages, and only against directories that already hold them, so a capability query still cannot trigger a language-pack download. Its memo is keyed on those languages together with the override, TESSDATA_PREFIX and XBERG_CACHE_DIR rather than on the override alone, so two callers whose resolution differs in any of them no longer serve each other's answer. supports_language without a config now probes for the language it was asked about; supported_languages without a config still answers for the English datapath. (GH#1891)
  • (pdf): an implausible text page the automatic OCR fallback will not route now always reports a warning. The warning's condition was a second, re-derived form of the predicate that decides whether the near-empty Auto fallback runs. The two agreed until ocr_near_empty_fallback gained an explicit false (GH#1752): with an ocr block and the fallback switched off, pages flagged as reading like no real language were not routed to OCR and the warning saying so was suppressed, so the caller received wrong-mapped text with no diagnostic at all. ocr_near_empty_fallback: true with no ocr block, on a build with no registered automatic OCR backend, diverged the same way. The warning now calls the routing predicate itself rather than re-deriving it, and its text no longer names a missing ocr block as the only possible cause. (GH#1890)
  • (ocr): a right-aligned amount column is no longer split into two or three columns by digit count. Table column detection grouped every word by its left edge, but an amount column is right-aligned: 5, 73 and 1,234,567 share a right edge and their left edges differ by far more than table_column_threshold at 300 dpi. One column minted several, and what the existing disjoint-numeric-column merge did not fold back came out as a dropped trailing nil-dash column, a sparse row's missing interior cells, or a value in the label cell. Adjacent columns whose every word reads as a value -- a number, or a symbol a cell prints alone such as a nil dash -- now fold into one when their right edges coincide within the same threshold. The threshold itself is unchanged, since widening it merges genuinely narrow neighbouring columns. Folding changes only how many columns the grid has: a folded column keeps the left edge of every track it absorbed and still decides which words belong to it on the nearest of those, so no word moves between columns. (GH#1886)
  • (ocr): a reconstructed table no longer replaces its region's lines when it dropped their words. On the Tesseract route every hOCR line under a detected table's bounding box was removed from the page document unconditionally, with no check that the table's cells still carried those lines' words. A reconstruction that splits a wrapped label across two rows, glues a value onto a label, or loses the interior cells of a sparse row therefore deleted those words from the page: they were in neither the paragraphs nor the table. A table now claims its region only when its non-empty cells retain at least as many words as the lines it would remove -- the same absolute word-count rule the standalone-image rebuild already applied -- and a table that lost words leaves those lines in place while still being published in tables, so the structured form is never lost either. The cost for such a region is that its text appears both as paragraphs and as a table. (GH#1884)
  • (pdf): a bare CMYK JPEG 2000 codestream in an image that names no /ColorSpace keeps its black plane. ISO 32000-1 Table 89 lets a /JPXDecode image omit /ColorSpace, and such an image is given a placeholder /DeviceRGB until the codestream reveals its real component count. That placeholder was handed to the decoder as though the document had declared three components, which agreed with the decoder's own reading of a bare four-component codestream as RGB plus alpha, so the fourth plane was discarded as transparency: CMY was painted as RGB on the render path, and the image was dropped from every plate on the separation path. The component count now comes from the codestream's own SIZ header when the dictionary declares nothing. A bare codestream cannot mark a component as alpha -- that requires the JP2 channel-definition box -- so all of its components are now read as colour components; the consequence is that a bare four-component codestream that really was RGBA, in an image declaring no /ColorSpace, is read as CMYK and loses its alpha. A PDF carries transparency in /SMask, not in the codestream. Images that declare a /ColorSpace, and JP2-boxed streams, are unaffected. (GH#1883)
  • (pdf): a JPEG 2000 image with an /Indexed colour space now renders and extracts through its palette. The decoder applied the palette box inside the JPEG 2000 data and failed the whole image when lossy coding pushed an index sample outside that palette, so a scanned page stored this way rendered blank, and image extraction dropped the image. ISO 32000-1 §7.4.9 puts the dictionary's /ColorSpace in charge, so the decoder now reads the raw indices, clamps each to the last entry of the palette, and looks it up in the dictionary's /Indexed palette, as it already did for other index streams. An /Indexed JPEG 2000 image with no palette box of its own was painted as greyscale index values and is now looked up too. An /Indexed JPEG 2000 image the decoder did resolve used to extract as CMYK and now extracts as RGB, like every other /Indexed image. An opacity channel in the JPEG 2000 data is dropped, as it already was. Such an image extracts at the JPEG 2000 data's own size when the dictionary's /Width or /Height disagrees, and a colour-key /Mask on it is tested against its indices. (GH#1885)
  • (pdf): a lossy palette JPEG 2000 image with no /Indexed colour space now decodes. When the JPEG 2000 data carries its own palette and the dictionary names no /Indexed colour space, the decoder resolved that palette with an unclamped lookup, so one index sample that lossy coding pushed past the palette failed the whole image and the image was dropped. xberg now reads the palette from the JPEG 2000 header and resolves the index plane through the same clamped lookup it uses for /Indexed images. The image keeps the colour space of the palette, so a CMYK palette extracts as CMYK and paints the process plates. The separation renderer now uses that decoded colour space when the image dictionary omits /ColorSpace, instead of classifying the image as unknown and skipping it before decoding. This applies when the dictionary names no /ColorSpace, or names a device colour space with the palette's component count. A declared colour space that the palette does not fit still fails as before, and a palette with entries deeper than 8 bits keeps the decoder's own lookup. (GH#1903, GH#1922)
  • (pdf): an expanded /Indexed image reports 8 bits per component. After an /Indexed image's indices were expanded to 8-bit colour samples, bits_per_component still reported the dictionary's 1, 2 or 4-bit depth, so the field disagreed with the pixel data. It now reports 8, for uncompressed index streams and for /Indexed JPEG 2000 images. (GH#1901)
  • (pdf): a transparent pixel of a drawn image shows the page under it. The renderer drew each image into the page with its opacity stored beside unchanged colour, but the page canvas expects colour already scaled by opacity. A pixel that a soft mask (/SMask) or a colour-key /Mask made transparent therefore added its own colour to the page under it: a white image with a transparent region painted that region white on any background, and a half-transparent pixel came out too light. xberg now scales each image pixel's colour by its opacity before drawing it, and an image drawn smaller than its own size is resized with that scaled colour. (GH#1905)
  • (pdf): a stencil image mask no longer paints its whole rectangle on a separation plate. The separation renderer drew an /ImageMask stencil with the ink tint as unscaled colour beside the stencil's opacity, but the plate canvas expects colour already scaled by opacity. Each pixel the stencil leaves transparent therefore added the full tint to the plate under it, so the whole rectangle of the mask read as ink. xberg now scales the tint by the stencil opacity, and the plate under an unmarked pixel keeps its value. (GH#1906)
  • (pdf): a JPEG 2000 image with /SMaskInData 1 or 2 now renders with its own transparency. When the image dictionary sets /SMaskInData to 1 or 2, the opacity channel inside the JPEG 2000 data is the image's soft mask, but the renderer dropped that channel and drew the image opaque. xberg now keeps the channel and applies it as the soft mask when the image has no /SMask entry. An /SMask entry still takes precedence, and /SMaskInData 0 or absent leaves the channel unused as before. Image extraction is unchanged. (GH#1902)
  • (pdf): a colour-key /Mask on a 1, 2 or 4-bit image, or on an image with a /Decode array, is now applied. The ranges of a colour-key /Mask are written in the image's own sample values, but the extractor unpacks 1, 2 and 4-bit samples to 8 bits and folds a /Decode array into them, so the renderer skipped the mask and painted the pixels the file marks transparent. xberg now keeps the raw samples of such an image beside the unpacked ones and tests the mask against them, for every colour space. A colour-key /Mask on a 16-bit image is still skipped. (GH#1904)
  • (pdf): an /Indexed palette over a DeviceN colour space is now read one byte per colorant. The extractor read each palette entry as four bytes whatever the number of colorants, so a palette over a DeviceN space with two, three, five or more colorants mixed neighbouring entries. Each entry now holds one byte per colorant the space names, and such an image paints its colorant plates. (GH#1913)
  • (pdf): a JPEG 2000 /Indexed image now paints the separation plates. The separation renderer looked up the palette only for an uncompressed index stream, so a JPEG 2000 coded /Indexed image left every plate empty. Its index plane is now looked up in the palette the same way, at the JPEG 2000 data's own size. (GH#1916)
  • (pdf): a spot ink used only through an /Indexed palette now gets its separation plate. The plate list for a page comes from the spot inks its colour spaces name, and that list never looked inside an /Indexed colour space. An image whose palette is over a /Separation or /DeviceN space therefore got no plate for those inks, and its ink was lost. The base of an /Indexed colour space now names its inks like any other space. (GH#1915)
  • (docker): images no longer share OCR cache entries across builds at the same crate version. The cache key folds in a per-build discriminator derived from git rev-parse HEAD, but .dockerignore excludes **/.git/, so the derivation always failed inside a Docker build and the discriminator fell back to an empty string -- collapsing every image and published artifact built at one version into a single cache key space, including a force_republish rebuild at an unchanged version. Every Dockerfile that compiles Rust -- the core, full and cli images, and the eight that build published artifacts (the musl xberg-cli and cargo binstall binaries, the musl Python wheel, libxberg_ffi for the Java/C# natives, the musl and manylinux Elixir NIFs, and the musl and manylinux Node addons) -- now takes XBERG_BUILD_ID as a required build argument and fails the build rather than default to empty; the CI and publish workflows resolve it from the commit being built and pass it through. (GH#1845)
  • (ocr): ocr_near_empty_fallback and ocr_scanned_page_quality_gate now get the per-page byte offsets their routing needs, even without an ocr block. The text pass only tracked page boundaries when force_ocr_pages was set or an ocr block was present, so a caller who set either of the two settings introduced for GH#1752 alone got no per-page offsets at all: the per-page quality gate silently collapsed to one whole-document verdict, and a single flagged page escalated to whole-document OCR instead of routing just that page. Boundary tracking now also turns on for either setting. (GH#1752)
  • (pdf): per-page hierarchy blocks report the measured font size instead of a hardcoded 12pt. PageContent.hierarchy.blocks[].font_size was set to a literal 12.0 for every heading and paragraph block, even though the PDF structure pipeline already computes each paragraph's real dominant font size -- it was read into a debug log and then discarded. Each block now reports the font size actually measured for its element, falling back to 12pt only when none was recorded. (GH#1844)
  • (pdf): a JPEG 2000 CMYK image that names no /ColorSpace now picks up the document's /OutputIntents profile. ISO 32000-1 Table 89 lets a /JPXDecode image omit /ColorSpace, and such an image is given a placeholder /DeviceRGB until the codestream reveals its real component count. The ICC decision was made from that placeholder and never redone, so a four-component codestream was corrected to /DeviceCMYK but reached extraction and rendering with no profile, and its colour was converted through the §10.3.5 additive clamp instead of the document's press characterisation. The same image declaring /ColorSpace /DeviceCMYK was always handled correctly. (GH#1839)
  • (pdf): an image handle for an /ImageMask with no /ColorSpace now reports DeviceGray, matching the decoded image. /ColorSpace is forbidden on an image mask (§8.9.6.2), and extraction has always defaulted such an image to /DeviceGray before decoding it. The cheap, pre-decode image-handle enumeration used a plain /DeviceRGB default for any image with no /ColorSpace, so a mask's handle disagreed with the image decode() later produced from the same XObject. Both the XObject and inline-image handle builders now default a missing /ColorSpace the same way extraction does. (GH#1825)
  • (ocr): numeric_repair now reaches a scanned PDF page's structured output, not just its flat text. On the PDF mixed-OCR route the repair was applied to each page's flat OCR string, which is what a Plain-format extraction renders -- but every other renderer rebuilds each OCR'd page from its structured per-page document, and every table on every renderer comes from there, so Markdown, HTML and JSON output and all reconstructed table cells kept the unrepaired numbers. The per-page documents' element text, table cells, header columns and table markdown are now repaired too, under the same opt-in flag and the same markup-renderer exclusion as before. An element carrying a partial-range bold or italic annotation is left unrepaired, because the repair changes byte offsets the annotation is anchored to. (GH#1840)
  • (ocr): the Tesseract language-availability probe now honours a configured tessdata_path. supports_language and supported_languages always probed the no-override tessdata search chain, even when the caller's OcrConfig.tessdata_path names a different directory -- so a backend reported as unable to handle a language a real OCR job, which does honour that override, loads without trouble. The probe now takes the same override a job passes, and its language memo is keyed by it so two configs with different overrides do not serve each other's answer. supports_language/supported_languages and list_ocr_backend_capabilities/ocr_backend_supports_language keep their previous no-override behaviour; use the new _for forms when a configured tessdata_path should be honoured. (GH#1857)
  • (pdf): the page-count and page-render entry points reject a wrong PDF password instead of ignoring it. xberg_native_pdf's authenticate reports a rejected password as Ok(false) rather than as an error, and this path mapped only the error and discarded the bool, so every wrong password authenticated silently -- pdf_page_count and render_pdf_page_to_png would read an encrypted document with any password at all, including from a binding that mangled the argument in transit. The document extraction path always read the bool and is unaffected. Supplying no password is still not an error: a PDF's page tree is structure rather than encrypted data, so the count reads without authenticating. (GH#1879)
  • (api): an optional reference the JSON contract omits is described in the OpenAPI document as a direct $ref rather than oneOf[$ref, null]. utoipa types every Option<T> as a union with {"type": "null"}, but 41 of the document's 45 optional references carry skip_serializing_if = "Option::is_none", so they are absent from the payload rather than null; the union also pushed each field's description onto a branch, where an OpenAPI generator does not read it. Those 41 properties are now direct references carrying their own description, and they stay out of required. The four that genuinely serialise null -- ElementMetadata.coordinates, DocumentRevision.anchor, ImageMetadataType.dimensions and ImagePreprocessingMetadata.new_dimensions -- keep the union, and inner primitive nullability such as ExtractionConfidence.ocr_aggregate is unchanged. Wire payloads do not change, and no generated binding's types change, because binding optionality is derived from the Rust Option<T> and not from the schema annotation. A client generated from the published spec will see these fields lose their nullable wrapper while remaining optional. (GH#1841)
  • (ocr): a table region whose OCR text was all filtered out no longer publishes an empty table. Table-structure recognition decided between emitting a table and falling back to the region's raw text on the reconstructed cell grid's geometry alone. When the overlap filter that assigns OCR elements to a region discarded every one of them -- its test is against the whole table bounding box, so a region whose word boxes and rasterised crop disagree about coordinates loses all of them -- a geometrically valid grid still produced a fully shaped table of empty strings whose markdown is non-empty, because a header separator alone is not empty. The caller published that table, and because the outcome counted as recognised, the text-preserving fallback added for GH#1622 never ran, so the region's words were absent from the table and from the unrecognised-text output. Such a region now emits no table, leaving its words in the paragraph stream. (GH#1813)
  • (pdf): a JPEG 2000 image with an alpha channel extracts in its own colour space instead of being mis-typed or dropped. The decoder took the codestream's raw component count as the colour-component count, and an alpha channel is a component like any other: an RGBA codestream reported four components and was mapped to DeviceCMYK, and a greyscale-with-alpha codestream reported two and was rejected outright, so the image never reached the page. The codestream's own alpha flag now decides, and the alpha channel is dropped rather than counted -- carrying it through as a soft mask is a separate change. Covers both the JP2 container and the bare codestream form. (GH#1850)
  • (pdf): a multi-page forced-OCR extraction no longer loses every render warning. The engine's render-warning capture buffer is per-thread, and the force_ocr and force_ocr_pages routes render each batch across a rayon pool. A batch of one page does not split, so it rendered on the calling thread and its warnings were collected; a batch of two or more was handed to pool workers whose buffers nothing drained. The batch size is the resolved thread count, so on any multi-core host every document of two or more pages silently dropped missing-glyph-ink, blank-image and fallback-font notices, while the same document extracted with one thread kept them. Each page's warnings are now drained on the thread that rendered it and merged back for the caller, matching what the layout-detection route already did, and in page order, so the same input reports the same warnings on every run. Capture is opt-in and unchanged: a caller that never installs the render diagnostics is unaffected. (GH#1847, GH#1851)
  • (ocr): an invalid candle backend_options or paddle_ocr_config now fails the extraction before any page runs. These settings were parsed by the backend on each page, so on the automatic OCR route a rejected value became a per-page failure, and the extraction returned the native text with only a warning. Configuration validation now runs each backend's own check, for the configured backend and for every pipeline stage, so the caller gets a validation error like any other invalid setting. This is a behaviour change: the error now comes even for a document that needs no OCR. A page that fails at run time, for example on a security limit, still returns the native text with a warning. Options for a custom plugin backend are still checked only on each page. (GH#1829)
  • (ocr): values in a shaded table row no longer glue into one cell across underscore marks. Tesseract reads the edge of a shaded row as runs of underscores and fuses them onto the values beside it, or into one word that spans two values. The fused word's box then covered the gap between two columns, so two values landed in one cell and the cell kept the underscores. Before table reconstruction, each word is now cut at a leading or trailing underscore run and at an interior run of two or more, and each piece keeps the part of the box its characters cover. A single underscore inside a word stays. The OCR text outside tables is unchanged. (GH#1833)
  • (ocr): values in a shaded table row no longer glue into one cell across a stray mark. Tesseract also reads the edge of a shaded row as a word of its own, a tall = or a thin dash, in the gap between two values, and the cell merge then joined both values into one cell. Before table reconstruction, a word in a Tesseract table region is now dropped when its height is under a quarter or over 1.8 times the median height of its row and its confidence is under 35. Both tests must hold: a value read at a low confidence at normal height stays, and so does a thin dash read at a moderate confidence, such as the dash between the two dates of a statement period. The OCR text outside tables is unchanged. (GH#1858)
  • (ocr): a word of a multi-word row label no longer starts a table row of its own when shading stretches its box. Shading can stretch the box Tesseract reports for one word over the next row, so the centre of that box passed the row threshold and the word left the row of its label. Before table reconstruction, the words that Tesseract reads on one text line now take the vertical box of the line's word of typical height, so they share one row. The reported table bounding box uses the same boxes. The native PDF table path does not change. (GH#1834)
  • (ocr): a scanned table is no longer discarded because a misread row minted a phantom column. Table post-processing already folded away a column whose data cells were entirely empty, but a column that was merely almost empty had no such path and instead reached the column-density gate, which rejects the whole table. On a scanned page that is routine: a shaded row read as a few stray glyphs lands at x-positions that mint a header-less column, and on the reported page three stray cells out of twenty-two rows were enough to discard an otherwise well-formed 23-row grid, at every segmentation mode and with or without shaded-row normalisation. Such a column is now folded into its neighbour, carrying its cells' text rather than dropping it. This is a rescue path only: it is reached solely for a column the density gate is about to reject on, and a table containing one returned nothing at all before. A column with its own header is still treated as a legitimately sparse column and left alone. (GH#1797)
  • (ocr): the embedded-image OCR retry stops claiming the rasterizer failed, and stops overwriting better text. The retry fires both when a page's raster came back blank and when a perfectly-drawn page simply yielded no OCR text, but its warning reported every case as image XObjects "the PDF rasterizer could not draw" -- sending a reader looking for a drawing bug that had not happened. Two of the repository's own tests encoded that wording against fixtures whose pages rasterize correctly. The trigger now reports the raster verdict separately and the warning names the cause that was actually measured. Second and more serious: the retry's text replaced the page's own whenever it was merely non-empty, and the retry fires for a page carrying up to 200 characters of real text, so one character of noise from an embedded thumbnail could silently overwrite a terse but correct transcription. The retry's output is now adopted only when it reads more than the page attempt did. The last-resort pipeline route, which runs after every OCR stage has already failed and has no page raster to inspect, keeps the previous wording rather than infer a signal it never measured. (GH#1826)
  • (ocr): a two-word column header split across two columns no longer shifts every value in a scanned table. When the gap inside a header like Year 1 reaches the cell-merge threshold, OCR reports the two words as separate cell tokens and each mints its own column; the right-hand one holds a header fragment and no data. Post-processing already folded such a column away, but only while its data cells were entirely empty -- and a single stray glyph from a shaded row defeats that, so the phantom column survived and every value in the table sat one or more places left of the column it belonged to, while the OCR itself had read it correctly. A column whose header is non-empty and which carries text in at most one row in ten is now folded into its neighbour, header and cells together, so Year and 1 rejoin as one column. On the reported page this moves an absolute 18 of 138 known values into the correct cell to 98, and 36 to 102 under sparse-text segmentation; the 36 that remain are rows whose label OCR mangled, not values in the wrong place. The bar is deliberately far below a legitimately sparse labelled column -- the bank-statement DEPOSIT column of GH#1649 carries two rows in nine -- but unlike the header-less case this does restructure a table that was accepted before, since the column-density gate exempts a column with its own label. (GH#1832)
  • (ocr): PDF table-structure recognition no longer runs on the async runtime's worker threads. Checking a TATR model out of the shared pool can block on a std::sync::Condvar -- for pool exhaustion, or behind a competing model load -- and on a miss it performs model-file I/O and ONNX Runtime session construction; the recognition call itself is synchronous CPU work with no await points. Both ran directly on a tokio worker thread, which parks that thread rather than yielding the task, so on a small worker pool several concurrent layout-enabled extractions could occupy every worker in native blocking calls with nothing left to poll the task that would release a model back to the pool. Both now run through spawn_blocking, the same boundary the RT-DETR layout engine has always used. Nothing is serialized and no pool capacity changed. Reported as an indefinite stall on two concurrent layout-enabled extractions of a scanned PDF (GH#1812), which remains open until the reporter's reproduction is re-run against a build carrying this change.
  • (ocr): numeric_repair no longer rewrites the coordinates in hOCR and TSV output. TesseractConfig::output_format selects which renderer the backend returns, and under "hocr" or "tsv" the OCR text is the markup -- so the repair's thousands-separator rule re-punctuated the bare integers inside every bbox and every TSV coordinate column, turning bbox 1234 567 into bbox 1,234 567. At the default 300 dpi a Letter page is 2550x3300 px, so four-digit coordinates are the norm rather than an edge case, and a downstream hOCR or TSV parser then read a truncated coordinate or failed outright. Both the standalone-image and the PDF route now skip the repair for a markup renderer; the recognized text in elements and table cells is repaired as before. The three shapes the repair rewrites that are not amounts are now documented on the field rather than left to be discovered: a bare 4-9 digit integer is grouped whatever it denotes (Year 2019 becomes Year 2,019, and likewise a ZIP code, page number or part number, with only a literal FY prefix exempt), a US-convention decimal with exactly three fraction digits becomes an amount (0.125 becomes 0,125), and a lone digit before an already-grouped number is joined (5 12,000 becomes 512,000), which is indistinguishable from two adjacent table cells. The option remains off by default. (GH#1836)
  • (ocr): normalize_shaded_rows says so when it cannot take effect. The per-band step operates on the grayscale conversion, but with binarization_method set to "none" or "off" and contrast_enhance on, the non-binarized contrast path re-reads the original image and discards that result -- so the option did nothing at all, silently. That combination now logs a WARN naming the reason, and the field documents the limitation. The processing itself is unchanged on every path. (GH#1837)
  • (pdf): a two-column page whose table spans the full width no longer reads column by column. The corridor search that rescues a split landing inside a table is scoped to the band around that table, so on a page whose table spans the full width the only corridors in scope are the gaps between the table's own columns -- and the search moved a split that was already the page's gutter into one of them. Every prose line below the table then straddled the new split, and the two columns came out interleaved line by line. A split is now kept when a real population of lines outside the band agrees it is already the gutter, measured as both a count and a rate: on a page where the band is the whole page there is nothing outside it to consult, and a disagreement count of zero there means nothing was examined rather than that everything agreed. (GH#1801)
  • (pdf): a full-width figure legend written as several font runs no longer tears at a two-column XY-cut split. The column cut assigned each font run to a side by its own left edge, so a legend line broken into fragments at its own run boundaries -- a bold label, a kerned superscript, a Table 2 prefix -- had its later runs cut off and emitted after the opposite column's body instead of staying with the rest of the line. A contiguous run of full-width lines is now peeled off as its own band before the column cut runs at all, and any full-width line that still reaches the cut is assigned whole, to whichever side holds the majority of its inked width, rather than split at the cut point. An ordinary two-column page's own full-width furniture -- a title, an abstract header, a running header, a footer -- is unaffected: peeling requires at least two such lines in a row, not one. (GH#1808)
  • (pdf): a sparse table beside a prose column no longer redirects a misplaced column split into the table's own cell gap. When a split lands inside a column, the rescue search that moves it to the nearest empty corridor picked the WIDEST corridor with no further check, and a sparse table's own cell gap is routinely wider than the page's real gutter (a row that fills two of five columns opens a gap between them wider than any gutter). The redirect already rejected exactly this shape -- a hanging-label indent, or a corridor whose flanks do not both read as columns -- but only on its second, narrower search; the first, wider search applied neither check. Both searches now share the same qualification, so the true gutter is chosen even when a table's own gap is the wider candidate. (GH#1800)
  • (pdf): a two-column page's last line before a short band no longer welds to the next column's first line. The paragraph grouper decided whether to start a new element from font, weight, role, list markers and vertical spacing alone -- nothing tested whether two candidate lines stood on opposite sides of the page's own column gutter. A left column's closing line and the right column's opening line, whose boxes happened to overlap a couple of points in the vertical direction, were merged into one full-width element spanning both columns. Two lines with no horizontal overlap at all, on either side of the page's single column corridor, now always start a new element; a full-width line (a title, a caption, a footer) straddles the corridor by construction and is unaffected. Recognising such a page as two columns in the first place no longer depends on each line being one span: a producer that writes one Tm/TJ per word used to contribute no prose-line evidence at all, because that test was applied per span rather than per visual line. (GH#1806)
  • (pdf): a two-column page whose second column is short no longer welds the two columns' headings together and reads row by row. Two independent gates refused the page, and fixing either alone left the other defect in place. The column-gutter detector's three sub-detectors all measure MASS -- what share of the page's spans cluster at each left edge, central density, a whole-page classifier verdict -- and a right column that fills only its top third presents none of it, so the detector returned nothing even though the corridor between the columns is empty from the top of the page to the bottom. The XY-cut's heading-run pre-pass is inert without a gutter, so it folded the right column's bold heading into the left column's chapter heading and emitted one list item spanning both columns. A fourth detector now asks only what a gutter is: the page's sole empty x-interval, with a real population of spans on each flank, and narrow enough relative to the content it splits that a label/value form's much wider corridor does not qualify. It declines outright on a page carrying two or more multi-column table rows, so a grid's own recurring cell gaps -- which are empty x-intervals too -- are never read as a column corridor. Separately, the reading-order selector required the two sides' prose-line counts to stand within a ratio of each other, and refused this page at 10 lines against 73. A short second column is an ordinary layout -- a sidebar, a declaration, a signature block beside running text -- and how much text a side holds says nothing about whether the two sides are columns; the vertical-overlap test immediately after it already establishes that they stand beside each other, so the ratio is removed rather than loosened. (GH#1809)
  • (ocr): a scanned PDF page is OCR'd the way the same page is OCR'd as an image. Under force_ocr and force_ocr_pages, a page that is one full-page image rendered at 150 dpi and ran Tesseract in automatic layout mode, while the same page as a PNG kept its own resolution and ran in sparse-text mode. On a scanned table the difference was table values: the automatic mode dropped values the sparse mode reads, and rendering a 196 dpi scan at 150 discarded detail that the OCR preprocessor then interpolated back. Such a page now renders at its image's own density, kept between 150 and 300 dpi, and uses the sparse-text segmentation mode when the caller sets no psm. A scan counts as such when its raster covers at least 80 percent of the page, or at least a quarter of it while the page's own text layer is no more than a stamp of 400 glyphs, which is how a scanned sheet painted inset with margins looks. Every other page keeps the 150 dpi default and the automatic mode, a configured images.target_dpi still decides the render resolution, and an explicit tesseract_config.psm is never overridden. tesseract_config.preprocessing.target_dpi now says that it resamples the OCR image and does not change the render. (GH#1786)
  • (ocr): a scanned page gets the same OCR render density and segmentation mode with layout detection on as with it off. GH#1786 made a scanned PDF page render at its own raster density and use the sparse-text segmentation mode, but only fixed the routes that render pages themselves. With layout detection on, OCR runs against the layout pass's own rasters, which still rendered at the 150 dpi default and left the page on the automatic segmentation mode. Both routes now share the same scan-density render decision and the same whole-page-raster check, so a scanned page's OCR defaults no longer depend on whether --layout is set. (GH#1828)
  • (pdf): unmapped Type 3 text no longer hides behind a short readable header. Extraction retains one visible placeholder per painted glyph whose font has no Unicode mapping, while blank d0/d1 CharProcs become whitespace. The fabricated-text gate can now route a page dominated by those glyphs to OCR without counting blank spacing procedures. A synthetic two-header page has 130 painted fallback glyphs and 126 readable non-whitespace characters and routes to OCR; a six-header control with 378 readable characters remains native. (GH#1782)
  • (pdf): a table drawn with horizontal rules only is no longer lost when the same page has a fully gridded table. The default table detector returned as soon as its grid pass found a table, so the horizontal-rule pass never ran on that page and the later bordered and heuristic passes skip any page that already has a table; the rules-only table came back as paragraphs. The horizontal-rule pass now also runs on the rules no found table covers. Two more cases in that pass are fixed along the way: a booktabs header row alone between the top rule and the midrule joins the body instead of being dropped, and a table ruled under every row forms one table instead of one-row fragments. Across the 379 test_documents PDFs only issue-466-example changes (its "No vertical lines" table is recovered). (GH#1799)
  • (pdf): a heading that spans several shaded table columns is no longer dropped. Three things cost grouped headings: repeated inset background edges inside an already-ruled band were treated as real dividers, a cell's right edge was chosen once with no retry when a divider ended at the header row, and a heading spanning two narrower body cells matched neither of them because grouping required an exact full-edge match. Inset backgrounds are now ignored inside a ruled band, the column search retries wider candidates, and a spanning heading is attached to an established body grid by a separate pass over already-formed groups -- the exact-edge rule the sweep-line prune depends on is unchanged. Header-only fragments stay separate so they cannot suppress fallback extraction of an unruled body. (GH#1803)
  • (pdf): an arrow-bullet list is emitted as a list instead of being folded into the paragraph above it. A producer that marks list items with ➢ (U+27A2) had those items run into the preceding paragraph, and they appeared as neither list items in Markdown nor ListItem nodes in the document structure, while the same document's • lists were handled correctly. Bullet-glyph membership was enumerated independently at thirteen sites across list detection, list normalisation, paragraph splitting, markdown rendering and the two "this is a list, not a table" guards, and those sets had drifted apart -- ➤ (U+27A4) appeared in exactly one of them and ➢ in none. The list-detection and normalisation sites now share one glyph set -- the union of what they already accepted between them -- so a bullet added once is recognised everywhere rather than at whichever sites a fix happened to touch. Two consequences worth stating: the heading/body classifier previously accepted four of these glyphs and now accepts all eleven, which can only make it harder to read a bullet-prefixed line as a heading; and the guard that stops a list being rebuilt as a table accepts the same set, which can only prevent a table, never create one. The sets that are deliberately narrower are left narrow and now say why, and the equivalent guard in the native-pdf crate keeps its own list because it cannot import this one. (GH#1790)
  • (pdf): a valid multi-page PDF no longer logs a warning for every page-tree branch the walk passes. Finding page N by walking the page tree used the same error variant for "not found under this branch, keep looking" as for a genuine page-tree fault; a fix for a deep-tree abort turned that ordinary signal into a WARN log, so a flat n-page tree logged up to n(n-1)/2 warnings extracting every page. The walk now returns None for an ordinary miss and reserves Err for a real fault (a load failure, a cycle, the depth limit, or a malformed node); a genuinely broken branch still warns. Extraction output is unchanged. (GH#1798)
  • (pdf): a font whose /Encoding is a /Differences dictionary no longer logs "dictionary used where stream expected" on every load. Two sites sped up CID-CMap and writing-mode detection by calling the stream decoder on the raw /Encoding object without checking it was a stream first; a /Differences dictionary is a completely normal encoding and was never a stream to begin with. The decoder now runs only when /Encoding actually is a stream. (GH#1795)
  • (pdf): scan detection now scores a page's native text by how much of the page it covers, not by whether it is present or how many glyphs it has. Three related gaps in scanned_confidence are fixed by the same change: a full-page scan carrying a header/footer scored identically to a decorative background under a real page of text (both 0.50, GH#1779); a page whose image covered less than 80% of it scored 0.0 regardless of any other evidence, so a real-world corpus's 112 image-only pages (22-76% coverage, 41-92 native glyphs each) were unreachable at any threshold (GH#1793); and a full-page raster with any visible glyphs at all was capped at 0.65, under the 0.70 default, even with a bilevel codec and a scanner producer (GH#1752). A header, footer, or stamp covers on the order of 0.1-0.5% of a page's area regardless of glyph count; a real page of text covers upward of 9%. Text at or under 2% of the page area now counts as "no substantive visible text" for scoring purposes, the same as an absent or invisible text layer, and the image-coverage floor below which scoring is skipped entirely is lowered from 80% to 15% so a partial-coverage scan is no longer zeroed before its text layer is even inspected. A page whose native text does cover a real share of the page -- a decorative background under body text, or body text next to a figure -- is unaffected. (GH#1779, GH#1793, GH#1752)
  • (pdf): numeric footnotes stay separate in CLI and binding output. A small footnote after a number no longer turns comma 3 plus note 5 into comma 35. The separator requires a matching, smaller note below the reference; numeric scripts without that evidence keep their existing joins. (GH#1771)
  • (ocr): an opt-in preprocessing step recovers table rows on a shaded/filled background. Whole-page Otsu binarization applies one threshold to the entire page, so a subtotal or total row rendered on a light or dark fill is thresholded away in full. ImagePreprocessingConfig::normalize_shaded_rows (default false) finds each shaded horizontal band, determines its own polarity, and stretches its contrast to plain dark-on-white before binarization runs, leaving the rest of the page untouched. It defaults to off rather than on: the same per-band step that recovers light- and dark-fill rows measurably regresses a mid-grey fill it does not model, so enabling it is a per-document decision, not a new default. (GH#1785)
  • (ocr): an opt-in repair for numeric OCR tokens that are clearly numeric but mis-punctuated. Tesseract reliably drops a thousands separator (1172 for 1,172), misreads a comma as a decimal point (7.812 for 7,812), or splits a number across a rendering gap (2 2,411 for 22,411) in financial tables. OcrConfig::numeric_repair (default false) fixes these three narrow shapes after OCR runs. It defaults to off: the repair assumes the US/UK number convention (comma groups, period decimals) and has no table-column context available to distinguish that from 1.234,56-style European formatting, so it can misjudge a genuine decimal on a page that uses the other convention. Enable it only for documents known to use the US/UK convention. (GH#1789)
  • (ocr): ocr-wasm-only builds compile with warnings denied. The automatic-backend lookup now
    compiles only when a PDF or native embedded-image OCR consumer is enabled. (GH#1830)
  • (pdf): text in a table cell that spans several rows or columns is no longer dropped. A detected cell's occupancy was recorded only for the grid slot it starts in, so a word whose centre fell in any other part of the same spanning cell was discarded with no warning; intermediate grid boundaries introduced by neighbouring cells made parts of a cell appear empty. Each span is still assigned exactly once, and text outside any detected cell stays excluded. The spanning cell is still emitted as one row per row-band it covers rather than as a single merged cell -- only the text loss is fixed, not the reported geometry. (GH#1802)
  • (pdf): a document whose cross-reference stream is truncated no longer returns mostly empty pages. When the final /XRef stream fails to decode, the parser rebuilds the table by scanning the file for literal N G obj headers. That scan cannot see an object packed inside an /ObjStm container, which in an incrementally updated PDF routinely includes every page's font dictionary. The reference then resolved to null -- legitimate per PDF 32000-1 7.3.10 for a deleted object, and therefore silent -- so pages extracted no text while the document reported success. Two reported files returned 86 of 88 and 22 of 24 pages empty; both now recover every page. The object-stream sweep that already existed for the analogous mis-flagged-free case now also runs once when the xref came from reconstruction. A reference still unresolvable after that sweep raises an XrefRecovery warning, which is deliberately not raised for an ordinary null resolution. (GH#1774)
  • (pdf): a fixed-pitch producer's word gaps survive extraction. A PDF that places every glyph in its own cell with one Tm and one Tj, and marks a word gap as an empty cell carrying no space glyph, extracted as a single welded token (INVOICENUMBERANDDATE). The operator path that batches such runs adjusts the buffer's accumulated width and consults no spacing heuristic, so the geometric space test never saw the pair. Since a monospace font's space-glyph advance is its character-cell pitch, a gap of at least half a cell beyond the buffer's own advance can only be a skipped cell. The gate applies only to monospace horizontal runs on this path and fired on none of 229 corpus documents, so proportional-font extraction is unchanged. (GH#1770)
  • (pdf): page attributes now inherit past dangling references. A missing indirect /MediaBox, /CropBox,
    /Resources, or /Rotate entry on a page or intermediate page-tree node no longer masks a valid ancestor
    value. Lazy and bulk page walks agree on the nearest valid ancestor and keep sibling inheritance separate.
    (GH#1775)
  • (pdf): a render warning says glyph ink is missing only when a glyph was not painted. Every warning the PDF engine logged during a page render reached processing_warnings as "could not paint one or more glyphs ... the glyph ink is missing", including font-load notes that drop nothing, such as the Type 3 glyph-name fallback. A warning is now worded by what the engine was doing when it logged it: a dropped glyph or a font that could not be found or loaded for rendering keeps the glyph-ink wording, an unrenderable image keeps its image wording, and any other warning reads "Page N rendering logged a warning and continued" with its own cause. (GH#1794)
  • (pdf): a Type 3 font's glyph advances follow its /FontMatrix instead of a hardcoded 1/1000 em. The page renderer and the text rasteriser both advanced the pen by width * size / 1000, which is correct only for the 1000-unit em that Type 1 and TrueType fonts use. A Type 3 font declaring /FontMatrix [0.01 0 0 -0.01 0 0] measures its glyphs in 100 units, so every advance was understated about tenfold and the glyphs stacked on nearly one x position. The scale factor was already parsed and stored as font_matrix_a and simply never read. On the reporting document, 503 such fonts across 165 pages: a rendered page went from 6,408 dark pixels spanning 75 px to 16,705 spanning 610 px, and because the collapsed raster also defeated OCR the pages returned noise and 0 of 20 sampled values. (GH#1780)
  • (ocr): the OCR result cache honours use_cache: false and keys on the recognition model actually used. Two defects made it serve text that did not belong to the request. A top-level use_cache: false bypassed only the whole-extraction cache, leaving TesseractConfig.use_cache -- which defaults to true -- caching underneath it, so a caller asking for a fresh result still got a stored one; the flag is now propagated into the backend_options["use_cache"] channel the backend already reads. And the cache key covered the configured tessdata path string but not the directory it resolved to, so two runs against different TESSDATA_PREFIX models collided on one entry and the second silently received the first model's text; the resolved directory is now part of the key. Anyone who has compared OCR configurations or models in one process should discard measurements taken before this change. (GH#1787)
  • (ocr): a PNG carrying no resolution tag is no longer assumed to be 72 dpi and resampled before recognition. An image with no pHYs chunk was treated as 72 dpi, so a genuine 300 dpi scan was upscaled about 1.24x into interpolated pixels before OCR -- meaning identical pixels produced different text depending only on whether the encoder happened to write a density tag. On a dense reported page that cost 58% of the words, 1,018 down to 431. The resolution is now inferred from the pixel dimensions against standard page sizes at plausible scan resolutions; dimensions matching no standard page size still fall back to the historical 72 dpi assumption, unchanged. (GH#1788)
  • (pdf): a JPEG 2000 image with no /ColorSpace entry is decoded instead of skipped. ISO 32000-1 Table 89 allows the entry to be absent for JPXDecode, because the colour space is carried inside the codestream -- and the decoder already reads it from there, ignoring the dictionary value. The entry was nevertheless required for every non-mask image, so the page handed OCR a blank region while extraction reported success, and neither force_ocr nor force_ocr_pages helped. Only JPXDecode gets the exemption; a missing /ColorSpace is still an error for every other filter, pinned by a FlateDecode control test. (GH#1781)
  • (ocr): a table is no longer discarded because one OCR-split label became its own column. When recognition split a multi-word row label across two cells, the fragment formed a column of its own that was empty in every other row, and the column-sparsity check then rejected the whole table rather than the spurious column -- on one reported set, only one of four inputs survived. Such a column is now folded back into its left neighbour before the check runs, and only when it has no header of its own and every cell in it is word-like, so a genuinely sparse numeric column is still left to the checks that exist for it. A malformed grid is still rejected. (GH#1797)
  • (ocr): force_ocr renders its page batches across the thread budget. The route that recognises every page rendered them one at a time and only parallelised recognition, so raising max_threads did not shorten the render half at all -- 17.5 s against 8.2 s on a 40-page scan at sixteen threads, for byte-identical output. The sibling route that takes an explicit page list was already parallel; both now use the same bounded pool, and remain sequential on wasm32, which has no thread pool. (GH#1796)
  • (pdf): a decoder stops at its output limit instead of allocating past it first. RunLengthDecode and LZWDecode built their entire output and only then compared it against the configured maximum, and the compression-ratio guard could never fire for run-length data at all, since that filter's densest encoding is 64:1 against a 100:1 check. A chained [FlateDecode, RunLengthDecode] stream could therefore hold about 16 GB transiently. Both now check after each decoded unit, bounding the overshoot to one unit. This is transient allocation only -- accepted output could not exceed the cumulative ratio before this change either. (GH#1764)

Don't miss a new xberg release

NewReleases is sending notifications on new releases.