Added
-
ServerConfiggainedjob_timeout_secs(default 600 seconds, override via
XBERG_JOB_TIMEOUT_SECSorserver.job_timeout_secs) as the configurable fallback timeout
forPOST /extract-asyncjobs whose request does not pin downextraction_timeout_secs. A
per-requestextraction_timeout_secsstill always overrides it, and an explicit
extraction_timeout_secs: nullstill falls back to this server cap rather than running
unbounded. Previously this fallback was a hardcoded 300 seconds, inconsistent with the 600
second default used everywhere else. See the Changed section for the source-compatibility impact. -
A long-running process can now release the embedding and reranker models it no longer
uses.embeddings::evict_modelandreranking::evict_model(and the same functions in
sparse_embeddingsandlate_interaction) drop one model,clear_engine_cachedrops
every model in a cache, andxberg::clear_engine_cachesdrops all of them.
set_engine_cache_limitbounds the number of resident engines in a cache and drops the
least recently used one first. The default stays unbounded, so existing callers see no
change (GH#1626).
Fixed
-
OutputFormat.customin the Python binding always returnedNone, even when the value genuinely
was a custom format. The accessor derived a discriminator from the enum's JSON, which works for
every tagged representation but not for this variant, whose payload is untagged and so carries no
discriminator to find. It now matches on the variant directly and returns the label. -
The Python type stub declared a
type: strattribute on 23 enum classes that have no such
attribute at runtime. The stub emitted it for every data enum, while the runtime only exposes it
for enums carrying an explicit serde tag -- so for externally tagged enums (EntityCategory,
PiiCategory,OutputFormat) a type checker acceptedcategory.type, which raised
AttributeErroron use. Stub-only change; no runtime behaviour moved. -
The Ruby binding's
FormatMetadata.from_hashandDiffLine.from_hashread a_0key that no
longer exists on the wire, so every variant deserialized with anilpayload. All 24 call
sites are corrected: the 21FormatMetadatavariants now build their payload from the flattened
hash, and the 3DiffLinevariants read thetextkey the enum's#[serde(content = "text")]
actually emits (#1594). -
The Python type stub declared format-metadata payloads as
_0(for example
_0: ExcelMetadata), a key present neither on the wire nor on the runtime object -- the runtime
per-variant getters such as.excelwere always correct. The stub now names the real fields, so
type checkers stop reporting valid code as an error. Runtime behaviour is unchanged
(#1594). -
The Elixir
Xberg.FormatMetadatatypespec documented ametadata:payload key that matched
neither the NIF struct nor the serialized wire. It now matches the struct the NIF actually
returns (#1594). -
A Type0 (composite) font's content-stream character codes are now translated to CIDs before glyph
widths and vertical metrics are looked up, instead of being used as if they already were CIDs.
The two coincide only forIdentity-H/Identity-V, which is presumably why this went unnoticed.
A PDF using a non-Identity predefined CMap (UniCNS-UCS2-H,UniJIS-UCS2-H,UniGB-UCS2-H,
UniKS-UCS2-H, or their-V/UTF16counterparts) or an embedded/EncodingCMap stream got a
wrong width for nearly every glyph -- some over-advancing by up to 4x through the/DWfallback,
others under-advancing -- which rendered as stretched or overlapping text and, in extracted text,
could split one sentence into a spurious extra paragraph. CIDs now resolve from an embedded
/EncodingCMap stream's ownbegincidrange/begincidchardata when present (including a
variable-width codespace), else from the font's/CIDSystemInfocharacter collection for the
four Unicode-keyed predefined families above, else viaIdentity-H/Identity-Vas before.
Measured over the 230-document local PDF corpus: 227 byte-identical -- the expected result, since
most PDFs use Identity-H -- and 2 changed, both merging text that a stale glyph position had
fragmented. Legacy multi-byte predefined CMaps this crate carries no code-to-CID table for
(90ms-RKSJ-H,GBK-EUC-H,B5-H, theUTF8family,UniJIS-UCS2-HW-*,UniJISPro-*, and
others) are unchanged -- still wrong, not newly broken -- and now log once per font instead of
failing silently (GH#1631). -
A reconstructed PDF table cell now reads left to right instead of in the order its words happened
to arrive. The reported symptom was a sub/superscript printing after the rest of the cell --
eta_S %came out aseta % S,Q_HE GJasQ GJ HE-- because a script is drawn as its own
content-stream segment a fraction of a point below the line it annotates, so every reading-order
sort upstream placed it after the whole line. The same defect also transposed values between
columns when two columns were merged into one cell: on a balance sheet whose header reads
2017 2016, the row beneath it emitted the 2016 figure first, silently attributing each year's
number to the other year. A cell's words are now grouped into visual lines and ordered left to
right within each line. Grouping first is load-bearing -- ordering by horizontal position alone
interleaves the two halves of a wrapped cell. Each word's own whitespace is also collapsed, so a
segment carrying a trailing space no longer stacks it on the separator. Table cells recovered by
OCR go through the same ordering (GH#1628). -
Image OCR now honours a PNG's embedded
pHYspixel density instead of assuming 72 DPI. A genuine
300-DPI PNG submitted withtarget_dpi = 300was resized anyway, because the extractor decoded,
resized and re-encoded the image -- discarding the density chunk -- before the OCR backend, and
therefore before the existingocr.backend_options["source_dpi"]override, ever saw it. Embedded
density is now resolved at the extractor boundary and at the backend from one shared
implementation, with the explicit override still taking precedence over it. An image carrying no
density metadata still defaults to 72 DPI and still resizes (GH#1630). -
A body paragraph is no longer deleted for repeating text that appears elsewhere on the same page.
The secondstrip_repeating_textpass keyed on lowercased paragraph text with no check that a
table was involved, so a sentence matching an earlier title -- differing only in case, with no
table on the page at all -- was silently removed. The pass now runs only on pages that have a
detected table, removes a paragraph only when that table's own cells carry the same text, and
compares case-sensitively. Measured over 230 PDFs: 44 documents changed, 3472 words recovered and
32 lost, both loss cases inspected and benign (one is a restructure whose total content grew, the
other two mojibake tokens) (GH#1623). -
OCR text is no longer discarded when a scanned page region is detected as a table but its cell
grid cannot be recognised.recognize_single_tablereturned nothing whenever TATR failed,
produced no rows or columns, or the grid failed validation, which threw away every OCR element
that had been assigned to that region. A region that cannot be recognised as a table now falls
back to emitting its text in reading order, and only when that text is not already carried by
one of the page's paragraphs, so nothing is duplicated. Together with the restructuring-heuristic
retention guard below, recognised OCR text is no longer silently lost on the layout path
(GH#1622). -
extraction_confidenceno longer reports a failed structured extraction as fully
schema-valid. The pipeline passedSchemaCompliance::AllValidunconditionally, which is 40% of
the combined score under the default weights, so a run whose LLM call failed -- or that was
built without theliter-llmfeature, or ran on wasm -- scored exactly as high as one that
validated. A requestedstructured_extractionthat leaves nostructured_outputnow scores
AllInvalid. Extractions with nostructured_extractionconfigured are unaffected and keep
their previous score;ConfidenceSignalsis unchanged in shape, so no serialized form moves
(GH#1624). -
detect_mime_type_from_bytesno longer refuses text that is not valid UTF-8. A byte buffer with
no filename or declared type -- a Windows-1252 or ISO-8859-1 CSV export, say -- returned
UnsupportedFormateven though the extractors that would receive it decode legacy encodings
throughencoding_rs. Such content is now reported astext/plain, the same answer the UTF-8
path already gave for the same document, so the two encodings of one file behave alike. Content
holding a NUL byte, or with too few printable bytes to read as prose, is still rejected
(GH#1625). -
The Go binding no longer discards the message of every error the native layer reports. Each
known error code was mapped to a typed sentinel (ErrTimeout,ErrParsing,ErrOcr, and ~20
more) and returned before the message was ever read, so the detail the native layer had
already produced -- observed durations, limits, plugin names, counts -- was dropped for all of
them; only unrecognised codes kept their text. A timeout surfaced as the sentinel's own
placeholder-stripped text,extraction timed out after ms (limit: ms), which reads as a
formatting bug but is the whole message the binding ever had, and left callers unable to tell
which timeout had fired. The message is now read first and returned alongside the sentinel, so
errors.Is(err, xberg.ErrTimeout)still matches whileerr.Error()carries the real
interpolated text. Go was the only binding affected; C#, Java and Zig already read the message
before switching on the code. Regression in 1.1.0, when the typed sentinels were introduced. -
The Python package's public option classes regained
from_json.from xberg import ExtractInputresolves to a generated dataclass that shadows the native class at the same
name, and that dataclass carried none of the native class's methods, soExtractInput.from_json(...)
raisedAttributeErrorwhilexberg._xberg.ExtractInput.from_json(...)worked -- the same
name meaning two different things depending on the import. 134 public classes were affected.
The dataclasses now delegatefrom_jsonto the native class, so both import paths behave the
same. Other native-only methods on those classes (validate,is_empty, the
PaddleOcrConfig.with_*builders) are still absent from the dataclass twins and are tracked
separately. -
An extraction cancelled by
extraction_timeout_secsnow actually stops its per-page PDF OCR
work. The timeout firescancel_token.cancel()at every timeout site, but nothing in the OCR
page fan-out read the token, so pages kept being OCR'd after the caller already had its
Timeouterror -- burning CPU and holding OCR concurrency permits, which degrades later
extractions in a long-lived process (a server, or anything extracting in a loop). The token is
now checked both before spawning a page and inside each spawned task, because the spawn loop
finishes almost immediately while tasks queue on the OCR semaphore long after it. A cancelled
run also reportsCancelledinstead of tripping the all-pages-failed guard and reporting a
wholesale OCR backend failure. -
PDF no longer promotes ordinary body text to a heading. Two gates decide headings
independently and neither tested the line's shape, so any line past the title-length floor
could be promoted. The sentence-boundary check that should have caught this looked for a
literal". "followed by a capital, but paragraph text joins a block's physical lines with a
newline, so every sentence boundary landing at a line end was invisible to the gate while the
renderer joined the same lines with a space and displayed it -- the gate and the output
disagreed about what the text was. Boundaries are now found across any whitespace, and a line
that is mostly bare numerals is treated as a flattened data row rather than a heading. Across
490 documents: 484 unchanged, 5 with fewer headings, 0 with more (GH#1599). -
Hardened the document-global heading/list heuristic's safety check on the scanned-PDF
layout-markdown path (use_layout_for_markdown/ layout detection, force-OCR route). That
heuristic rebuilds paragraphs from bare line geometry with no knowledge of the ML layout
regions the OCR path already classified, and can silently drop a line its own font-clustering
pass treats as furniture or noise; the guard against this only checked that the whole
document still had one non-empty element, so a single surviving word anywhere passed it even
if an entire page's body vanished. The guard is now a per-restructuring canonical-character
retention check against the lossless OCR assembly, and falls back to that lossless assembly
whenever any content would otherwise be lost. Compares characters rather than word tokens: a
restructuring pass legitimately re-wraps text across the line boundaries it reads (measured
case: "list of findings" split across a line came back "list offindings", one dropped space),
and a word-token comparison read that benign re-wrap as content loss and rejected legitimate
heading/list promotion along with it. This closes a real gap in the guard's own logic; it was
not reproduced against a specific "entire page lost" report and should not be read as a
confirmed fix for one (GH#1622). -
PDF table/paragraph assembly (
assemble_page_elements_with_tables) now suppresses a
paragraph whose words a positioned table's own grid fully carries, so a recognized table no
longer also renders its flattened source text as an ordinary paragraph immediately next to
the grid -- observed directly (not inferred) on a scanned-PDF fixture with layout detection
enabled, where a table's status-row text appeared once as prose and once as a correctly
gridded table. Mirrors the GH#1616 precedent from the other direction: suppression requires
the table's own cell/markdown content to actually account for every one of the paragraph's
words (an order-insensitive multiset match, since a reconstructed grid can reassemble the
same words in a different order than the source paragraph), not geometry alone, so a
paragraph carrying text the grid does not represent still survives. Note: this closes the
duplication for paragraph/table pairs that share a coordinate space (the native-PDF table
path). Investigating this also surfaced a separate, unresolved coordinate-space mismatch
between OCR/TATR-recognized table bounding boxes and OCR paragraph bounding boxes on the
force-OCR + layout-detection route specifically, which currently prevents this same guard
from geometrically matching on that route; fixing that is out of scope here and is not yet
done (GH#1622). -
PDF no longer deletes text a table's bounding box covers but its grid leaves out. Suppression
of text a table already renders was decided on geometry alone, and a reconstructed grid need
not span every printed column inside its own bounding box. On a four-column fault-finding grid
reconstructed with two columns, every run in the two omitted columns vanished from the
document — not in a cell, not in any element, nowhere. A covered run is now suppressed only
when the table actually carries its text (GH#1616). -
PDF no longer cuts a numbered heading that wraps onto a second line. The wrap exemption
compared the two lines' right edges, and a wrap's last line is short by definition, so it could
never fire: the heading kept only its first line and the rest of its title was emitted as body
text. A heading's own continuation is now recognised by its left edge, which is the title's
hanging indent rather than the margin body text returns to. Regression in 1.1.5 (GH#1615). -
An extraction that never requested OCR no longer fails when no OCR backend is registered.
ocr-pipelinecan be enabled without any backend —ocrimpliesocr-pipeline, not the
reverse — and in that build the automatic scanned-page trigger aborted an ordinary PDF
extraction withOCR backend 'tesseract' not registered. Automatic triggers now check
availability and skip with a warning; an explicitforce_ocr,force_ocr_pages,
ocr_inline_imagesor caller-suppliedocrconfig still fails loudly (GH#1610). -
Legacy binary
.pptnow reports which slide each embedded picture belongs to. Pictures were
read from the OLEPicturesstream, which stores blips in save order and names no slide, so
every extracted image carried no page number and every image node was emitted after the last
slide. A slide whose only content is a picture therefore produced nothing at all on its own
number and read as a blank slide, and captions or any other data keyed on an image's page were
filed against the end of the deck. The owning slide is now resolved through the drawing that
references the blip; a picture no live shape references is still extracted, without a slide
(GH#1620). -
Legacy binary
.pptno longer extracts deleted slide revisions or presents slides in the
wrong order. The format is append-only across saves, so editing a deck leaves superseded
copies in the stream; treating everySlidecontainer as a slide produced 190 slides for a
96-slide presentation, numbered by byte order. Live slides and their order now come from the
persist chain (Current User→UserEditAtom→PersistDirectoryAtom) and the document's
slide list, falling back to the previous behaviour if the chain cannot be read in full. Slide
numbers are the page every element and chunk of a deck is cited by, so both defects reached
consumers as wrong page numbers (GH#1614). -
PDF de-hyphenation no longer welds a compound whose own hyphen falls on a line break. Two
sites decide whether a trailing hyphen survives; only one consulted the lexical evidence, so
long-term,cost-effectiveandantigen-presentingcame out aslongterm,costeffective
andantigenpresenting— tokens that do not exist, and so unreachable by any lexical search.
The assembly site now asks the same question the paragraph site already asked, weighing both
the static compound list and the witnesses collected from the document itself. A hyphen the
wrap genuinely inserted is still removed (GH#1613). -
Legacy binary
.pptno longer loses slide titles. PowerPoint keeps a slide's text in two
places, and the extractor read only one: titles held in the document-level outline
collection (SlideListWithText) landed in the loose-text bucket, which is discarded whenever
any slide exists, so they were absent from the output entirely. Outline text is now attributed
to its slide by persist order and merged in, skipping any line the slide's own drawing already
carries so a title drawn on the canvas is not duplicated (GH#1612). -
The documented install versions for Java, Kotlin Android, Swift, Zig and the spring-ai
integration no longer lag the release. These snippets sit outsidetask version:sync, which
covers the generated API-reference badges but not hand-authored install directives, so they
had been telling users to install 1.1.3 (GH#1593 covers the same class of staleness in
test_apps, which is still open). -
OcrConfigno longer rejects valid Tesseract language codes such asfao(Faroese) with
Invalid language code 'fao'. Use ISO 639-1 or ISO 639-3 codes.. Config validation checked
the language against a general-purpose allowlist that was missing 66 codes Tesseract actually
supports, while a separate, Tesseract-specific list already carried them; the two lists had
never been reconciled. Config validation itself only started running for configs loaded from
files, JSON overrides, or set programmatically in 1.1.0 (previously it ran only in tests), which
is when this allowlist gap first became user-visible. Both validators now read from one shared
list of Tesseract-supported codes, so this class of divergence cannot recur (GH#1621).
Changed
-
Breaking (Java binding): enum constants now follow Java's own convention and are
SCREAMING_SNAKE_CASEinstead of carrying Rust's PascalCase verbatim --LinkStyle.Inline
becomesLinkStyle.INLINE, across roughly 75 generated enums. The JSON wire value is
unchanged; only the Java identifier moves, so serialized documents and stored payloads are
unaffected. Update references to the constants themselves; generated default values
(ChunkType.Unknown,OutputFormat.Plain) moved with the declarations. -
Breaking (Ruby binding): an externally tagged enum variant carrying a single payload
(EntityCategory::Custom,PiiCategory::Custom,OutputFormat::Custom) now serializes the way
the Rust core always did --{"custom" => "my-label"}-- instead of wrapping the payload in an
extra object keyed by a synthesized positional name,{"custom" => {"_0" => "my-label"}}. Code
readinghash[:custom][:_0]should readhash[:custom]. The previous shape matched no other
binding and no core output. -
Breaking (PHP binding):
EntityCategory,PiiCategoryandOutputFormatchange from
constants-only classes to classes with static factories, because the old shape could not carry a
payload at all: a caller-supplied label was silently discarded in both directions, so
Custom($label)always round-tripped as an empty string. UseEntityCategory::custom($label)
andEntityCategory::person()in place of the oldXberg\EntityCategory::PERSONconstants; the
label is readable from the readonly$customproperty. -
Breaking (Node binding): the JSON surface is camelCase throughout, nested types included, and
FormatMetadatais a flat discriminated union keyed byformatTypewhose variant payload fields
sit directly on the object ({ formatType: "excel", sheetCount: 2 }). The binding previously
exposed two parallel shapes for one Rust type -- an idiomatic camelCase interface beside a
snake_case structural twin -- and only the latter was reachable from a result. -
Breaking (Node, Swift bindings):
FormatMetadatanow serializes flat in every binding --
{"format_type": "pdf", "page_count": 12, ...}-- matching what the core Rust enum has always
serialized (#[serde(tag = "format_type")]), what the OpenAPI discriminator describes, and what
the REST API serves. Two bindings disagreed with that wire and have been corrected:-
Node nested the payload one level down under a property named for the variant, so
doc.metadata.format.pdf.pageCountbecomesdoc.metadata.format.page_count. Note the field
names are snake_case, unlike the camelCase Node uses elsewhere: the variants carry
mutually incompatible field types (headersisstring[]for text and an object array for
HTML), so no single Node class can describe them and the value is passed through as serde
emits it.formatis typed as a discriminated union inindex.d.ts, so narrowing on
format_typestill gives a fully typed payload. This also resolves the Node and WebAssembly
bindings disagreeing with each other -- WASM was already passing serde's shape through, so the
two now emit an identicalformatobject for the same document. -
Swift's
FormatMetadatawas atypealiasto an opaque bridge class carrying no payload,
so the JSON inMetadata.formatcould not be decoded into anything useful. It is now a real
Codable/Sendableenum with one case per format, making the payload reachable:if let json = doc.metadata?.format, case .excel(let meta) = try formatMetadataFromJson(json) { print(meta.sheetCount) }
Metadata.formatstill hands back the serialized JSONString; what changed is that
formatMetadataFromJsonnow yields a pattern-matchable enum carrying the payload instead of
an opaque handle. The wire was already correct here -- the Swift type system was the part
that was missing.
Python, Go, Java, C#, Kotlin, PHP and Ruby are unaffected on the wire: they either already
emitted the flat shape or expose native per-variant accessors over it
(#1594). -
-
Breaking (Rust source, Java):
ServerConfigaddsjob_timeout_secs. Exhaustive Rust struct
literals must set the field or use..ServerConfig::default(), and the Java record's canonical
constructor gains a sixth component, sonew ServerConfig(host, port, corsOrigins, maxRequestBodyBytes, maxMultipartFieldBytes)no longer compiles -- useServerConfig.builder(),
which is unaffected. Every other binding is source-compatible: the field is last and defaulted in
the Python dataclass (= 600), the Kotlin data class (= 600L) and C# ({ get; init; } = 600);
a defaulted keyword in Ruby and PHP; an optional pointer withomitemptyin Go; and an additive
xberg_server_config_job_timeout_secsgetter in the C FFI (gated onapi-types). Deserializing
callers are unaffected everywhere -- the field carries#[serde(default)]. -
Public binding-facing structs in this crate are deliberately not
#[non_exhaustive]: alef
generatesimpl From<Mirror> for xberg::Twith a struct literal in roughly ten binding crates,
and#[non_exhaustive]forbids that cross-crate (E0639) -- including the..Default::default()
spread.Defaultplus#[serde(default)]is the forward-compatibility mechanism instead, and a
field addition is recorded here as a labelled source break rather than prevented by the type
system.#[non_exhaustive]is reserved for types excluded from binding generation. -
Retroactive note for 1.1.4:
Metadata#formatin the Ruby binding changed shape and no
changelog entry recorded it at the time. The format-specific payload had been nested under a
_0key (format.fetch(:_0).fetch(:title)); since 1.1.4 the payload's fields sit directly
alongside theformat_typetag (format.fetch(:title)). Ruby callers written against the
older shape raiseKeyErroron_0. The binding has emitted the flat shape since 1.1.4; the
generated Ruby e2e specs were still asserting the nested one, which is why this went unnoticed
for two releases. Only the Ruby binding is affected. Part of GH#1594, which also tracks the
Swift binding still discarding the payload entirely -- that half is not yet fixed.