What's New
Anchor-based auto-detection (Approach 3)
Tables are now auto-discovered without explicit headers. The library finds the row with the most short, non-numeric, well-spaced spans across all pages, then uses zone-based mapping to handle PDFs where header X positions don't align with data positions (e.g. BCR bank statements).
AutoTune
FindTableAcrossPages with AutoTune: true tries a grid of MergeGap/MaxRowGap values and picks the result with the best log(rows) * fillRate² score. No manual parameter tuning needed.
AnchorColumn
Content-based row merging for formats where Y-gap signals can't distinguish transactions (e.g. BCR statements with uniform 8pt line spacing). Rows where the anchor column is empty merge into the previous row — but only if they have description text, not summary amounts.
Example
// Fully automatic — discovers columns, tunes parameters
tbl := pdf.FindTableAcrossPages(pages, &pdf.TableOpts{AutoTune: true})
// For formats needing content-based merging
tbl := pdf.FindTableAcrossPages(pages, &pdf.TableOpts{
AutoTune: true,
AnchorColumn: "Data operatiunii",
})Full Changelog: v0.5.1...v0.5.2