🌀 Release fingerprint — v0.10.2..v0.11.0
8 commits · 1 contributor(s) · +5004/-988 across 35 files · 2026-09-11 → 2026-09-23
🎨 Generated from this release's commit signal — how it's made →
Text extraction now reads rotated text, superscripts and letter-spaced text the way a reader sees them. The command-line tools are folded into one gopdf command. RemoveText now also removes marked-content replacement text, which could still expose a removed string.
One command
go install github.com/razvandimescu/gopdf/cmd/gopdf@latest installs one binary with every capability behind subcommands:
gopdf tables extract a table as aligned text or CSV
gopdf merge combine PDFs and images into one PDF
gopdf pages keep only some pages of a PDF
gopdf watermark stamp an image across every page
cmd/merge, cmd/watermark and cmd/extract_tables are removed. Anyone installing them by path should install cmd/gopdf instead. watermark now takes its input as a positional argument.
Image pages come with library additions:
- Baseline JPEGs are embedded without re-encoding.
- EXIF orientation is honoured.
Image.DPIandImage.DisplaySizeare new.PageBuilder.DrawImageis new.- A merge now writes the same bytes on every run.
Text extraction
- Rotated text is read along its own baseline, whether it runs up the page, down it, upside down or at a slant. It used to come out one glyph per line.
- Superscripts and subscripts join their line:
W/m2K,H2O. Before, they formed a separate line above or below. - Letter-spaced text stays one word. A gap counts as a space only above 0.15 em of the drawn size, so tracked text such as
S s _ 4 0now readsSs_40. - Valid UTF-8 on every page. Codes that ToUnicode and the font's encoding left unmapped are decoded through the encoding the font implies. The WinAnsi and MacRoman tables are complete, and glyph names are decoded by the Adobe Glyph List's rules.
- Inline images no longer end early when their data happens to contain
EI, which used to lose the rest of the page. - Streams with spaces after the
streamkeyword are read. Before, such a page extracted as empty text with no error. - Search locates matches by glyph, using the same recorder as
RemoveText, so its rectangles are exact for rotated text too.RedactTextdraws its boxes from them. Page.OutlineHintis new. It reports when a page's filled paths repeat like glyphs, which means an empty extraction may be text drawn as outlines rather than an empty page.
Removal
When RemoveText removes glyphs inside a marked-content section, it also strips the section's /ActualText, /Alt and /E. MuPDF, Acrobat copy/paste and screen readers read that text in place of the glyphs, so before this fix the removed string was still readable. Other keys, such as /MCID, are kept. Replacement text stored in the structure tree is still not removed. The README table lists what is and is not removed.
Fixes that change existing behaviour
Text(),TextLines,TablesandFindTablereturn an error for a page whose content cannot be read. They used to return empty results.Document.Textnames the page that failed.TextSpan.FontSizeis the em height as drawn on the page, not theTfoperand. The two differ when a matrix scales the font.TextLine.Yis the baseline of the line's largest text, measured where the line starts.BuildLinesno longer reorders the slice it is given.- Table auto-detection no longer reports running text as a table. When column names average under four characters, the candidate is treated as prose.
-headersbypasses detection for tables whose columns really have one- or two-character names.
No exported library symbols were removed.
CI now tests the go.mod floor (Go 1.22) and the latest stable Go.
Full changelog: v0.10.2...v0.11.0
