github razvandimescu/gopdf v0.11.0
v0.11.0 — one CLI, rotated and raised text, replacement text removed

latest releases: v0.11.2, v0.11.1
18 days ago

release fractal

🌀 Release fingerprint — v0.10.2..v0.11.0

8 commits · 1 contributor(s) · +5004/-988 across 35 files · 2026-09-11 → 2026-09-23

🎨 Generated from this release's commit signal — how it's made →

Text extraction now reads rotated text, superscripts and letter-spaced text the way a reader sees them. The command-line tools are folded into one gopdf command. RemoveText now also removes marked-content replacement text, which could still expose a removed string.

One command

go install github.com/razvandimescu/gopdf/cmd/gopdf@latest installs one binary with every capability behind subcommands:

gopdf tables     extract a table as aligned text or CSV
gopdf merge      combine PDFs and images into one PDF
gopdf pages      keep only some pages of a PDF
gopdf watermark  stamp an image across every page

cmd/merge, cmd/watermark and cmd/extract_tables are removed. Anyone installing them by path should install cmd/gopdf instead. watermark now takes its input as a positional argument.

Image pages come with library additions:

  • Baseline JPEGs are embedded without re-encoding.
  • EXIF orientation is honoured.
  • Image.DPI and Image.DisplaySize are new.
  • PageBuilder.DrawImage is new.
  • A merge now writes the same bytes on every run.

Text extraction

  • Rotated text is read along its own baseline, whether it runs up the page, down it, upside down or at a slant. It used to come out one glyph per line.
  • Superscripts and subscripts join their line: W/m2K, H2O. Before, they formed a separate line above or below.
  • Letter-spaced text stays one word. A gap counts as a space only above 0.15 em of the drawn size, so tracked text such as S s _ 4 0 now reads Ss_40.
  • Valid UTF-8 on every page. Codes that ToUnicode and the font's encoding left unmapped are decoded through the encoding the font implies. The WinAnsi and MacRoman tables are complete, and glyph names are decoded by the Adobe Glyph List's rules.
  • Inline images no longer end early when their data happens to contain EI, which used to lose the rest of the page.
  • Streams with spaces after the stream keyword are read. Before, such a page extracted as empty text with no error.
  • Search locates matches by glyph, using the same recorder as RemoveText, so its rectangles are exact for rotated text too. RedactText draws its boxes from them.
  • Page.OutlineHint is new. It reports when a page's filled paths repeat like glyphs, which means an empty extraction may be text drawn as outlines rather than an empty page.

Removal

When RemoveText removes glyphs inside a marked-content section, it also strips the section's /ActualText, /Alt and /E. MuPDF, Acrobat copy/paste and screen readers read that text in place of the glyphs, so before this fix the removed string was still readable. Other keys, such as /MCID, are kept. Replacement text stored in the structure tree is still not removed. The README table lists what is and is not removed.

Fixes that change existing behaviour

  • Text(), TextLines, Tables and FindTable return an error for a page whose content cannot be read. They used to return empty results. Document.Text names the page that failed.
  • TextSpan.FontSize is the em height as drawn on the page, not the Tf operand. The two differ when a matrix scales the font.
  • TextLine.Y is the baseline of the line's largest text, measured where the line starts.
  • BuildLines no longer reorders the slice it is given.
  • Table auto-detection no longer reports running text as a table. When column names average under four characters, the candidate is treated as prose. -headers bypasses detection for tables whose columns really have one- or two-character names.

No exported library symbols were removed.

CI now tests the go.mod floor (Go 1.22) and the latest stable Go.

Full changelog: v0.10.2...v0.11.0

Don't miss a new gopdf release

NewReleases is sending notifications on new releases.