paperless-gpt now connects documents that belong together. A reminder is linked to the invoice it is about, an amendment to its contract, a letter to the case file it cites, automatically, through paperless-ngx's own Document Link fields.
Under the hood, PDF rendering moves from MuPDF to PDFium. That keeps the published binary under permissive licenses, and it removes the last system library from the build. Around both, a set of fixes for silent data loss in tags, custom fields and the history view.
The theme is connection: between documents, and between what paperless-gpt does and what you can see and undo.
Heads-up: PDF rendering now uses PDFium. For almost everyone this changes nothing, but read the Known issues at the bottom.
🔗 Related documents link themselves
Office paperwork comes in chains. A reminder cites an invoice number, a credit note cites the invoice it corrects, a delivery note cites the order, and every letter in an insurance claim or a dispute cites the same file number (Aktenzeichen). Until now, connecting them was manual work.
paperless-gpt now reads those references and fills a paperless-ngx Document Link custom field with the documents they point to. paperless-ngx shows the link on both sides, so the invoice lists its reminder and the reminder lists its invoice.
- Accounts payable: order, delivery note and invoice end up connected.
- Contracts: a contract carries its amendments, renewals and the termination letter.
- Case files: every letter that cites the same file number joins the case, whoever sent it.
- Reminders: checking whether a reminder is already paid becomes one click.
It is built to run unattended. The model only extracts the reference numbers; paperless-gpt resolves each one itself with an exact whole-word match against your archive. References that are too short, or that match more than three documents (your own customer number, say), are skipped, a document is never linked to itself, and anything that doesn't resolve is left unlinked rather than guessed.
Setup: create a Document Link custom field in paperless-ngx, select it under Settings → Custom fields in paperless-gpt, done. No new environment variable.
Thanks @andre161292 for building it, and @LembkeM for describing exactly this in the feature request. (#1097, #1098, closes #1020)
📄 PDF rendering: PDFium replaces MuPDF
MuPDF is licensed under the AGPL-3.0 and was statically linked into the binary, which made every published image an AGPL-bound work although paperless-gpt is MIT. PDF pages are now rendered with PDFium (Apache-2.0, the engine behind Chrome's PDF viewer) through go-pdfium's WebAssembly runtime.
- Licensing: the paperless-gpt binary and images are now built from permissively licensed components only.
- No system PDF library:
mupdf/libmupdf-devare no longer needed to build from source. A C compiler is still needed for SQLite. - Same output: checked side by side against MuPDF on a real archive of 726 documents and 1,626 pages. Page counts and sizes match exactly, rendering speed is the same, and filled-in form fields are rendered, as before. Memory peaks slightly higher while rendering; see Known issues.
- Ready before the first document: with OCR enabled, the renderer is compiled in the background while paperless-gpt starts, on your machine and for your CPU. Without that, the first page after every restart waited for the compilation, about 20 seconds on a small home server; it now renders in under a second.
- Go 1.27.1 for the build.
🛡️ Fixes for silent data loss
- Append mode no longer deletes custom fields. paperless-ngx replaces the whole custom-field list on update, and for auto-tagged documents paperless-gpt was merging against an empty list. So
append, documented as "never overwrites", removed every field you hadn't selected. It now merges into the document's current fields. Thanks @shiaho777, and @schubimice for the report. (#1085, closes #1031) - Tags your API user can't see are no longer removed. With a scoped-down API token, tags owned by other users couldn't be resolved, and an update meant only to swap the trigger tag could strip a document of all of them. They are now preserved untouched. Thanks @shiaho777, and @MSPTW for the report. (#1099 from #1084, closes #1012)
- No more permanent fail tag for "no date found". When the model couldn't determine a created date, the field was skipped locally but still counted as rejected by paperless-ngx, so
FAIL_TAGwas applied and could never clear. It now only marks fields paperless-ngx actually rejected. Thanks @shiaho777, and @dertbv for the report. (#1086, closes #1053)
🕘 History you can read, and undo that works
- History shows names, not IDs. Tag changes used to read
Previous: [Inbox paperless-gpt-auto]→New: [3 2]. Both sides now show names, and so do correspondent and document type. Thanks @zhzy0077, and @kkettinger for the report. (#1101 from #1078, closes #842) - Undoing a tag change works again. It had been failing with an error for every entry, because history was written in one format and read in another.
- Older history entries are restored safely. Entries written before this release store tag lists in a format that can't tell a tag called
Boîte de réceptionapart from three separate tags. paperless-gpt now matches them against your existing tags, and refuses the undo with a clear message instead of guessing when it can't.
🔬 OCR
- Non-PDF originals no longer break OCR. With
PAPERLESS_ARCHIVE_FILE_GENERATION=never, paperless-ngx serves the original file, which may be an image. Images are now processed as a single page; other formats fail with a message that explains why. Images are converted to JPEG on the way, so providers that check the image type (Anthropic does) accept them, and theIMAGE_MAX_*size limits apply to phone photos too. Thanks @zhzy0077! (#1100 from #1079)
🏷️ Better suggestions
- The correspondent is the sender, not the subject. Smaller models sometimes named the patient instead of the hospital that wrote the letter, or returned a chatty preamble that ended up saved as a correspondent name. The default prompt now rules out both. Thanks @Earthlike1104! (#1081)
Default prompt changes only reach new installations. If you've used paperless-gpt before, your copy in
prompts/is kept. Compare it withdefault_prompts/correspondent_prompt.tmplto pick up the change.
⚠️ Known issues
PDFium can fail to draw text in certain embedded fonts. In the 726-document comparison, one born-digital invoice rendered with whole lines of text missing. Two other renderers draw the same file correctly, and current native PDFium shows the same gap, so the problem lies in PDFium itself rather than in paperless-gpt.
It only matters when LLM-based OCR runs on such a document: the missing lines are missing from the OCR result. The text paperless-ngx extracts on its own is unaffected. Born-digital PDFs already contain their text, so PDF_SKIP_EXISTING_OCR=true skips them entirely. If OCR output looks incomplete compared to the document, please open an issue. We'll collect affected cases and report them upstream.
Memory peaks a little higher while rendering. PDFium copies each page out of its WebAssembly sandbox, so a burst of OCR or page previews briefly holds more memory than before: about 320 MiB instead of about 270 MiB right after rendering five pages at full OCR resolution. It is released again while paperless-gpt is idle: a home server we tested had gone from 304 MiB back to 42 MiB when checked 45 minutes later, without intervention. The prepared renderer itself stays loaded while OCR is enabled and accounts for only about 30 MiB. If you run paperless-gpt with a tight container memory limit, leave a little headroom.
🙏 Credits
- @andre161292: automatic document linking (#1097)
- @shiaho777: three data-loss fixes for custom fields, invisible tags and the fail tag (#1084, #1085, #1086)
- @zhzy0077: readable history with working undo, and non-PDF originals (#1078, #1079)
- @Earthlike1104: the correspondent prompt fix (#1081)
- @LembkeM, @schubimice, @MSPTW, @dertbv, @kkettinger: reports precise enough to fix from
Full Changelog: v0.28.0...v0.29.0