I'm pleased to announce the release of pandoc 3.12,
available in the usual places:
Binary packages & changelog: https://github.com/jgm/pandoc/releases/tag/3.12
Source & API documentation: http://hackage.haskell.org/package/pandoc-3.12
This release focuses on performance. Here are some benchmarks showing
the improvement since pandoc 3.10.2:
There shoud be even more dramatic gains in image-heavy documents, because
of optimizations in ImageSize.
In addition to performance improvements, there are many other small
improvements and bug fixes. See the changelog for full details.
Most of the performance gains and and many of the bugs were found
with the help of Claude Fable.
API changes:
- Text.Pandoc.ImageSize: ImageType now derives Eq.
- Text.Pandoc.Sources: new functions takeWhileP, takeWhile1P.
- Text.Pandoc.Shared: new functions stringifyInlines,
compactifyTable. FromText constraint added to the signatures
of htmlAttrs and tagWithAttrs.
Other changes of note:
-t xmlnow respects--standalone, and produces a
fragment (just the blocks) if it is not provided, like-t native.- The default CSS now includes dark mode support.
Thanks to all who contributed, especially new contributors Markdown reader:
Typst reader:
LaTeX reader:
Docx reader:
ODT reader:
HTML reader:
Muse reader:
Org reader:
RST reader:
CommonMark reader:
AsciiDoc reader:
Djot reader:
Vimwiki reader:
Typst writer:
RST writer:
Docx writer:
TEI writer:
LaTeX writer:
RST writer:
Native writer:
XML writer:
EPUB writer:
Powerpoint writer:
ANSI writer:
HTML writer:
Org writer:
Org reader and writer:
Text.Pandoc.ImageSize:
Text.Pandoc.SelfContained:
Text.Pandoc.UTF8:
Text.Pandoc.Class:
Text.Pandoc.MediaBag:
Text.Pandoc.XML.Light:
Fix escaping of repeated Make Avoid round-trips in Parse XML fragments from the event stream instead of using xml-conduit’s document parser, which requires a single root element. Our earlier woraround with a wrapper element was fragile. Behavior changes:
Use Expose ConfigPP now has a field Text.Pandoc.Sources:
Text.Pandoc.Translations:
Text.Pandoc.Chunks:
Text.Pandoc.Data:
Text.Pandoc.Shared:
Text.Pandoc.Writers.Shared:
Text.Pandoc.Parsing:
HTML template:
reveal.js template:
flake.nix: parse allow-newer and allow-newer-deps in stack.yaml.
Make Depend on commonmark 0.3.1, commonmark-extensions 0.2.7.3, commonmark-pandoc 0.3.0.2 (major performance improvements).
Depend on released asciidoc 0.1.1 (major performance improvements).
Use released texmath 0.13.3 (major performance improvements).
Depend on released djot 0.1.4.3 (major performance improvements).
Use released citeproc 0.14 (major performance improvements).
Use released doclayout 0.6.
Use released zip-archive 0.5 (major performance improvements).
Use released skylighting-0.15 (major performance improvements).
Depend on released doctemplates 0.11.1.
Depend on released typst 0.12.
Require text >= 2.0.
Bump upper bound for unicode-data.
Allow crypton 2.0.x.
Allow Diff 2.0.
Add Fix Add Remove tested-with from cabal file. We tend not to keep it up to date.
Fix typo in Lua filter example (#11875, Andonome).
Fix a bug in jats-reader.xml (duplicate attribute)
Andonome, Clar Fon, Gaurav Vijay Jadhav, Robert Szarka,
Samuel Huang, Yusuf Efe, and zenor0.
Click to expand changelog
alert is now added to the produced Divs.
base64DataURI.
bareURL before trying uri/emailAddress.
parseWithString' when parsing a block quote or list item. Otherwise it can happen that by the time parseWithString' is called, the position has already been set to the next file on the command line. Fixes an odd bug with rebase_relative_paths (#11888).
rebase_relative_paths: recognize URLs with unknown schemes (#11858).
takeWhile1P in the hot inline parsers str, code, enclosure, and mmdShortSubscript.
parseWithString' when parsing a block quote or list item (#11888). This ensures that rebase_relative_paths will see the right source file.
form: "prose" citations to AuthorInText (#11846, Samuel Huang).
pInline, only perform the label-target check for ref elements, not for every inline element.
highlight as a mark span (#11879, Samuel Huang).
#par (explicit paragraph element).
\qed to produce U+00A0 (nbsp) instead of BEL.
unescapeURL handling of escaped backslash.
\newif names begin with “if”.
## in tokenizer.
retokenizeComment.
untokenize linear instead of quadratic.
macroDef fast on non-macro-defining commands.
peekTok instead of going through satisfyTok.
doMacros state update for non-macro tokens.
inlineCommands, etc.) once per parse.
\iftrue etc.
comment-id instead of id in AST for comments.
textProperties on paragraph styles (#2623). Previously these just got ignored, not applied to the paragraph’s text.
<input> is a void element.
</tr> in tables. The </tr> closing tag is optional in HTML.
<col> elements.
raw_html for inline <style> elements.
<ol class="fancy lower-roman"> got DefaultStyle, because the whole class attribute was compared against the known style names. Check each class individually. As a side effect, an unrecognized class no longer prevents falling back to the style attribute.
htmlTag: Don’t copy the remaining input on each invocation. A space was appended to the remaining input to guarantee a TagPosition token after the parsed tag; since the input is a strict Text, this copied the entire remaining input every time htmlTag was called (e.g. for every inline HTML tag in a markdown document), giving quadratic behavior in tag-dense documents. Instead, handle the case where the tag is the final token by computing the end-of-input position directly.
pSpanLike. The inline dispatcher already knows which span-like element it is looking at, so there is no need for pSpanLike to try a parser.
pre/code attribute precedence for first-wins dedup.
pSatisfy as a single parsec primitive.
pTagText: when the text contains no character that could parse as anything but Str, Space, or SoftBreak under the enabled extensions (and we are not in a pre element), return B.text directly.
pTagContents to try pStr and pSpace before the math, smart punctuation, and raw TeX parsers. This is safe because pStr cannot consume the special characters that start those parsers, and it avoids most guard checks on the slow path. This makes the html reader benchmark about 33% faster and halves its allocation.
str earlier in inline parser. This makes the reader 2x faster and reduces heap allocation by 60%.
.pdf as an image format (#11859). This matches the behavior of Emacs, which will render [[file:foo.pdf]] as an image.
#+OPTIONS: ^:nil disable all sub-/superscript parsing.
- and _ in inline footnote labels. Org footnote labels may contain word-constituent characters, hyphens and underscores.
lookupGE to find the next anonymous key. Parsing a document with 16000 anonymous links drops from 5.2s to 2.0s.
tex_math_gfm handling.
map id in sourceToToks in the common case where the source starts at line 1.
nowrap to just footnote label, not body. This bug surfaced after nowrap was fixed in doclayout.
_ to start all bookmark names (#11845). This ensures that they are “hidden” and will not be read by screen readers.
withDirection.
withDirection for Space and SoftBreak.
convertSpace linear instead of quadratic. On an ad hoc benchmark (4 paragraphs of 80K words each), conversion time drops from 3.2s to 1.1s, and time no longer depends on paragraph length (16 x 20K words previously took 1.8s, now also 1.2s).
comment-id instead of id in AST for comments. Note that the Docx writer will still interpret an id attribute for legacy compatibility, so if you use markdown files that specify id, they should still work.
rend, not rendition attribute, on milestone (#11842, Yusuf Efe) rendition takes pointers to rendition descriptions, while rend is the free-text attribute, which is what a plain “line” value needs.
footnotehyper for notes in longtable. Instead, generate them manually as we do for floating tables. This removes our dependency on footnotehyper, and resolves a compatibility problem with endfloat (#11857).
stInternalLinks state field. The field has not been read since the writer switched to always adding hypertargets, but we were still doing a full-document query to populate it.
sectionHeader. The note-free and link-free variants of the heading text were rendered for every heading, even though the former is only needed when the heading contains a note, image, or identified span, and the latter only for unnumbered, listed headings.
inlineListToLaTeX. The strut-insertion and quote-kerning fixups were separate list traversals, allocating an intermediate list each, and this function is called at every level of inline nesting. Combine them into a single pass.
lstinline delimiter.
SOURCE_DATE_EPOCH is set.
showHex instead of (slower) printf in toLabel.
\ markers. With --reference-links this could make the inline reference and its definition render differently, producing a broken RST reference.
ppDoc, which shows the document, tokenizes and re-parses the result into a generic Value, and lays that out via Text.PrettyPrint.HughesPJ. The layout step dominated the cost of the writer (and of any pipeline producing native output). We now build a width-cached layout tree directly from the AST and render it with a small renderer that reproduces HughesPJ’s layout algorithm exactly (including the ribbon computation with ribbonsPerLine = 1.2 and the treatment of glued closing delimiters), so the output is byte-for-byte identical to before. Verified against the old binary on all golden .native files and the markdown test corpus at many column widths, in both standalone and plain modes. The writer is about 4x faster. Also remove the pretty and pretty-show dependencies.
--standalone. When standalone is selected, we get a full Pandoc element with xml header and metadata. When not, we get a fragment – just the blocks.
ppcElement, using a configuration that treats elements with inline content as inline tags, so that no significant whitespace is added inside them. Consecutive text nodes are merged, SoftBreak is written as a literal newline, and whitespace runs that would not survive a roundtrip (e.g. " \n" or "\n\n") are encoded as Space and SoftBreak elements.
typst:property). Use the common convention of encoding such characters as _xHHHH_, where HHHH is the hexadecimal code of the character: the writer encodes attribute names (foo:bar becomes foo_x003A_bar) and the reader decodes them.
("k","") disappeared and did not round trip.
runPure (writeHtmlStringForEPUB ...). Since every one of these invocations starts with a fresh CommonState, the translations YAML file was re-read and re-parsed for every TOC item, which accounted for a significant part of the EPUB writer’s run time on documents with many sections.
intrinsicEventsHTML4. The list had onmouseout twice and was missing onmousemove.
<p></p> for paragraphs with no rendered content.
--id-prefix in EPUB3 footnote section id.
strToHtml more efficient. Replace the T.groupBy-based implementation, which allocated a list of Text fragments and round-tripped through String, with a simple T.break scanner.
#+begin_export html for raw HTML blocks. #+begin_html was removed in Org 9.0 (2016).
Str "." and Str ")", since T.all isDigit "" is True. Require at least one digit.
<<>> for spans with no id.
= delimiter for inline code containing =. Org has no escape mechanism inside verbatim text, so =code with == did not parse as verbatim. Fall back to the equivalent ~...~ delimiter when the content contains = (and no ~).
org-link-escape does: backslash-escape brackets and double backslash runs occurring before a bracket or at the end of the target. Link descriptions cannot contain escapes; instead, like org-link-make-string, insert a zero-width space between consecutive closing brackets and before a closing bracket at the end of the description.
escapeString allocated one Doc node per character for any string containing a non-alphanumeric character. Split on the (rare) special characters instead and emit intervening text as single literals. No change in output.
length stNotes + 1 walked the accumulated note list for every footnote, making note numbering quadratic in the number of notes.
^[ \t]*,*(\*|#\+), i.e. lines already starting with commas before * or #+ get an additional comma; unescaping removes one comma from such lines. The writer previously left a literal ,#+foo line unescaped, and the reader then stripped its comma when reading the result back, corrupting the code on round trips. Writer and reader now both handle runs of commas, matching Emacs.
pdfSize. Treat malformed streams as a parse scanner (so we keep scanning the rest for a /MediaBox).
viewBox.
findSvgTag by using a single pass. Up to 60X faster on files with few < characters.
writerDpi for AVIF images instead of hardcoding 72. With the default options this changes the assumed resolution from 72 to 96 dpi.
numUnit. So e.g. width="3 cm" is now recognized.
sizeInPoints.
makeDataURI.
@import fallback output.
</script check case-insensitive.
url(#...) occurrences in SVG attributes. Previously only the first got prefixed rewritten.
isHtml5 field from ConvertState.
role and aria-label when inlining SVGs. Do not add them to other elements with src attributes.
toText and toTextLazyunconditionally ran a CR-removing filter over the input, allocating a full copy of the document even in the common case where no CRs are present. Check for a CR first (B.elem, a fast memchr) and reuse the input buffer unchanged if none is found; for the lazy variant, do this chunk-wise to preserve laziness. On a 10 MB LF-only input this makes toText over 4x faster; when CRs are present the extra scan is not measurable.
readFile exception-safe. Use withFile instead of openFile so the handle is closed even if reading throws.
toTextM. Skip the CR-filtering copy when the input contains no CRs, as already done in Text.Pandoc.UTF8.toText.
runSilently error-safe. Previously, if the action passed to runSilently threw an error that was later caught, the verbosity remained pinned at ERROR and all previously accumulated log messages were lost. Now the original log and verbosity are restored even when the action fails.
isRelativeToParentDir. Compare the first path component rather than just looking at a prefix, to correctly handle paths like ..foo/bar.yaml.
extractURIData. The base64 indicator in a data URI is the final parameter of the media type and may follow other parameters, e.g. charset. Previously, the code expected ;base64 to be the only parameter.
extractURIData.
toTextM errors. Scan for the first invalid UTF-8 sequence and report its actual position and byte.
setNoCheckCertificate. The HTTP manager is created lazily with TLS settings based on stNoCheckCertificate and then cached in CommonState, so changing the option after the first request had no effect. Discard the cached manager when the option’s value changes.
addToFileTree.
openURL into Text.Pandoc.Class.IO.HTTP. Commit 455bea9 added the new module but did not register it in pandoc.cabal or remove the original definitions.
logOutput: avoid multiple hPutStrLn, which can cause confusing interleaving.
data: and file: URI schemes case-insensitively.
mediaContents is left lazy so contents need not be forced at insert time.
Text.Pandoc.URI.isURI in canonicalize. Network.URI.isURI treats Windows drive-letter paths like c:/foo.png as URIs.
.. as a path component. The insertMedia check used isInfixOf, so a harmless name like foo..bar.png was silently renamed to its content hash.
. and .. components in canonicalize. normalise does not remove redundant path components, so img/../a.png and a.png were distinct keys. Use makeCanonical (as PandocPure’s FileTree already does for its path-indexed map), which also handles duplicate and trailing slashes, replacing backslashes with slashes first.
mediaPath collisions between keys. The friendly mediaPath was derived by percent-unescaping the key, so distinct keys like a%20b.png and a b.png produced the same mediaPath (“a b.png”) and silently clobbered each other on extraction (and inside docx/epub archives). Now the original name is only kept if the key contains no percent sign, so mediaPath equals the key and distinct keys yield distinct paths; anything percent-encoded gets a content-hash name. Hashed names can only coincide for identical contents, which is harmless.
]]> in CDATA.
escStr more efficient.
ppCDataS prettify path. This makes showCData and ppcCData unused, so they are removed.
text-builder package for rendering, instead of text’s lazy Text builder. On a large document this cuts docx conversion time by about 13%; output is byte-for-byte identical.
ppcTopElement from Output.
inlineTag that checks for inline tags. Inline tags are printed on one line and not indented, by default. Export useInlineTags, prettyConfigPP.
takeWhileP/takeWhile1P combinators [API change]. Character streams over Sources previously had to be consumed one character at a time via satisfy, at a cost of several allocations and monadic binds per character. The new combinators scan a whole run of matching characters with a single parser invocation using T.span, while replicating the exact semantics of T.pack <$> many/many1 (satisfy f), including empty-chunk handling, position updates at chunk boundaries, and parsec’s error messages.
takeWhileP/takeWhile1P combinators across readers.
setTranslations now keeps an already-loaded translation table when the language is unchanged, instead of unconditionally clearing the cache and forcing a re-read and re-parse of the translations YAML file on the next translateTerm.
nav-path attribute and rmNavAttrs walk. This no longer did anything; output is unaffected.
compactifyTable for tables produced by all readers. Remove old ad hoc paraToPlain at the table cell level. This should ensure that we don’t get tables that mix Plain and Para (#11864). Such tables tend to look funny when rendered in docx and other formats.
getDataFileNames with -embed_data_files. Previously it was not looking in the right directory and not recursing.
taskListItemFromAscii: Fix incorrect treatment of [ ] as checked.
stringifyInlines, a single-pass stringify for inlines [API change]. This is about 6x faster than stringify for long inline sequences
stringify by making it accumulate [Text] and concatenate once at the end, instead of mappending at every node.
stringifyInlines instead of stringify where possible.
compactifyTable. [API change] This converts cells that consist in a single Para block to a Plain, provided the table contains only such cells (or empty cells).
decrementTrailingRowSpans. Entries whose RowSpan fell to 0 were left in the map, relying on every consumer to guard against them. Delete them instead.
toTaskListItem.
endsWithPlain look inside DefinitionList. endsWithPlain recursed into the last item of BulletList and OrderedList but ignored DefinitionList, so list items ending with a compact definition list were treated as loose by the RST, Org, and Haddock writers.
htmlAttrs. This adds a FromText constraint to htmlAttrs and tagWithAttrs [API change].
ensureValidXmlIdentifiers.
splitSentences.
lookupMetaBool: treat empty block or inline list as False.
htmlAttrs: escape id and class attributes, like the others.
stripLeadingTrailingSpace - strip multiple Space, if present.
toSubscript: handle minus sign.
ensureValidXmlIdentifiers for Figure and table sub-elements, resolving a bug that produced broken internal links in the HTML4/XHTML, EPUB, DocBook, TEI, ICML, FB2, and ODT writers.
uriScheme by using a trie. Up to 15% faster in URL-heavy documents with autolink_bare_uris.
$ checks in math when delim is \( or \\(.
anyOrderedListMarker more efficient.
embed_data_files flag default to True. Remove flag settings from cabal.project. This makes it possible to override it on the command line.
tools/diff-golden-tests.sh.
tools/diff-zip.sh on non-Darwin.
tools/benchplot.js. This creates a nice graph comparing two benchmarks.
typst-properties.md: fix fill syntax in Typst property examples (#11855, zenor0).