This release contains breaking changes. It focuses on speed improvements and corner cases.
The main change is that the Modest backend is no longer available.
It is outdated and unmaintained, contains bugs, and does not follow modern HTML5 standards.
The second is that parse_fragment() is gone. It re-implemented fragment parsing by guessing
whether the input was a document or a fragment; LexborHTMLParser(html, is_fragment=True) already
does this properly, so the guessing layer has been removed.
The rest of the changes fix corner cases where the bugs were happening rarely,
usually when heavily modifying the tree.
- Breaking: Remove the Modest backend.
selectolax.parseris now a stub that raisesImportError
on import. Use the lexbor backend (from selectolax.lexbor import LexborHTMLParser) instead. - Breaking: Remove
parse_fragment(). - Breaking: Fix
css_matchesandany_css_matchesmissing matches outside the first top-level node of an HTML
fragment. They now search the same scope ascss. - Breaking: Fix
tagsandstrip_tagsdoing nothing on an HTML fragment.tagsreturned an empty list and
strip_tagsremoved nothing, because the fragment was not reachable from the document node. - Breaking: Fix the internal
<html>wrapper of an HTML fragment being reported as a match. - Breaking: Fix
scripts_containandscript_srcs_containmissing matches outside the first top-level
node of an HTML fragment. They now search the same scope ascss. - Breaking: Fix
scripts_containandscript_srcs_containanswering from another scope's cached
result. The cache is keyed by the node the search is rooted at rather than the node it was called on. - Breaking: Fix
select()missing matches outside the first top-level node of an HTML fragment.
It now searches the same scope ascss. - Breaking: Fix
traverse()covering only the first top-level node of an HTML fragment, and
yielding nothing when the fragment starts with a text node. - Breaking: Fix
attribute_longer_thanandany_attribute_longer_thanreturning inconsistent results - Breaking:
attributesandattrsnow report an attribute's qualified name instead of its local name.
An element carrying bothhrefandxlink:hrefused to collapse them into a singlehrefkey. - Breaking: Fix
iter()skipping the remaining children when a node is removed during iteration - Breaking: Fix an HTML fragment whose first node is text dropping the rest of the fragment.
LexborHTMLParser('a<span>s</span>', is_fragment=True).text()returned'a'instead of'as'. - Fix
text()silently returning truncated text instead of raising when a fragment fails to be collected. - Fix
scripts_containandscript_srcs_containsometimes returning wrong results due to HTML mutations. - Fix a single undecodable byte in an untrusted document making
html,inner_html,html_pretty,
attributes,attrs,id,tag,text_lexborandtext_contentraiseUnicodeDecodeError.
Bytes Lexbor passes through verbatim are now substituted with U+FFFD, the waytext()already did. - Fix
text_lexborsometimes holding temporary memory longer than needed - Improve memory consumption, potential stack overflow and slow extraction in
merge_text_nodes(). It is now up to 10 times faster. - Improve speed of
__eq__ - Improve speed of
attrs.items()andattrs.values(). - Fix the
inner_htmlsetter attaching element children to non-element nodes - Fix
headandbodydangling after settinginner_htmlon the<html>element - Fix the
inner_htmlsetter freeing the replaced children - Fix
rootgoing stale on an HTML fragment. - Fix a segfault when walking the tree from the root of an HTML fragment that had been detached from the
fragment, for example byunwrap(). - Avoid segfaults when hitting OOM
- Improve performance of the
textmethod. Text fragments are now concatenated as raw bytes, up to 5x faster. - Fix
skip_emptybeing ignored bytext(deep=True) - Fix
text()raisingUnicodeDecodeErroron undecodable bytes whendeep=False. It now substitutes
U+FFFD, like the deep path always did - Fix handling of the
idmethod on text nodes - Fix
attrs[key] = valueraisingAttributeErrorinstead ofTypeErrorwhenvalueis not a string - Prevent segfaults when instantiating
LexborNodeorLexborAttributesdirectly. - Fix
unwrap()corrupting the tree in some cases. - Fix
headandbodygoing stale once<head>/<body>is removed from the document. - Fix
attrsreading freed memory when it outlives the node it was obtained from. - Fix memory leak in
attrs[key] = None; it leaked the value buffer header on every call. - Fix possible memory leak in
clone() - Add
encoding=TruetoLexborHTMLParser, which detects the encoding ofbytesinput and
transcodes it to UTF-8 before parsing.