pypi multidict 7.1.0

2 hours ago

7.1.0 is a correctness and housekeeping release. It changes what
keys() yields, fixes a long list of crashes and
inconsistencies that Python code run in the middle of an operation could
trigger, stores the common all-str multidicts in half the space, and
deprecates multidict.upstr.

Behaviour changes. keys() and iteration over a
multidict yield each key once, as spelled in its first item, so
len(d.keys()) counts distinct keys; items(),
values() and len(d) still see every item. A
str subclass's own lower() is no longer called to compute the
case-insensitive identity of a CIMultiDict key.

Robustness. Finalizers, __eq__(), __repr__(), lower() and
argument iterators that read or mutate the multidict in the middle of
__init__(), update(), merge(), del d[key],
popall(), popitem(),
to_dict() or repr() no longer crash the C
extension, lose or duplicate pairs, or make the two backends disagree. Two
data races against lock-free readers on the free-threaded build are gone
too.

Memory. A MultiDict whose keys are all exact
str, and a CIMultiDict whose keys are all exact
istr, are stored in a compact table without a separate
identity and hash per entry, which cuts their size by about 45%.

Bug fixes

  • Fixed building the C extension on platforms where char is unsigned,
    such as ARM, PowerPC, s390x and RISC-V, and with GCC 9, which flags
    CPython's own headers from 3.12 on; both failed on a sign conversion
    warning since -Wconversion took effect. GCC 9 builds the extension
    without -Wconversion -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1759, #1760.

  • Fixed getversion() returning only the low 32 bits of
    the version on Windows and on 32-bit platforms, where two different
    versions could compare equal, and fixed a race in the pure-Python
    implementation on GIL builds that could give two multidicts mutated at
    the same time from different threads the same version
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1607, #1610.

  • Fixed in on an items view truncating the length of a non-tuple,
    non-list operand to 32 bits in the C extension, which let an object
    reporting 2**32 + 2 items be indexed as a pair and raise
    TypeError instead of returning False -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1611.

  • Fixed adding to a full table of the C extension, which compacted it in
    place whenever it held deleted entries and so regained only as many
    slots as were deleted: a delete-then-add loop on a full table rebuilt
    it on every add. The table was resized by the number of live entries
    instead, as the pure-Python implementation did, which also let it
    shrink after deletions. Fixed extend(),
    update() and merge()
    reserving room by comparing table sizes, which ignored the room taken by
    deleted entries; the reservation checked the free room instead, in both
    implementations -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1615, #1617.

  • Fixed calling __init__() again on a live multidict losing and leaking
    whatever a finalizer of an old value added to it in the C extension, and
    miscounting len() in the pure-Python implementation, by installing
    the new contents and size before releasing the old pairs
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1619, #1626.

  • Fixed two data races on free-threaded builds against lock-free readers:
    copy() and __init__() from a multidict of the same kind read the
    source's reader count while its readers updated it, and __init__()
    on a live multidict rewrote the case-insensitivity flag its lookups
    read -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1621.

  • Fixed MultiDict.__init__() called on a CIMultiDict making it
    case-sensitive in the C extension, and CIMultiDict.__init__()
    rejecting a MultiDict there, by taking the mode from the
    instance rather than from the class whose __init__() was called, as
    the pure-Python implementation already did -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1622.

  • Fixed the C extension dropping the existing items when a multidict was
    re-initialized from itself, or from a proxy of itself, together with keyword
    arguments, as in d.__init__(d, key=value); the pure-Python
    implementation was not affected -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1666.

  • Fixed a use-after-free in popitem() when
    building the result ran Python code that mutated the multidict: a
    str subclass key's __str__() on
    CIMultiDict, or on Python 3.10 and 3.11 a garbage
    collection's finalizers, by removing the pair first, and fixed the
    pure-Python popitem() miscounting len() in the same case
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1625.

  • Fixed del md[key] and popall() also
    removing, and popall() also returning, the pairs a finalizer of a
    removed key or value added while the call was still running, by
    releasing the removed pairs only once every match is gone, and fixed the
    pure-Python implementation raising IndexError or miscounting
    len() in the same case, for md[key] = value too
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1627.

  • Fixed a crash when Python code run between the items of
    update() (the argument's own iterator, a key's
    lower() or a finalizer) read the multidict, where the pure-Python
    implementation returned None for the pairs the call was about to
    remove, and fixed update() and merge()
    overwriting or losing pairs when that code changed the multidict. An
    update from another multidict raised RuntimeError when a key's
    lower() mutated the source -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1628, #1744.

  • Fixed a use-after-free in to_dict() on Python
    3.10 and 3.11, where allocating a value list can run a garbage collection
    whose finalizers mutate the multidict, by refusing that mutation with
    RuntimeError as it is refused from a key's __hash__()
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1630.

  • Fixed repr() of a multidict reading freed memory on the GIL build
    when a key's or value's __repr__ called an extend(), update()
    or merge() that grew the table and then failed, by bumping the version
    on every resize -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1658.

  • Fixed CIMultiDict returning a key added as a
    str subclass with the spelling of the subclass's own
    __str__() instead of the key's, and stopped the C extension from
    replacing such a key with its istr while reading it, which
    could run the key's finalizer inside the read. Documented that
    CIMultiDict converts keys to istr
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1640.

  • Fixed the istr keys that
    extend(), update(),
    merge() and the MultiDict
    constructor took from a CIMultiDict: they carried the
    key's own spelling as their case-insensitive identity instead of its
    lower-cased form, so a CIMultiDict keyed by one of them
    found it under no spelling of the key. Changed the pure-Python
    implementation to hand over those keys as istr too, as
    the C extension did -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1647.

  • Fixed a lock-free lookup on the free-threaded build reading the hash
    through a key of a compact table that a concurrent delete had already
    freed, by taking a reference to such a key first unless it is the very
    object being looked up -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1653.

  • Changed CIMultiDictProxy.__init__() called without an argument to
    report the same TypeError message as MultiDictProxy and the
    pure-Python implementation -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1677.

Features

  • Made MultiDict in the C extension store a mapping whose
    keys are all exact str in half the space, keeping no separate
    identity or hash per entry, and did the same for a
    CIMultiDict whose keys are all exact
    istr, taking each identity from the key's canonical
    form, which cut its size by about 45%. Copying such a mapping became
    about 55% cheaper in instructions and clear()
    about 27%; any other key moves the mapping to the full layout, and
    clear() lets it start compact again. A lookup
    also stopped taking a reference to the identity of an exact str
    or istr key, which made d[key] and key in d up
    to 7% faster on the GIL build and up to 20% cheaper in instructions on the
    free-threaded one -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1644, #1648, #1651.

  • Sped up getall(), del d[key], replacing a
    key with d[key] = value and update() by 5
    to 14 percent, and up to 26 percent in the pure-Python implementation,
    when no key repeats, which each hash table started to track; the
    tracking made add() up to 10 percent slower,
    and building a small multidict from items up to 18 percent, both up to
    16 percent in the pure-Python implementation -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1654.

  • Replaced the set of visited entries that getall(),
    to_dict() and the items view's & kept with a
    check that each match's entry index is above the previous one's, since
    equal keys are first reached in insertion order. A
    getall() of a key with many values needed up to
    half fewer instructions -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1752, #1756.

  • Sped up update() and
    merge(): they stopped growing the table up front
    when it could hold the argument even empty, as dict.update() does,
    stopped holding extra references to the keys and values of a dict
    argument, and walked the hash chain for each item without a function
    call. An update of keys already present needed 13% to 29% fewer
    instructions -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1747, #1755, #1757.

  • Made d[key] = value in the C extension keep the collector for
    duplicate keys off the stack until a duplicate turns up, which made
    replacing a key about 4-6% and adding a new one about 2-3% cheaper in
    instructions -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1650.

  • Sped up building an istr from a str in the C
    extension by copying the string directly instead of going through
    str.__new__(); reading a str key of a
    CIMultiDict for the first time and
    popitem() became about 30% cheaper in
    instructions -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1645.

  • Stopped a multidict with no watchers from testing for them on every
    record made by del d[key], d[key] = value and pop(): each
    operation checked once and ran a copy without the watcher code, which
    recovered most of the cost the watchers C API had added to them. Changed
    free-threaded builds to reserve versions for each multidict in blocks of
    256, so a mutation advanced a counter in the object itself instead of
    looking up a thread-local one, which cut 2 to 5 percent of the
    instructions of every mutation -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1604, #1605.

  • Stopped clearing each field of a hash table that is being freed, and read
    its entry count once instead of after every release, cutting about 8% off
    tearing down a table of 200 entries in clear() and on deallocation;
    replaced Py_CLEAR() on success paths with Py_DECREF() or
    Py_XDECREF(), which made building a multidict from items up to 1.7%
    and comparing one with another mapping up to 4% faster
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1639, #1642.

Deprecations (removal in next major release)

  • Started emitting DeprecationWarning on access to multidict.upstr,
    the alias for istr deprecated since 2.0, and removed it
    from multidict.__all__
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1680.

Removals and backward incompatible breaking changes

  • Changed keys() and iteration over
    MultiDict, CIMultiDict and their
    proxies to yield each key once, as spelled in its first item, so that
    CIMultiDict([('X-Foo', '1'), ('x-foo', '2')]) iterates over 'X-Foo'
    alone and len(d.keys()) counts distinct keys; items(),
    values() and len(d) kept every item
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1654.

  • Stopped calling a str subclass's own lower() for the
    case-insensitive identity of a key in CIMultiDict and in
    IStr_FromUnicode(); the identity came from str.lower() instead,
    so computing it ran no user code -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1746.

Packaging updates and notes for downstreams

  • Moved the project metadata from setup.cfg to pyproject.toml per
    621, and removed the now-empty setup.cfg along with its obsolete
    [bdist_wheel] universal flag. Adopted 639 license metadata:
    declared the license as the SPDX expression Apache-2.0 and moved
    license-files into the [project] table, which raised the
    build-time requirement to setuptools >= 77.0. Built distributions
    started carrying License-Expression in place of the legacy
    License field -- by @aiolibsbot.

    Related issues and pull requests on GitHub:
    #1646.

  • Added the benchmarks/ scripts to the source distribution, so that the
    test suite run from an unpacked sdist also covers the benchmark driver
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1624.

Contributor-facing changes

  • Stopped the wheel jobs that run under QEMU from running the pure-Python
    half of the tests marked expensive, and the tests marked threaded
    on GIL builds, which had grown enough since 7.0.0 to take the manylinux
    riscv64 job past its two-hour timeout; native jobs still run them
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1762.

  • Added tools/check_inlining.py and an Inlining CI job, which
    fail when GCC stops inlining a helper that a hot path depends on, or
    makes an out-of-line copy of an inline function from the CPython headers
    or from pythoncapi_compat.h such as Py_DECREF(), since each time
    that happened a benchmark regressed on code the change never touched
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1612, #1686.

  • Stopped passing -Wno-conversion to the C extension build, which
    had silently turned -Wconversion off, and added the explicit casts
    it then required -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1611.

  • Dropped the macOS and Windows legs of the CI test matrix, since
    cibuildwheel already ran the same test suite against the very wheels
    those jobs installed. Moved every artifact the CI test jobs consume into
    one first stage, the source distribution together with the pure-Python
    and Linux binary wheels, and switched the remaining test jobs to that
    source distribution instead of a repository checkout. The wheel builds
    for the other platforms stopped gating the test matrix, the
    windows-11-arm wheel build got limited to release tags again, and
    x86_64 macOS wheels started being tested under Rosetta outside pull
    requests only -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1575, #1638.

  • Added CodSpeed benchmarks for mutating a multidict that a watcher is
    attached to and for re-initializing a populated multidict, started
    benchmarking the FT build alongside the GIL one, and split each build's
    benchmark CI job into two parallel shards of similar length, selected by
    a new benchmark_shard_2 pytest mark -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1608, #1620.

  • Changed benchmarks/callgrind_driver.py to run its Callgrind children from a
    fixed-length staging directory holding only the modules they import, so that
    two checkouts of one commit at different paths no longer measure up to 3%
    apart on allocating operations, and pinned mimalloc's arena address and purge
    timer, which had moved free-threaded counts between two runs of one tree.
    Stopped the driver from collecting whole-process counts, which cannot be
    compared with bracketed ones, unless it is given --whole-process, and
    made its checks before measuring, including --self-check, cover only
    the implementations selected with --impl, so a pure-Python install can
    be measured -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1623, #1624, #1637.

  • Added a test that runs every mutating method of both multidict classes
    with keys, values and argument items whose finalizers mutate the
    multidict, and checks the C implementation against the pure-Python one,
    along with reusable helpers for writing such tests, and a threaded stress
    test that re-initializes a shared multidict whose old values each have
    such a finalizer while other threads run lock-free lookups and iterate it
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1621, #1631, #1633.

  • Made the ThreadSanitizer run reliable: added
    tools/tsan_suppressions.txt for the faulthandler watchdog's
    traceback dumps, used by the CI job and by the documented local command,
    which also gained --no-cov; and changed the two pure-Python deadlock
    tests for reciprocal view operations to fail only when their workers stop
    making progress for 20 seconds, rather than when they have not finished
    within 20 seconds; and rewrote the retry loop of the test for
    reference leaks on memory errors so that it takes the same branches on
    every Python version
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1655.

  • Renamed the hypothesis-freethreading CI job to hypothesis-ft
    (Hypothesis FT), on a par with the GIL one, and removed the update
    argument of the debug-only ASSERT_CONSISTENT() check, which stopped
    doing anything once update() kept its doomed
    entries whole -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1609, #1629.

  • Moved the sanitizer build recipes and the performance measurement
    procedure out of AGENTS.md into Claude Code skills under
    .claude/skills/, and dropped its file layout table, so that less of
    it is loaded into every agent session
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1692.

Miscellaneous internal changes

  • Reworked how the C extension spends GCC's inlining budget, which it sits
    at as a single translation unit. Dropped the ALWAYS_INLINE macro and
    inline from almost every helper, leaving the decision to the
    compiler, and turned the hash table index lookup, store and probe helpers
    into macros. Moved the table resize and reservation, clear(),
    popitem(), the equality comparison, the iterator and view
    constructors, the end of a lock-free read and other rarely run code out
    of line, and made the proxy methods tail-call the shared ones. Marked the
    error-raising helpers, __sizeof__, __reduce__() and the proxy
    constructors' error paths cold, and kept NOINLINE only for the few
    helpers GCC re-inlined without it. Moved argument parsing into the method
    entry points and returned the keyword-bound arguments by value, which
    dropped a stack protector canary from every get(), getone(),
    getall(), add() and setdefault() call
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1613, #1614, #1661, #1664, #1667,
    #1670, #1671, #1672, #1683, #1690,
    #1691, #1693, #1694, #1695, #1697,
    #1698, #1699, #1700, #1701, #1702,
    #1703, #1704, #1705, #1707, #1709,
    #1711, #1712, #1713, #1714, #1717,
    #1718, #1719, #1720, #1722, #1723,
    #1724, #1726, #1727, #1730, #1732,
    #1733, #1734, #1735, #1737, #1739,
    #1740, #1749, #1750, #1751.

  • Shrank the C extension by merging duplicated code: the bodies of
    __init__(), extend(), update() and merge(), of pop()
    and popone(), the per-class copies of the entry points, the update
    loops and the tp_vectorcall slots, which read the class at run time
    instead, the repeated tails of the table rebuilds and of getall() and
    popall(), and the module setup, which became two table loops. The
    set operators and comparisons of the keys and items views were rewritten
    around shared helpers compiled for size. When they landed, the view
    rewrite made the compiled extension about 6% smaller and reading the
    class at run time about 15%, which left room for the compact tables; the
    view set operators became up to 17% slower and update() and
    merge() up to 3% slower -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1669, #1673, #1675, #1677, #1678,
    #1679, #1706, #1715, #1716, #1725,
    #1728, #1729, #1731.

  • Reorganized how the C extension's hash table lays out and reaches its
    entries: split the entry into a key and value layout that every table
    kind shares and a layout that adds the identity and its hash, computed
    the entry size with sizeof, told a deleted entry by its key alone,
    replaced the md_pos_t cursor and the entry stepping helpers with
    typed, indexed access, and walked a table in one loop per table kind and,
    for compact tables, per class. repr(), the views' | and -,
    the update from another multidict and the comparison with a plain
    mapping moved onto the shared walk over all entries
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1641, #1649, #1656, #1659, #1660,
    #1662, #1663, #1665, #1674, #1676,
    #1741, #1742, #1743, #1745, #1748.

  • Merged the C extension's two ways of dropping a hash table it no longer
    used, freeing it at once on GIL builds and retiring it on free-threaded
    ones, into one helper -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1618.

  • Upgraded pythoncapi_compat, and stopped copying the fast
    path of PyUnstable_TryIncRef() into the free-threaded C
    extension, where it read CPython's private reference count fields, by
    calling the function itself instead -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1687, #1688.

  • Removed dead code from the C extension: a self-update branch no caller
    could reach, version guards that are always true on the supported Python
    versions, unused helpers and single-use forwarding wrappers. Renamed the
    header-local helpers to carry a leading underscore, moved
    istr_canonical() next to the istrobject struct, and folded the
    pure-Python update() match test back into one condition
    -- by @asvetlov.

    Related issues and pull requests on GitHub:
    #1635, #1668, #1689, #1753, #1754.

  • Added const qualifiers to internal C helpers that read table, reference list,
    and watcher metadata -- by @anshurajbisoyi98-ctrl.

    Related issues and pull requests on GitHub:
    #1643.


Don't miss a new multidict release

NewReleases is sending notifications on new releases.