github NVIDIA/cccl python-1.2.1
CCCL Python Libraries (1.2.1)

3 hours ago

CCCL Python Libraries (v1.2.1)

Previous release: v1.2.0.

These are the release notes for the cuda-cccl Python package version 1.2.1. This is a maintenance release with no new public Python APIs. The main changes are a self-contained Windows wheel, a workaround for a numba-cuda-mlir regression, clearer diagnostics for unsupported operations on structured data, and reduced Python-side overhead in cuda.compute calls.

Installation

Please refer to the install instructions here.

Packaging and dependencies

  • Windows wheels now bundle msvcp140.dll (#11495)

    The Windows cuda-cccl wheel is repaired with delvewheel to include the MSVC C++ runtime DLL needed by cccl.c.parallel.dll. It no longer relies on whichever msvcp140.dll happens to be installed on the user's machine. The Windows cuda.compute test lanes now run without installing that DLL separately. Other packages, such as CuPy, may still have their own runtime requirements.

  • numba-cuda-mlir 0.5.3 excluded (#11617)

    The cu12, cu13, sysctk12, and sysctk13 extras now require numba-cuda-mlir>=0.5.2,!=0.5.3,<0.6. Version 0.5.3 can fail when a stateful operator with the same state dtype and shape is compiled a second time in one process.

Bug fixes

  • Clearer errors for unsupported built-in operations on structs (#11404)

    On the current V1 backend, using a built-in OpKind with a struct or other opaque storage type now raises a descriptive TypeError directing users to provide a custom operator, instead of failing later with a generic NVRTC compilation error. This does not change operations that use a built-in comparator on ordinary keys while carrying struct values.

Performance

  • Less per-call overhead in cuda.compute (#11497)

    Shared iterator handling now caches whether an iterator represents a pointer and reuses its Pointer wrapper. Device-array pointer lookup also caches the appropriate accessor for each array type. These changes apply across algorithms without changing their public APIs or results. The PR measured approximately 9–20% less host-call overhead in representative reduce, scan, segmented-reduce, and histogram benchmarks; those measurements exclude GPU kernel execution.

Notes

  • v1.2.1 includes the Python-package changes since v1.2.0; there were no intervening cuda-cccl releases.

Don't miss a new cccl release

NewReleases is sending notifications on new releases.