CCCL Python Libraries (v1.2.1)
Previous release: v1.2.0.
These are the release notes for the cuda-cccl Python package version 1.2.1. This is a maintenance release with no new public Python APIs. The main changes are a self-contained Windows wheel, a workaround for a numba-cuda-mlir regression, clearer diagnostics for unsupported operations on structured data, and reduced Python-side overhead in cuda.compute calls.
Installation
Please refer to the install instructions here.
Packaging and dependencies
-
Windows wheels now bundle
msvcp140.dll(#11495)The Windows
cuda-ccclwheel is repaired withdelvewheelto include the MSVC C++ runtime DLL needed bycccl.c.parallel.dll. It no longer relies on whichevermsvcp140.dllhappens to be installed on the user's machine. The Windowscuda.computetest lanes now run without installing that DLL separately. Other packages, such as CuPy, may still have their own runtime requirements. -
numba-cuda-mlir0.5.3 excluded (#11617)The
cu12,cu13,sysctk12, andsysctk13extras now requirenumba-cuda-mlir>=0.5.2,!=0.5.3,<0.6. Version 0.5.3 can fail when a stateful operator with the same state dtype and shape is compiled a second time in one process.
Bug fixes
-
Clearer errors for unsupported built-in operations on structs (#11404)
On the current V1 backend, using a built-in
OpKindwith a struct or other opaque storage type now raises a descriptiveTypeErrordirecting users to provide a custom operator, instead of failing later with a generic NVRTC compilation error. This does not change operations that use a built-in comparator on ordinary keys while carrying struct values.
Performance
-
Less per-call overhead in
cuda.compute(#11497)Shared iterator handling now caches whether an iterator represents a pointer and reuses its
Pointerwrapper. Device-array pointer lookup also caches the appropriate accessor for each array type. These changes apply across algorithms without changing their public APIs or results. The PR measured approximately 9–20% less host-call overhead in representative reduce, scan, segmented-reduce, and histogram benchmarks; those measurements exclude GPU kernel execution.
Notes
- v1.2.1 includes the Python-package changes since v1.2.0; there were no intervening
cuda-ccclreleases.