This is the release note of v14.0.0a1. See here for the complete list of solved issues and merged PRs.
💬 Join the Matrix chat to talk with developers and users and ask quick questions!
🙌 Help us sustain the project by sponsoring CuPy!
✨ Highlights
This is the first alpha release of the CuPy v14 series, containing:
- New type promotion rules and behaviors aligned with the NumPy 2 specification.
- 42 new NumPy/SciPy-compatible APIs, including
cupy.concat,cupyx.scipy.interpolate.CubicSpline,cupyx.scipy.spatial.Delaunay,cupyx.scipy.ndimage.find_objects, andcupyx.scipy.special.lambertw. See the Comparison Table for the detailed coverage.
Binary packages are available for testing. Try installing now by:
$ pip install cupy-cuda12x --pre -U -f https://pip.cupy.dev/pre🛠️ Changes without compatibility
- CuPy v14’s behavior will be aligned with NumPy v2.
- Type promotion rules are now NEP50 compatible. See Changes to NumPy data type promotion.
intis now 64-bit (int64) on Windows. See Windows default integer.- APIs removed in NumPy v2 (see Changes to namespaces) were marked deprecated in CuPy v14. Although they are kept available in v14 for smooth migration, they are planned to be removed in the next major release (CuPy v15).
- The behavior of
copyargument has been changed (#8545). See Adapting to changes in thecopykeyword.
- Support for Python 3.9, NumPy 1.22 and 1.23, SciPy 1.7, 1.8, and 1.9 has been dropped. (#8491)
cupy.random.choicemay return different results from CuPy v13. (#8483)- Building CuPy from source code now requires Cython 3.0. (#8457)
cupyx.scipy.linalg.{tri,tril,triu}APIs were removed from CuPy to follow the latest SciPy’s specification. Usecupy.{tri,tril.triu}instead. (#8499)- NumPy fallback mode (
cupyx.fallback_mode) has been removed as discussed in #8497. (#8816) - Legacy DLPack APIs (
cupy.toDlpackandcupy.fromDlpack) are now marked deprecated. Usecupy.from_dlpackinstead. See the documentation for the usage. (#8831)
📝 Changes
New Features
- Add
KDTreetocupyx.scipy.spatial(#7671) - Add
neighborsoption toRbfInterpolator(#7864) - ENH: cupyx/signal: add
sweep_poly(#7873) - Add 2D Delaunay triangulation (#7985)
- Add
cupyx.signal.pulse_compressionfrom cuSignal's non SciPy-compat API (#8022) - Add
LinearNDInterpolatortocupyx.scipy.interpolate(#8035) - Add
cupyx.signal.convolve1d3ofrom cuSignal's non SciPy-compat API (#8037) - Add
cupyx.signal.{firfilter,firfilter_zi,firfilter2}(#8052) - Add
cupyx.signal.{pulse_doppler, cfar_alpha}(#8057) - Add
cupyx.signal.{complex_cepstrum,real_cepstrum,inverse_complex_cepstrum,minimum_phase}(#8062) - Add
cupyx.signal.mvdr(#8077) - ENH: signal: add
lanczosandkaiser_bessel_derivedwindows (#8081) - Add
cupyx.signal.ca_cfar(#8087) - Add
cupyx.signal.convolve1d2o(#8101) - Add
cupyx.signal.freq_shift(#8128) - Add
lambertwfunction (#8140) - Add
cupyx.signal.channelize_poly(#8141) - Add cupyx.scipy.interpolate.CubicSpline (#8175)
- Add
apply_over_axesAPI (#8177) - Add
cupy.put_along_axisAPI (#8199) - Add
CloughTocher2DInterpolatortocupyx.scipy.interpolate(#8208) - Add
NearestNDInterpolatortocupyx.scipy.interpolate(#8220) - Add
NdBSplinetocupyx.scipy.interpolate(#8223) - ENH: cupyx/scipy/interpolate: add *UnivariateSpline for 1D smoothing splines (#8267)
- Add NdBSpline based interpolation methods to RGI (#8276)
- ENH: cupyx/interpolate: port
interp1dfrom scipy (#8289) - Add batched solve_triangular (#8329)
- Add Incomplete Elliptic Integrals to special (#8425)
- Support system allocated memory (#8442)
- Add CUDA graph debug function (#8502)
- Add
siciandshichito special for sine and cosine integrals (#8620) - Update
unique_xxx(nep52) (#8665) - Add
cupyx.scipy.ndimage.find_objects(#8916)
Enhancements
- Support for break and continue keywords in CuPy JIT (#8010)
- Make
cupyx.signal.radartoolsprivate (#8047) - Remove usages of
numpy.float_andnumpy.complex_(#8050) - Support cusparseLt 0.6.1 (#8074)
- Add incontiguous support for cutensor functions (#8149)
- Add complex support for the digamma function (#8163)
- Fix
expm(complex matrix)(#8206) - Add CutensorMg support (#8212)
- Add
cudaStreamCreateWithPriority(#8219) - Add the
nearestmethod for percentile/quantile estimation (#8224) - Various Jitify improvements (#8235)
- Support fallback algorithm for spgemm (#8252)
- Bump to cuTENSOR 2.0.1 (#8282)
- Preload cuTENSORMg (#8283)
- Use
weakref.finalizeinstead of__del__forRandomState._generatordestruction (#8315) - Support ROCm 6 (#8319)
- cupyx: cleanup use of deprecated NumPy functionality (NumPy 2.0 compatibility) (#8320)
- Add wright_bessel function to special (#8324)
- MAINT: fft, linalg: add
__all__lists (#8333) - Cuda 12.5 Tests (#8337)
- Add axes support in ndimage filters module (#8339)
- MAINT: interpolate: update RBF to scipy 1.13 (#8343)
- Make CuPy import under NumPy 2.0 (#8346)
- Lazy-preload NCCL (#8360)
- Fix
map_coordinatesrecompilation condition (#8378) - Disable jitify for cub & Bump CCCL (#8412)
- Use custom less instead of specializing thrust (#8446)
- Port to Cython 3.0 (#8457)
- Avoid using Jitify everywhere inside CuPy (#8467)
- Get rid of
pkg_resources(#8480) - Drop support for Python 3.9, NumPy 1.22 and 1.23, SciPy 1.7, 1.8 and 1.9 (#8491)
- Remove deprecated
cupyx.scipy.linalg.{tri,tril,triu}(#8499) - Use
.toarray()instead of.Aattribute (#8508) - Support
halfoption inscipy.signal.minimum_phase(#8510) - Increase
MAX_NDIMto 64 (#8511) - Support CUDA 12.6 (#8513)
- Fallback to system headers for future CUDA 12.x versions (#8518)
- Extend runtime header search logic to conda (#8519)
- Support
copy=Noneincp.array/cp.asarray/cp.asanyarray(#8545) - Fix dtype rule of
cupy.scipy.stats.entropyfor SciPy 1.14 (#8547) - Support setuptools 74.0.0 or later (#8583)
- Add
NCCL_ERROR_REMOTE_ERRORto the set of errors from NCCL (#8662) - Replace
numpy.ComplexWarningwithcupy.exceptions.ComplexWarning(#8676) - ENH: Implement dlpack v1 (#8683)
- Fix some NumPy 2.x CI failures (cont.) (#8695)
- Bump CUDA version in cuda11x-cuda-python CI (#8737)
- [ROCm 6.2.2] Conditionally define CUDA_SUCCESS only if it's not (#8793)
- Remove fallback mode (#8816)
- Raise user warning in both
{to,from}Dlpack& Update the Interoperability page (#8831) - Use a custom Min/Max instead of specializing CUB (#8846)
- Updating pylibraft
pairwise_distanceto cuvs (#8847) - add axes support for additional functions in cupyx.scipy.ndimage (from SciPy 1.15.0) (#8858)
- Raise VisibleDeprecationWarning for wavelet functions (#8865)
- Support CUDA 12.8 + Blackwell GPUs (sm_100, sm_120) (#8899)
- Bump library installers for CUDA 12.8 (#8914)
- Use CCCL 2.8.x branch + Use
CUPY_CACHE_KEYin hash keys (#8919) - Use NVIDIA CCCL 2.8 latest w/CUDA 12.3 fix (#8924)
- Use C++17 in JIT compile (#8940)
- Restore CUB histogram and bincount (#8950)
- Broaden usage of C++17 (#8952)
cupyx.scipy.distance: initialize output array with empty instead of zeros (#8971)cupyx.scipy.spatial.distance.cdistremove explicit zeroing of user-provided output array (#8988)- Fix rocThrust build for ROCm 6.3 (#9022)
- Allow discovering cuTENSOR using major version (#9030)
- Update cutensornet accelerator based on cuquantum-python 25.03 deprecation (#9045)
- Support FIPS enabled machines with MD5 hashing (#9053)
- Refactor hashing (#9057)
Enhancements for NumPy & SciPy compatibility:
- Fix
scp.signal.{medfilt,medfilt2d}to raise ValueError for complex64 inputs (#8059) - Deprecate
cupyx.scipywavelet functions (#8061) - Fix
csrmatrix.__pow__to raise ValueError for non-int other (#8063) - Fix
cupyx.scipy.special.betaincfor invalid inputs (#8065) scipy.special.{btdtr,btdtri}are deprecated since SciPy 1.12 (#8066)- Fix
boxcox_llffor SciPy 1.12 changes (#8095) - NEP50 (#8323)
- Resolve Ruff
NPYerrors - fix exception imports andasfarrayusage in test code (#8455) - Fix
sparse.linalgfunction signatures following SciPy 1.14 (#8526) - NumPy 2.0 compatibility: (partially) sync with NEP52 (#8531)
- Fix dtype rule of special functions for SciPy 1.14 (#8532)
- Fix
cupy.histogramarg order to match NumPy (v1.24+) (#8559) - Make
cupy.linalg.solvecompatible withnumpyv2 (#8629) - Silence
FutureWarningemitted whenrcondis missing (#8638) - Fix some NumPy 2.x CI failures (#8690)
- Support
kindarg. in sorting methods (#8708) - Fix
cupy.percentilefor NumPy 2.x (#8726) - Fix some NumPy 2.x CI failures (cupyx) (#8727)
- Skip some tests incompatible with NumPy 2.2 (#8817)
- Fix scipy.spmatrix.sign for complex dtype inputs (#8822)
- Fix return type of
cupy.wherefor scalar arguments for NumPy 2.0 (#8835) - Fix
cupyx.scipy.special.logsumexpfor NumPy 2.0 (#8836) - Fix
cupy.cov(#8839) - Fix
cupy.histogramddfor NumPy 2.x (#8873) - Raise ValueError upon attempts to create 3-dim sparse array (#8877)
- Disable contiguous_check for COO/dense matmul test (#8878)
- Skip a test for invalid scipy return value of invalid COO matmul (#8879)
- Support empty tuple indexing for sparse matrix (#8882)
- Fix
fft.fhtfollowing bug fix in SciPy 1.15 (#8883) - Deprecate
cupyx.scipy.linalg.kron(#8885) - Add
special.sph_harm_yand deprecatespecial.sph_harm(#8898) - Fix boxcox_llf for SciPy 1.13 (#8907)
- Update cupyx.scipy.special functions for SciPy 1.15 (#8908)
- Fix dtype rule of stats.entropy for SciPy 1.15 (#8911)
- Fix dtype rule of boxcox_llf for SciPy 1.15 (#8913)
- Fix
cupyx.scipy.stats.zscorefor SciPy 1.15 (#8977) - Update
cupy.testingfollowing numpy (#9076)
Performance Improvements
- Speed up cupy environment duplicate detection (#8017)
- Improve efficiency of binary morphology functions by accepting a tuple for
footprint(#8312) - Improving performance of
cupy.random.choice(#8483) - Improve performance of
cupyx.scipy.ndimage.binary_fill_holes(#8956)
Bug Fixes
- Fix argmax/argmin for large reduction axis (#8031)
- Fix
lfilter_ziandsosfilt_ziwhen any IIR coefficient is zero (#8034) - Fix
cupyx.scipy.fft.{dst,dstn}in type 2/3 (#8058) - Update
_nccl_comm.py(#8089) - Fix Flags not to allow setters (#8104)
- Do not use
from-import(#8109) - Fix overflow indexing ndarray generated with
as_strided(#8150) - Prevent angular brackets from appearing in Jitify's cache filename (#8154)
- Set
-archin the compiler options unconditionally (#8157) - Allow
cupy.show_config()without CUDA (#8183) - Fix jitify warmup kernel (#8200)
- Fix: always switch to the submodule dir before checking git tag/commit (#8201)
- Fix: remove unnecessary include that causes deployment issue (#8202)
- Fix CUB
min/maxinitial values (#8203) - Fix build system for Thrust detection (#8229)
- Fix overflow of index calculation in random generator API (#8243)
- Fix Generator API parallelism (#8244)
- Fix jitify warmup kernel - Cont'd (#8250)
- Fix bounds issue in Clough-Tocher computation (#8277)
hipPointerGetAttributesreturns error when pointer is unregistered in ROCm 5.7 (#8335)- Fix CUB build error on win-64 (#8345)
- Re-enable NVTX range coloring for NVTX3. (#8353)
- Delaunay: Prevent checking neighboring triangles if nTri <= 1 (#8363)
- Fix
ndarray.get()not honoring current stream when layout is not contiguous (#8369) - Fix spline temp container size in make_interp_spline (#8388)
- MAINT: Avoid using np.compat.integer_types (#8409)
- Fix type dispatcher for arm64 (#8410)
- Fix
RandomState.seed()for NumPy 2 compatibility (#8429) - Fix
copytofor NumPy 2 compatibility (#8430) - Update compiler.py to avoid the popup of the nvcc.exe console (#8432)
- Address KeyErrors from importlib_metadata (#8441)
- Fix the size of temporary CUB output space to consider its alignment (#8444)
upfirdn:mode=None->mode="constant"(#8476)- Search header files from CTK wheel (#8489)
- Fix CUDA version condition to use headers from wheel (#8505)
- Fix order 'K' with shape given for
*_likearray creation (#8525) - Fix ROCm 4.3 binary package build broken (#8528)
- Fix cudart header detection for conda (#8530)
- Properly allocate in RNG when specified dtype is neither float32/float64 (#8538)
- Add nccl.broadcast 64-bit support (#8562)
- Support building CuPy with setuptools 74 (#8574)
- Guard for ROCm 6.x (#8610)
- Fix
HIP_VERSIONunit (#8613) - Switch to using platform.machine() instead of platform.processor() (#8655)
- Use
platform.machine()instead ofplatform.processor()(#8657) - Fix
sosfiltstate output shape when ndim < 2 (#8677) - Fix undefined inf/nan constant in CuPy JIT (#8707)
- Fix import order (#8723)
- Fix
bsplinekernel to avoid out of bounds error (#8760) - Fix race during SoftLink initialization (#8774)
- Fix
nanargminand nanargmax's parameter order and pass optional parameters (#8783) - Fix crashes of quantile and percentile (#8808)
- Fix handling of pinned memory (#8845)
- ndimage: fix bug in morphology
axessupport whenstructure=None(#8954) - Fix buffer protocol to raise TypeError when it is not meant to be supported (#8957)
- Use
/bigobjon Windows build (#8962) - Fix
cupyx.scipy.spatial.distance'scdistfor RAPIDS 24.12 compatibility (#8970) - JIT: Support empty return (#8989)
- Cast points when computing Delaunay triangulation on dtype != float64 (#9021)
- Fix compilation error of cupy.inf in fusion2 (#9034)
- Include
pxdfiles incupyx(#9081)
Code Fixes
- Refactor radartools (#8072)
- Refactor convolve1d3o (#8078)
- MAINT: cupyx/interpolate: remove a duplicate function (#8170)
- Do not use plain thrust namespace in complex clone (#8221)
- Use f-strings in ndimage kernel generation code (#8311)
- Upgrade pre-commit hooks to silence warnings (#8664)
- Resolve import loop (#8705)
- Enhance comments in percentile implementation (#8753)
- Resolve uncaught type warning (#8797)
- Switch from
.Aattribute to.toarray()method (#8810) - Add
cupy/_core/numpy_allocator.hto.gitignore(#8813) - Drop unneeded
bytescopy ofCUPY_CACHE_KEY(#8923) - Fix typo in
_cretate_frame_tree(#8942) - Fix get_typename to emit thrust::complex (#9040)
Documentation
- Generate signature for
ufuncdocumentation (#6137) - Use modern DLPack interface in torch interoperability document (#7988)
- Add a note on accelerating CI/CD (#8004)
- Update conda installation guide (#8129)
- Add CuPy logo in SVG (#8182)
- Fix
pdistdocstring in order to specify that the returned matrix is condensed (#8186) - Add CuPy Training Materials information to README.md (#8225)
- Fixed typo in
cupy.linalg.normdocs (#8264) - Replace license notice in cupyx.scipy.signal._spectral (#8268)
- Update document for CUDA 12.3 and 12.4 (#8280)
- Expose Delaunay in documentation (#8294)
- Add comparison table for
(cupyx.)scipy.sparse.*_matrix classesclass methods (#8295) - Find and fix typos with codespell (#8304)
- Update the version of cusparselt in the installation guide (#8362)
eigshdoc correction in_eigen.py(#8365)- Add NumPy 2.0 on document (#8368)
- Update API baseline on document (#8370)
- Add CUDA 12.5 to list of supported platform (#8424)
- typo: coping -> copying (#8426)
- Add docs about CUDA headers (#8593)
- Update badges in README (#8603)
- docs: update fft.rst (#8616)
- Update documentation to use
pre-commit(#8640) - Add tips on Windows development in Contribution Guide (#8702)
- Add notice about
cupy.array_apiremoval (#8750) - Remove "Type promotion" section from user guide (#8754)
- Add CUDA 12.8 to docs (#8961)
- Update list of supported versions (#8986)
Installation
- Patch the build system to better support conda-build (#7603)
- Skip
CUDA_PATHwarning in Conda installation (#8051) - Do not search for static libs (#8134)
- Bump version in main branch for v14 development (#8341)
- Update conda-build CUDA detection logic for Setuptools 72.2.0 (#8544)
- Remove unnecessary "install" package in Docker image (#8549)
- Use relative path of header files to generate cache key (#8927)
- Fix minimum CUDA version check and update comments (#8928)
Tests
- Revert CI timeout changes (#7903)
- Add import test without CUDA Toolkit (#7996)
- Bump stable branch to v13 (#8024)
- Update Compatibility Matrix (#8025)
- Remove some
signal.vectorstrengthxfail tests (#8060) - Fix scipy.linalg not to raise DeprecationWarning for zero-size inputs (#8064)
- Refactor radartools tests (#8071)
- Fix
scipy>=1.12filter condition (#8088) - Fix slow test (#8115)
- Support SciPy 1.12 (#8127)
- Fix invalid
vectorstengthtests (#8142) - Fix actions versions used in workflows to avoid node 16 deprecation warning (#8189)
- Add CI to test
cupy.show_config()pass without CUDA installed (#8190) - BUG: cupyx/scipy/signal: fix
mpmathtest (#8259) - Add support for CUDA 12.3 & 12.4 (#8265)
- Tentatively pin SciPy to v1.12 in CI (#8273)
- ndimage: more efficient testing of binary morphology functions (#8313)
- Skip some tests in aarch64 CI (#8411)
- Bump NumPy/SciPy versions in cuda-example CI (#8415)
- Fix CUDA 11.2 CI failure on Linux (#8436)
- Decrease number of threads to avoid "system error: excessive memory usage is detected" (#8461)
- CI: skip CUDA 12.1/12.2/12.3/12.4 CI on "mini" trigger (#8463)
- CI: Update micro versions of Python (#8493)
- Support SciPy 1.13 and 1.14 (#8500)
- Relax
test_firlsatol (#8509) - Fix
cupyx.scipy.sparsetests for SciPy 1.14 (#8527) - Skip
special.logsumexptest for empty input (#8533) - Skip
betaincinvtest with SciPy 1.14.1 (#8546) - Revert CI timeout bump (#8564)
- Add NumPy 2.x CI for Linux (#8565)
- Use
setuptools==73.0.1(#8567) - Fix FFT tests for NumPy 2.0 (#8584)
- Add self-hosted CI (#8612)
- Add CI for ROCm 6.2 (#8624)
- Update precommit (#8628)
- Renegerate rocm-6-2.Dockerfile (#8632)
- Skip tests if
scipyis not installed (#8634) - Use
with_requiresrather thanxfail(#8639) - Accept
OverflowErrorinTestCopytoFromScalarfor NumPy v2 (#8641) - Skip more tests if scipy is not installed (#8642)
- Replace
flake8withruff(#8671) - CI: Fix apt repository URL for Ubuntu 22.04 (#8713)
- Relax tolerance of
test_hilbertfor NumPy 2.0 (#8740) - Temporary skip for NumPy 2.0 tests (#8741)
- Remove ndarray.ptp from fallback tests (#8742)
- Bump SciPy version to 1.14 in Windows CI (#8759)
- Add NumPy 2.x CI for Windows (#8766)
- Fix some NumPy 2.x CI failures (cont.) (#8790)
- Set
numpy._set_promotion_state("weak")only for test with NumPy 1.x (#8809) - Add NumPy 2.2 to CI (#8815)
- CI: support "skip-ci" label (#8832)
- CI: Fix FlexCI compatibility (#8840)
- Support SciPy 1.15 (#8861)
- Support Optuna 4 (#8862)
- Fix test for csr_matrix.setdiag (#8874)
- Skip some signal tests for TypeError for inputs of
np.longlongdtype (#8884) - Disable
contiguous_checkfor somesignal.cont2discretetests (#8886) - Fix
test_interpndfor removed interpolate.interpnd namespace (#8887) - Add
testing.shaped_linspace(#8896) - Add CI for CUDA 12.8 (#8905)
- Fix splines tests to remove unexpected skips (#8918)
- Minor updates for sm120 (#8920)
- Add CI for Python 3.13 and mpi4py v4 (#8953)
- Increase host memory in Windows CI, free GPU memory in example code (#8966)
- Pass
localsdict toexec(#8982) - Mark xfails in some spline tests for SciPy 1.15 (#8984)
- CI: Do not run full CI on CUDA 12.0/12.1/12.2 + Windows (#8992)
- TST: Add test for fusion path with constant and scalar (#9013)
- CI: Pin setuptools version on Windows (#9038)
- Revert "CI: Pin setuptools version on Windows" (#9050)
- Add SciPy 1.15 CI for Windows (#9066)
Others
- Wait before published items shown on PyPI (#8302)
- Fix pull request project board workflows (#8366)
- Bump version to v14.0.0a1 (#8372)
- Add backport reminder (#8675)
- Fix script name of backport reminder (#8685)
- Add
workflow_dispatchevent to allow debugging workflow (#8687) - Update bug issue template (#8749)
- Update
pre-commithooks (#8838) - Allow specifying no libraries when generating wheel metadata (#9078)
👥 Contributors
The CuPy Team would like to thank all those who contributed to this release!
@99991 @acc-mu3n @andfoy @arkdong @asi1024 @Azusachan @bernhardmgruber @Berrysoft @bmerry @boku13 @cclauss @cjnolet @dagardner-nv @dakofler @EarlMilktea @eltociear @emcastillo @ev-br @gdaisukesuzuki @grlee77 @hauntsaninja @hmaarrfk @HollowMan6 @jakirkham @jemiryguo @johnnynunez @kmaehashi @leofang @littlewu2508 @lowener @martinResearch @MattTheCuber @miscco @mohitreddy1996 @monzelr @mroeschke @romerojosh @rongou @seberg @so298 @steppi @swelborn @syheliel @takagi @take-cheeze @tornikeo @xuefeng-xu @yangcal @YanivDorGalron @Zhyrek