This is the release note of v11.0.0b1. See here for the complete list of solved issues and merged PRs.
We are running a Gitter chat for general discussions and quick questions. Feel free to join the channel to talk with developers and users!
Notice (2022-04-05)
We have identified that this release contains a regression that prevents CuPy from working in older CUDA GPUs (Maxwell or earlier). We are planning to fix this issue in the next pre-release. See #6615 for the details.
Highlights
Increase coverage of cupyx.scipy.special APIs (#6461, #6582, #6571)
A series of scipy.special routines have been added to cupyx with optimized CUDA raw kernel implementations. loggamma, multigammaln, fast Hankel transformations and several other utility special functions are added in these series of PRs by @grlee77 and @khushi-411.
Support for CUDA 11.6
Full support for CUDA 11.6 has been added as of this release. Binary packages can be installed with the following commnad: pip install --pre cupy-cuda116 -f https://pip.cupy.dev/pre
Support for ROCm 5.0
Full support for ROCm 5.0 has been added as of this release. Binary packages can be installed with the following commnad: pip install --pre cupy-rocm-5-0 -f https://pip.cupy.dev/pre
Changes without compatibility
Use CUB by default (#6549)
CUB support in CuPy is now enabled by default. This results in faster general reductions and routines such as sum, argmax, argmin having increased performance. Notice that CUB may introduce some non-deterministic behavior and this can be disabled by setting the CUPY_ACCELERATORS="" environment variable.
Drop support for ROCm 4.0 (#6420)
CuPy v11 will drop support for ROCm 4.0. We recommend users to use ROCm 4.3 or 5.0 instead.
Changes
New Features
- Add
cupyx.scipy.specialstatistical distributions (#6461) - Add
cupy.real_if_closeAPI (#6475) - Add
cupyx.scipy.specialloggamma, multigammaln and fast Hankel transforms (#6528) - Add
cupyx.scipy.special.{i0e, i1e}(#6571)
Enhancements
- Update
cupy.array_api(#6486) - Fix for supporting ROCm 5.0 (#6524)
- Use CUB by default (#6549)
- Fix
cupy.copytoto take NumPy array scalars (#6584) - Implement
ndarray.ravel(order="K")(#6585) - Make einsum accept subscripts in numpy int (#6506)
Performance Improvements
- Support
cusparseSpGEMM()(#6511) - eigsh: Prefer gemv over gemm (#6570)
- Performance improvement of
cupy.in1d(#6583)
Bug Fixes
- Fix
cupy.fillto properly take zero-dimcupy.ndarray(#6481) - Fix error message in
vectorize(#6499) - Fix
cupy.cumsumon ROCm 5.0 (#6520) - Fix coo_matrix.diagonal (#6522)
- Fix array creation shape (#6545)
- Fix
outargs parser of ufunc (#6546) - Fix
may_share_memoryalgorithm (#6560) - Avoid using the same kernel from different devices in JIT (#6575)
- Fix cupy.full and cupy.full_like to make unsafe casting (#6587)
- Fix device context management in
MemoryAsyncPool(#6590)
Code Fixes
Documentation
- Fix documents for CUDA 11.6 (#6405)
- Remove description about issues from contribution guide (#6497)
- Documentation update for ROCm 5.0 (#6530)
Installation
- Skip appending
--compiler-bindirifcl.exeis already onPATH(#6510) - Bump version to v11.0.0b1 (#6601)
Tests
- Add FlexCI projects for Windows (#5889)
- Run cupy-benchmark on CI (#6417)
- Disable CentOS 8 test (#6492)
- Fix Dockerfile broken for array-api tests (#6508)
- CI: Trigger
pushevent of FlexCI via GitHub Actions (#6538) - Skip
async_malloctests on unsupported device (#6541) - Fix flaky test_inverse_indices_shape (#6551)
- Trigger CUDA 11.6 Windows CI when push/pull-request (#6553)
- CI: Fix event name in dispatcher (#6555)
- CI: Fix rule name in dispatcher (#6556)
Contributors
The CuPy Team would like to thank all those who contributed to this release!
@anaruse @asi1024 @emcastillo @grlee77 @khushi-411 @kmaehashi @leofang @Onkar627 @peterbell10 @pri1311 @Smit-create @takagi @toslunar @tushxr16