This is the release note of v9.0.0.
This release note only covers the changes since v9.0.0rc1 release. Read the blog for the details of new features introduced in CuPy v9!
We are running a Gitter chat for general discussions and quick questions. Feel free to join the channel to talk with developers and users!
Highlights
NVIDIA cuSPARSELt
CuPy now integrates the Python binding for the cuSPARSELt library that accelerates sparse matrix multiplications on NVIDIA Ampere GPUs. We are planning to start using it in CuPy sparse APIs to transparently improve performance.
RAPIDS cuGraph
cupyx.scipy.sparse.csgraph is added to the API with support for the connected_components method. The support for cuGraph is optional and can be installed through conda-forge or by manually building CuPy. Currently, PyPI wheels do not have built-in support for cuGraph.
Add MemoryAsyncPool to support malloc_async (#5034)
By using cupy.cuda.set_allocator(cupy.cuda.MemoryAsyncPool().malloc) it is now possible to use the stream ordered memory allocations introduced in CUDA 11.2.
APIs for creating NumPy arrays backed by pinned memory (#5100)
By using the cupyx.empty_pinned(), cupyx.empty_like_pinned(), cupyx.zeros_pinned() cupyx.zeros_like_pinned() it is possible to obtain NumPy ndarrays with their storage located in pinned memory to improve performance of data movement.
CUDA 11.0 and 11.1 wheels not available yet in PyPI (#4971)
In the meantime, they can be downloaded from the Assets section below. See #4971 for the detailed instructions.
Changes
See here for the complete list of solved issues and merged PRs after v9.0.0rc1 release. For all changes since v9 series, please refer to the release notes of the pre-releases ((alpha1, beta1, beta2, beta3, rc1).
New Features
- Support shared memory in CuPy JIT (#4977)
- Support cuSPARSELt (#4994)
- Add
randomfor uniform [0, 1) generation (#5003) - CUDA 11.2: Add
MemoryAsyncPoolto supportmalloc_async(#5034) - Add poisson distribution to random API (#5036)
- CuPy JIT: Print kernel code (#5038)
- Add gamma distributions to random API (#5086)
- Add APIs for creating NumPy arrays backed by pinned memory (#5100)
- Add SciPy compatible
connected_components(#5113)
Enhancements
- Disable CUB SpMV on CUDA 11.x (#4978)
- Move the NVTX module to
cupy_backends.cuda.libs(#5014) - HIP: add
-ftz=true(#5035) - CuPy JIT: Readable compile error messages (#5041)
- CuPy JIT: Use C++-like typing rule in 'cuda' mode (#5053)
- Mark
cupyx.jit.rawkernelas experimental (#5057) - Add PCI Bus ID to show_config (#5062)
- Print cuSPARSELt version in
show_config(#5065) - Give gufunc a name (#5085)
Bug Fixes
- Use THRUST_OPTIONAL_CPP11_CONSTEXPR (#5011)
- Disable cuFFT plan cache on CUDA 11.1 (#5068)
- Use async memcpy in
ndarray.copy(#5078) - CuPy JIT: Fix range type (#5081)
- Support PTDS in CuPy memory pool (#5082)
- Adjust PATH when preloading to load cuDNN v8 correctly on Windows (#5116)
Code Fixes
- Rename
cupy.coresubmodule tocupy._core(#4987) - Fix some internal
cpdeffunctions tocdefin_kernel.pyx(#5098)
Documentation
- Fix docs: cupy-cuda112 now on PyPI (#4990)
- Update installation guide for Conda-Forge (#4993)
- Document
cupyx.time.repeat(#5027) - Document
cupy.cuda.runtime.getDeviceProperties(#5029) - Doc: Add links to Anaconda, Gitter, StackOverflow (#5030)
- More documentation on the supported backends (#5039)
- Fix code block in installation guide (#5043)
- Document
CFunctionAllocatorandManagedMemory(#5059) - Improve the documentation on interoperability (#5064)
- CuPy JIT documentation (#5076)
- Improve comments for memory and stream API usage (#5079)
- Add user guide (#5109)
- Reorganize API reference pages (#5114)
- Point to the correct numpy random docs (#5115)
- Follow the latest NumPy/SciPy docs style (#5118)
- Add ROCm limitations to docs (#5119)
- Revise ROCm doc (#5123)
Installation
- Fix Windows dll loading for Conda (#5106)
Examples
- Update examples for current version of CuPy (#5009)
- Fix cuSPARSELt example not to use internal function (#5066)
Tests
- Tentatively pin CI to ROCm 4.0.1 (#4976)
- Update known base branches in flexCI config (#4980)
- Fix
cutensorimport in the test (#4981) - Update list of known branches (#4989)
- Make install_tests runnable without depending on current path (#4992)
- Fix
TestStreamcleanup (#5052) - Mark some memory tests as
testing.slow(#5063) - Refactor random tests (#5102)
Others
- Use bot mode in automatic backport (#5058)
Contributors
The CuPy Team would like to thank all those who contributed to this release!