This is the release note of v9.0.0a1. See here for the complete list of solved issues and merged PRs.
Highlights
CUDA 11.1 Support
Support for CUDA 11.1 is added in #4184, with CUDA 11.1, GeForce RTX 30 series and Quadro RTX series can now be used in CuPy.
Notes on Wheel Packages
Update (2020-11-25): cupy-cuda111 is now available on PyPI.
CuPy for CUDA 11.1 (
cupy-cuda111) wheel packages are currently only available for Windows. We are going to publish Linux wheels once we get approval from the PyPI team. Meanwhile, Linux wheels can be downloaded from the Assets section below (or pip install cupy-cuda111 -f https://github.com/cupy/cupy/releases/tag/v9.0.0rc1).
New Features
- Add compressed sparse
__setitem__(#3533) - Add
cupy.polyfit(#3747) - Support sparse pointwise division by vectors or matrices (#3838)
- Add
cudaGetDeviceProperties(#3858) - Support sparse pointwise maximum and minimum (#3860)
- Add all binary morphology functions to
cupyx.scipy.ndimage(#3907) - Support
cublasXgetrsBatchedand addcupy.cublas.batched_gesv(#3936) - Add
cupy.testing.shaped_sparse_random(#3944) - Add sparse pointwise equality & inequality functions (#3945)
- Add remaining grayscale morphology operations to
cupyx.scipy.ndimage(#3946) - Add
histogram2dandhistogramdd(#3947) - Add
cupy.gradient(#3963) - Add several functions to
cupyx.scipy.ndimage.measurements(#3979) - Add
cupyx.scipy.linalg.lu(#3995) - Add
cupy.apply_along_axis(#4008) - Add
cupyx.scipy.sparse.linalg.norm(#4017) - Add missing sparse matrix constructors (#4052)
- Add
cupy.cusolver.gels(#4064) - Add
@operator support tocupyx.scipy.sparse(#4075) - Add
cupy.nancumsumandcupy.nancumprod(#4077) - Add
orderoption incupy.testing.shaped_random(#4091) - Add
cupy.nanmedian(#4092) - Add complex dtype support in
cupy.nanminandcupy.nanmax(#4097) - Add
cupy.appendandcupy.resize(#4112) - Add
cupyx.scipy.sparse.linalg.eigsh(#4138) - Add support for CUDA 11.1 (#4184)
Enhancements
- Support list bins with
histogram(#3542) - Add a cuFFT plan cache (#3730)
- Support transforming NumPy arrays with multi-GPU
Plan1d(#3766) - Show numpy and scipy versions in
show_config(#3768) - Add cuTENSOR 1.2 support (#3884)
- Update FP16 header to CUDA 11.0 Update 1 (11.0.3) (#3888)
- Check format of sparse matrix in
numpy_cupy_array_equal(#3897) - Improve accuracy of
cupy.around(#3904) - Bump cuDNN version to v8.0.3 (#3985)
- Add complex dtype support to
cupyx.scipy.linalg.lu_factor/solve(#4002) - Add cython bindings to cuSPARSE
csrsv2/csrsm2related functions (#4031) - Support pickling
cupy.RawKernel(#4055) - Allow non-contiguous array input to binary morphology functions (#4058)
- Improve performance of binary morphology for fully nonzero structuring elements (#4059)
- Bump cuDNN to v8.0.4 (#4065)
- Add
*svdjBatchedprototypes (#4071) - Defer import in
cupy/_environment.py(#4162) - Record Cython build and runtime versions (#4164)
Performance Improvements
- Use cuTENSOR in
cupy.prod,cupy.max,cupy.min,cupy.ptpandcupy.mean(#3765) - Use
_csr_row_indexfor CSR matrix major-axis slicing with step (#3852) - Improve CSR matrix column fancy indexing (#3886)
- Use LU-decomposition based solver in
cupy.linalg.solver(#3942) - Improve
cupyx.scipy.sparseint x int indexing (#3981) - Avoid using
CUlinkStateunless absolutely necessary (#3992) - Improve
cupy.in1d(#4018) - Improve
cupy.cuda.cub.device_segmented_reduce()(#4161)
Bug Fixes
- Fix cooperative kernel launch (#3894)
- Fix dtype in CSR matrix division (#3905)
- Fix
csr2cscfor zero-size matrix (#3919) - Handle transfer to cupy view (#3928)
- Fix
_compressed_sparse_matrix._minor_slicefor step > 1 case (#3948) - Fix
csr_matrix._get_intXslicefor step < 0 case (#3951) - Fix
sparse.__getitem__not to return view of input (#3975) - ROCm: fix rocBLAS and rocSOLVER version displays (#3988)
- Add a kernel for integer GEMM (#3994)
- Fix typos in
cupy.cuda.cufft(#4014) - Fix managed memory leak (#4015)
- Fix potential segfault when reduction axis is empty (#4024)
- Use
__dealloc__instead of__del__for cdef class (#4036) - Fix typo in
_binary_erosion(#4038) - Fix CUB block reduction for F-order arrays with ndim > 2 (#4062)
- Add work-around for issue in
cutensorReductionof cuTENSOR 1.2.1 (#4081) - Handle
np.nanandnp.infconstant values properly in ndimage functions (#4083) - Fix
argmaxandargminfor F-order inputs (#4084) - Workaround
cudaPointerGetAttributeserror in CUDA 10.2+ (#4085) - Fix
argmax/argminin CUB block reduction for F-order arrays with ndim > 1 (#4096) - Fix
getDevicePropertiesfor HIP (#4108) - Add compute capability checking for
cublasGemmEx()(#4114) - Fix 64-bit int types in
type_dispatcher.cuh(#4124) - Fix mode='opencv' case in cupyx.scipy.ndimage.affine_transform (#4130)
- Add
compute_35for CUDA 11.0+ (#4137) - Fix device properties for cuda 9.2 (#4142)
- Fix
cupyx.seterr()whenlinalgnot supplied (#4150) - Fix broadcasting behavior in
ndimage.measurementsfunctions (#4151) - Fix
argwherefor 0d inputs (#4167) - Fix
nonzerofor 0d inputs (#4168) - Fix to use current stream properly with CUDA-related libraries (#4173)
Code Fixes
- Split cupy cuda header (#3616)
- Rename
cupy.iosubmodule tocupy._io(#3712) - Rename
cupy.logicsubmodule tocupy._logic(#3715) - Rename
cupy.manipulationsubmodule tocupy._manipulation(#3716) - Rename
cupy.mathsubmodule tocupy._math(#3717) - Rename submodules under
cupy.linalgpackage (#3741) - Rename
cupy.statisticssubmodule tocupy._statistics(#3774) - Rename
cupy.utilsubmodule tocupy._util(#3779) - Rename submodules under
cupyx.linalgpackage (#3784) - Refactor CSR sparse matrix row fancy indexing (#3865)
- Rename submodule under
cupy.profpackage (#3869) - Rename submodule under
cupy.fftpackage (#3870) - Hide private names in
cupy/__init__.py(#3871) - Rename
cupyx.rsqrtsubmodule (#3873) - Rename
cupyx.runtimesubmodule (#3874) - Rename
cupyx.scattersubmodule (#3875) - Rename submodule under
cupyx.scipy.fft(#3899) - Rename submodule under
cupyx.scipy.fftpack(#3900) - Rename submodules under
cupyx.scipy.sparse(#3901) - Rename submodules under
cupyx.scipy.special(#3902) - Hide private names in
cupyx/scipy/__init__.py(#3912) - Hide private names in
cupyx.time(#3965) - Hide private names in
cupy.cudnn(#3966) - Hide private names in
cupy.cusolver(#3967) - Hide private names in
cupy.cusparse(#3968) - Hide private names in
cupy.cutensor(#3969) - Move
_normalize_axis_indextocupy/core/internal.pyx(#4057) - Move
matmulfromcore.pyxto_routine_linalg.pyx(#4060)
Documentation
- Fix wrong curand enum names (#3840)
- Add
cupy.searchsortedto doc (#3908) - Update
cupyx.scipyAPI documentation (#3954) - Fix docs of cupyx.scipy.linalg.lu_factor (#4011)
- Improve the plan cache documentation (#4013)
- Update README and docs for unified tagline (#4047)
- Simplify ROCm install guide (#4048)
- Fix typo (#4053)
- Add note about starting nvprof with profiling off (#4144)
- Fix docstrings of
cupyx.scipy.ndimage.{minimum,maximum}_position(#4146)
Installation
- Add
CUDA_VERSIONdefine for Cython compilation (#3877)
Tests
- Code fix on tests for
cupyx.scipy.ndiamgestats functions (#3426) - Add different dtype input test in histogram (#3618)
- Fix 32-bit boundary test to run on Windows (#3859)
- Fix
cupy.ndimtest style (#3890) - Fix test fail when cudnn is unavailable (#3906)
- Add v8 to list of known branch in FlexCI script (#3911)
- Fix side effects in some tests (#3934)
- Fix some test to check compatibility with scipy's behavior (#3955)
- Refactor sparse indexing tests (#3958)
- Require SciPy 1.2 for sparse comparison (#4033)
- Add
generate_matrixtocupy.testing(#4070) - Make parameterized dtype test skip by
pytest.skip(#4094) - ROCm gpg url changed (#4127)
- Fix tests that have side effects (#4149)
- Enhance dtype error message in testing helpers (#4156)
- Fix
polyfittests tolerance (#4159) - Use
testing.assert_warns(#4169)
HIP/ROCm
- ROCm: Fix bugs and test suites to make ROCm/HIP happy - Part 1 (#3823)
- ROCm: Fix bugs and test suites to make ROCm/HIP happy - Part 2 (#3835)
- ROCm: Support rocTX (#3843)
- ROCm: Support rocFFT/hipFFT (#3896)
- ROCm: Support more hipBLAS/rocBLAS and rocSOLVER functions (#3950)
- ROCm: Support hipCUB/rocPRIM (#4027)
- ROCm: Support RCCL (#4099)
- ROCm: Build on latest ROCm (#4110)
Others
Contributors
The CuPy Team would like to thank all those who contributed to this release!
@anaruse @carterbox @cjnolet @Dahlia-Chehata @garanews @grlee77 @kalvdans @leofang @mrkwjc @saswatpp