This is a minor release. See https://github.com/cupy/cupy/milestone/8?closed=1 for the complete list of solved issues and merged PRs.
New features
Sparse matrix
cupy.sparse is a module that implements scipy.sparse API using CUDA and cuSPARSE. We now have basic features for using sparse matrices on GPU.
- CSR and CSC (#226)
- COO matrix (#234)
- Conversion method from compressed matrix (csr, csc) to coordinate format (coo) (#235)
- CSR and CSC copy (#236)
__add__,__radd__,__sub__and__rsub__for CSR and CSC (#238)- Fix
toarrayincupy.sparse.spmatrix(#312) - Return
NotImplementedinstead ofNotImplementedError(#330) - Use
csc2denseto convert csr-matrix to dense (#305)
We are planning to add more features to cupy.sparse in upcoming releases.
New memory allocator (#168)
The memory pool implementation is greatly updated. It is based on best-fit allocation with coalescing. When there are a large number of allocations with different sizes (e.g. NLP applications), the memory usage is improved and the number of re-allocations is reduced (which also reduces the running time).
For example, the memory usage of the sequence-to-sequence code using Chainer (chainer/chainer#2070) is reduced from 12GiB (which means the process is using all of the available GPU memory) to 3GiB, and the number of memory reallocations from 20 times to 0 times.
It may increase the memory usage in some cases, although the amount of additional usage is small in practice (see the benchmark results in #168).
You can use this memory allocator by calling cupy.cuda.set_allocator(cupy.cuda.MemoryPool().malloc) (when using Chainer, it is called by default).
Other features
- Implement
cupy.linalg.det(#96) - Support
cupy.sortto sort arrays along arbitrary axis (#229) - Implemented
RangeStartandRangeEndfor NVIDIA visual profiler (nvvp) (#246) - Introduce
cupy.is_available()which takes account of device availability (#247) - Implement
cupy.msort(#251, #329)
Bug fixes
- Fix
cupy.copytofunction to treat multiple GPUs correctly (#220) - Restore kernel type check (#253)
- Fix
deepcopywith multiple devices (#254) - Fix
cupy.argsortfor non-contiguous arrays (#284) - Fix
ldexpon Windows (#278)
Improvements
- Improve
cupy.argsortperformance (#285)
Installation
- Remove old cuDNN support (#219)
- Add compile options to build on Windows (#244)
- Remove duplicated build options (#280)
- Avoid creating garbage file on setup (#287)
- Fix setup for cusolver (#292)
- Use
cupy.cuda.thrust_enabledto check Thrust enabled (#224)
Documentation
- Updated difference with NumPy on reduction function behavior (#144)
- Fix spelling in tutorial (#268)
- Fix test instruction in README (#310)
- Fix links to GitHub source pages (#332)
Examples
- Add Gaussian Mixture Model (GMM) example (#29, thanks @KotaroSetoyama!)
- Make grid size to integer for SGEMM example (#289, thanks @yuyu2172!)
- Use absolute path in SGEMM example (#291)
- Updated README for SGEMM example (#245, thanks @yuyu2172!)
Tests
- Use
cupy.testing.for_all_dtypes(#269) - Enable style check for Python code in Travis (#273)
- Refactor
cupy.argsorttests (#282)
Others
- Small fixes for
cupy.argsort(#223)