pypi cupy 2.0.0b1
v2.0.0b1

latest releases: 14.2.0, 14.1.1, 14.1.0...
9 years ago

This is a minor release. See https://github.com/cupy/cupy/milestone/8?closed=1 for the complete list of solved issues and merged PRs.

New features

Sparse matrix

cupy.sparse is a module that implements scipy.sparse API using CUDA and cuSPARSE. We now have basic features for using sparse matrices on GPU.

  • CSR and CSC (#226)
  • COO matrix (#234)
  • Conversion method from compressed matrix (csr, csc) to coordinate format (coo) (#235)
  • CSR and CSC copy (#236)
  • __add__, __radd__, __sub__ and __rsub__ for CSR and CSC (#238)
  • Fix toarray in cupy.sparse.spmatrix (#312)
  • Return NotImplemented instead of NotImplementedError (#330)
  • Use csc2dense to convert csr-matrix to dense (#305)

We are planning to add more features to cupy.sparse in upcoming releases.

New memory allocator (#168)

The memory pool implementation is greatly updated. It is based on best-fit allocation with coalescing. When there are a large number of allocations with different sizes (e.g. NLP applications), the memory usage is improved and the number of re-allocations is reduced (which also reduces the running time).

For example, the memory usage of the sequence-to-sequence code using Chainer (chainer/chainer#2070) is reduced from 12GiB (which means the process is using all of the available GPU memory) to 3GiB, and the number of memory reallocations from 20 times to 0 times.

It may increase the memory usage in some cases, although the amount of additional usage is small in practice (see the benchmark results in #168).

You can use this memory allocator by calling cupy.cuda.set_allocator(cupy.cuda.MemoryPool().malloc) (when using Chainer, it is called by default).

Other features

  • Implement cupy.linalg.det (#96)
  • Support cupy.sort to sort arrays along arbitrary axis (#229)
  • Implemented RangeStart and RangeEnd for NVIDIA visual profiler (nvvp) (#246)
  • Introduce cupy.is_available() which takes account of device availability (#247)
  • Implement cupy.msort (#251, #329)

Bug fixes

  • Fix cupy.copyto function to treat multiple GPUs correctly (#220)
  • Restore kernel type check (#253)
  • Fix deepcopy with multiple devices (#254)
  • Fix cupy.argsort for non-contiguous arrays (#284)
  • Fix ldexp on Windows (#278)

Improvements

  • Improve cupy.argsort performance (#285)

Installation

  • Remove old cuDNN support (#219)
  • Add compile options to build on Windows (#244)
  • Remove duplicated build options (#280)
  • Avoid creating garbage file on setup (#287)
  • Fix setup for cusolver (#292)
  • Use cupy.cuda.thrust_enabled to check Thrust enabled (#224)

Documentation

  • Updated difference with NumPy on reduction function behavior (#144)
  • Fix spelling in tutorial (#268)
  • Fix test instruction in README (#310)
  • Fix links to GitHub source pages (#332)

Examples

  • Add Gaussian Mixture Model (GMM) example (#29, thanks @KotaroSetoyama!)
  • Make grid size to integer for SGEMM example (#289, thanks @yuyu2172!)
  • Use absolute path in SGEMM example (#291)
  • Updated README for SGEMM example (#245, thanks @yuyu2172!)

Tests

  • Use cupy.testing.for_all_dtypes (#269)
  • Enable style check for Python code in Travis (#273)
  • Refactor cupy.argsort tests (#282)

Others

  • Small fixes for cupy.argsort (#223)

Don't miss a new cupy release

NewReleases is sending notifications on new releases.