github ggml-org/llama.cpp b11513

pre-releaseone hour ago
Details

CUDA: improve top-k algorithm selection (#28713)

  • CUDA: radix top-k for large row counts

Replaces CUB's per-row DeviceTopKKernel with a grid-over-rows radix select,
gated on GGML_CUDA_TOPK_RADIX_MIN_ROWS. On qwen4exp at 34,816 tokens this cuts
top-k from 1,671,253 launches / 5,761.8 ms to 2,329 / 941.8 ms.

  • CUDA: select the TOP_K implementation by shape

Replace the nrows/ncols special case with the decision boundary from #28547
(as implemented in #29278): bitonic for short rows, radix select for several
long rows, and DeviceTopK or CUB argsort for a single long row. The
thresholds stay overridable at build time.

Two refinements on top of that boundary:

  • bitonic stays in use for rows up to a padded 1024 while the rows fit in one
    wave of blocks (nrows <= number of SMs); radix select pays a fixed cost of
    about a dozen launches that only amortizes over more rows
  • with DeviceTopK available, it handles up to two rows

Radix select now processes rows in chunks so its scratch memory stays bounded,
and the bitonic path keeps its chunking. HIP and MUSA keep their previous
thresholds.

Add perf cases around the bitonic/radix crossover to test-backend-ops.

  • CUDA: make top-k comments less verbose

  • CUDA: remove the TOP_K width limit from supports_op

  • CUDA: use DeviceTopK for single-row TOP_K if available

  • CUDA: avoid ncols overflow in the TOP_K bitonic check

  • CUDA: share the row chunking helper between argsort and top-k

  • CUDA: do the TOP_K radix blocks_per_row math in int64_t

  • CUDA: rename GGML_CUDA_TOP_K_NROWS_THRESHOLD_DEVICETOPK to GGML_CUDA_TOP_K_NROWS_THRESHOLD

  • CUDA: share one sort helper between the bitonic and CUB TOP_K paths

  • CUDA: update the TOP_K TODO, threshold and chunking comments

  • tests: add TOP_K cases that span several row chunks

  • CUDA: use int64_t col in the TOP_K radix loops, fix threshold comment

  • CUDA: limit TOP_K and ARGSORT support to ne[0] <= INT_MAX


Co-authored-by: praneshgo 227579474+praneshgo@users.noreply.github.com
Co-authored-by: Pranesh Gonegandla pgonegandla@nvidia.com

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.