github ggml-org/llama.cpp b10534

latest releases: b10539, b10538, b10537...
one hour ago
Details

CUDA: adding switch points per HW and quant type to tune the mvq->MMQ decode crossover (#26079)

  • CUDA: runtime GGML_CUDA_MMVQ_MAX to tune the mvq->MMQ decode crossover

Add a runtime override of the mul_mat_vec_q -> MMQ batch crossover
(default MMVQ_MAX_BATCH_SIZE). Lowering it routes batches above the
threshold from the CUDA-core vector kernel to the int8 MMQ tensor-core
path, which is faster once quantized decode becomes compute-bound at
B>1 (measured +23-41% at B=8 on RTX 5090 for Q4_K dense, no low-batch loss).

The value is parsed once and clamped to [1, MMVQ_MAX_BATCH_SIZE], since
mul_mat_vec_q asserts ncols_dst <= that; invalid input warns and falls
back to the default. The override is applied consistently in both the
mul_mat_vec_q and MUL_MAT_ID dispatch paths. Default behavior unchanged.

  • Added Blackwell specific switch point, to reduce dependence on runtime env var.

  • Add per-HW switch point values for DGX Spark and removing runtime env var

  • Adding switch points for Ada, tested on RTX 4090

  • Modifying DGX Spark numbers based on latest run and adding some comments and small functional changes relating to MoE

  • Reverting an unnecessary conditional

  • Update ggml/src/ggml-cuda/mmvq.cu


Co-authored-by: praneshgo 227579474+praneshgo@users.noreply.github.com
Co-authored-by: Oliver Simons osimons@nvidia.com

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.