github ggml-org/llama.cpp b11282

pre-releaseone hour ago
Details

musa : define CUDA_ARCH for device passes (#29508)

The MUSA vendor header never defined CUDA_ARCH, so every architecture
test in the shared ggml-cuda sources evaluated to 0. Kernel bodies gated on
the architecture therefore compiled to nothing, for example the q8_0 -> f16
dequantization kernel in convert.cu, whose NO_DEVICE_CODE fallback expands to
an empty body in host code.

Report the newest architecture like the HIP backend does and exclude the
NVIDIA-only features explicitly, as they are not usable on MUSA. Define it
for device passes only: CUB uses defined(CUDA_ARCH) to detect device
compilation, which is also how nvcc behaves.

Drop the now-redundant defined(CUDA_ARCH) checks in the architecture
comparisons: CUDA_ARCH is undefined in host passes for CUDA and MUSA, and
HIP defines it for every pass, so both forms select the same branch.

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.