github ggml-org/llama.cpp b11457

pre-releaseone hour ago
Details

cuda : add BF16 support for XIELU (#29955)

The XIELU CUDA kernel template is already generic over the element
type; only the F32/F16 type assertion and the else-if dispatch were
missing. Add the nv_bfloat16 branch to the launcher, and drop the
temporary supports_op gate in ggml-cuda.cu that rejected BF16+XIELU.

test-backend-ops gains two BF16 cases ([10,5,4,3] and [512,16,1,1]).
docs/ops/CUDA.csv and docs/ops.md are regenerated; the F32 xIELU row
flips from no to yes as well, i.e. the previous record was stale.

Tested:

  • Mac CPU: xIELU F32/F16/BF16, 6/6
  • Mac Metal: existing F32/F16, 4/4; BF16 still unsupported
  • RTX 4090 CUDA: xIELU F32/F16/BF16, 6/6
  • RTX 4090 CUDA BF16-only: 2/2
  • git diff --check passes

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.