github ggml-org/llama.cpp b11448

latest releases: b11455, b11454, b11451...
pre-release3 hours ago
Details

cuda: BF16/FP16 conversion to f32 chunking (#29442)

  • ggml-cuda: chunk large BF16/FP16 to F32 conversions

  • Update ggml/src/ggml-cuda/ggml-cuda.cu

Co-authored-by: Johannes Gäßler johannesg@5d6.de

  • Update ggml/src/ggml-cuda/ggml-cuda.cu

Co-authored-by: Johannes Gäßler johannesg@5d6.de

  • ggml-cuda: respect dst stride in chunked cuBLAS matmul

Co-authored-by: Johannes Gäßler johannesg@5d6.de

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.