github ggml-org/llama.cpp b10982

latest releases: b10985, b10984, b10983...
pre-release5 hours ago
Details

vulkan: support sparse Flash Attention (#28105)

  • vulkan: add sparse Flash Attention support for DSV4/GLM

  • tune implementation

  • add tests

  • avoid nondeterministic atomicAdd

  • add cm2 decode vector support

  • simplify logic and make variable names more consistent

  • add cm2 f16vec4 binding for decode vector

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.