github ggml-org/llama.cpp b10425

latest releases: b10427, b10426
2 hours ago
Details

sycl: fuse the gated-delta-net state writeback cpy (#26643)

Port of #23940.

Arc Pro B70, Qwen 3.6 27B Q4_K - Medium (48 of its 64 blocks run
gated_delta_net), -ngl 99 -fa 1 -ctk f16 -ctv f16 -b 2048 -ub 2048,
interleaved A/B passes of r=3:

tg128 23.91 / 23.90 / 23.90 -> 24.19 / 24.17 / 24.20 +1.2%
tg128 (rebuild) 23.81 / 23.81 -> 24.09 / 24.10 +1.2%
pp2048 1050.8 / 1053.9 -> 1053.8 / 1054.5 flat
2 seqs, tg128 32.73 / 32.75 -> 33.11 / 33.10 +1.1%

Website:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.