github ggml-org/llama.cpp b10427

one hour ago
Details

sycl: fuse mul_mat(gate) + mul_mat(up) + GLU for q4_K dense FFN (#26779)

Measured on Arc Pro B70 (Battlemage, Level Zero), llama-bench -r 20, two
interleaved rounds, tg128:

qwen2.5-3B-Instruct Q4_K_M    154.18 -> 158.53 t/s   +2.8%
gemma-2-2b-it Q4_K_M          162.45 -> 165.62 t/s   +2.0%

llama-batched-bench on qwen2.5-3B, S_TG by batch size:

  B=1   142.72 -> 147.57 t/s    +3.4%
  B=2   243.72 -> 268.26 t/s   +10.1%
  B=4   359.58 -> 398.02 t/s   +10.7%
  B=8   449.75 -> 505.63 t/s   +12.4%

Website:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.