github ggml-org/llama.cpp b10675

latest releases: b10678, b10677, b10676...
pre-release3 hours ago
Details

Vulkan: add hoisting support for row IDs and expert count in shaders (#26686)

  • vulkan: add hoisting support for row IDs and expert count in shaders

  • use hoisted row ids in coopmat2

  • vulkan: address review feedback on count_experts

  • use vk_op_count_experts_push_constants instead of a raw uint vector
  • apply the fastdiv trick to the ne00 div/mod in count_experts
  • compute the per-expert offsets with subgroupExclusiveAdd when the
    device supports it, keeping the serial path as fallback
  • document the data_d layout and the hoisted_row_id_words bound
  • drop a leftover debug print in ggml_vk_matmul_id
  • vulkan: use init_pushconst_fastdiv for count_experts push constants

  • vulkan: refine comments for row ID hoisting and data layout in count_experts shader

  • Whitespace


Co-authored-by: Jeff Bolz jbolz@nvidia.com

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.