github ggml-org/llama.cpp b10655

pre-release58 minutes ago
Details

Feature: Added LIGHTNING_INDEXER support for Deepseek V4 ops on Vulkan Backend (#27453)

  • vulkan: add LIGHTNING_INDEXER op

  • vulkan: updated lightning_indexer.comp and ggml-vulkan.cpp with 128-lane dot-product reduction moved from a shared-memory tree to subgroupAdd.

  • vulkan: cleanup; Skip bounds checks

  • vulkan: cleanup FA_K_ONLY

  • Revert "vulkan: cleanup FA_K_ONLY"

This reverts commit fdcbdd9.

  • vulkan: restore interleaved K/V buffer ordering

  • vulkan: Remove FA_K_ONLY

  • vulkan: Revert flash_attn_dequant

  • vulkan: Revert tests in backend-ops.cpp

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.