github ggml-org/llama.cpp b10883

pre-release59 minutes ago
Details

vulkan: use spec constant for matrix matrix multiplication A-type (#25773)

  • vulkan: use spec constant for mul mat type_a

vulkan: use map for mul_mm shapes

cleanup

fix indentation

fix cm2 and shmem init

fix cm2 spec constants

fix cm2 bindings

consolidate shmem tables and reduce size by type spec constant

fix compiler warning

fix missing Q2_0 type

fix unused warning when integer dot glslc support is missing

use minimal shmem size 8 instead of 1 to workaround cm2 compiler bug

fix missing Q2_0 type in cm2 matmul

fix types

  • remove LUT quants from unified shader

  • clean up

  • restore coopmat2 q4_k/q5_k optimization

  • split out q4_k/q5_k cm2 shader to fix Ampere regression

  • revert iq shmem table renames

  • simplify cm2 code with single uint8_t buffer

  • fix fp4 extension use switch being overwritten by generic shader

  • clean up

  • adapt TQ1_0 changes

  • adapt #27471 f16 Intel tuning changes

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.