github ggml-org/llama.cpp b10726

pre-releaseone hour ago
Details

AVX2: Speed up large batch size prompt processing of IQ models (#27402)

  • Batched gemm for grid IQ quants

Style updates and a bit more performance

Clean up comments

Move code around

Vectorize IQ panel decode, lower threshold for speedup

IQ panel: single-source gather layout, gate bias, vectorize interleave

Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer

Move IQ panel code out of repack into iqp.cpp, clean up comments

Another comment sweep

  • Add myself as iqp.* codeownder

  • Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size

  • Renaming and moving

  • The other half of renaming and moving

  • Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition

  • Update ggml/src/ggml-cpu/iqp.h

Co-authored-by: Georgi Gerganov ggerganov@gmail.com

  • Add iqp_rows work buffer

  • Revert "Add iqp_rows work buffer"

This reverts commit 4255429.

  • Add NUMA fallback

  • Add 10 row batch tests for IQP coverage on all grid IQ types

  • Swap assert for return false in support check

  • Move IQP mul_mat_id test


Co-authored-by: Georgi Gerganov ggerganov@gmail.com

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.