Details
AVX2: Speed up large batch size prompt processing of IQ models (#27402)
- Batched gemm for grid IQ quants
Style updates and a bit more performance
Clean up comments
Move code around
Vectorize IQ panel decode, lower threshold for speedup
IQ panel: single-source gather layout, gate bias, vectorize interleave
Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer
Move IQ panel code out of repack into iqp.cpp, clean up comments
Another comment sweep
-
Add myself as iqp.* codeownder
-
Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size
-
Renaming and moving
-
The other half of renaming and moving
-
Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition
-
Update ggml/src/ggml-cpu/iqp.h
Co-authored-by: Georgi Gerganov ggerganov@gmail.com
-
Add iqp_rows work buffer
-
Revert "Add iqp_rows work buffer"
This reverts commit 4255429.
-
Add NUMA fallback
-
Add 10 row batch tests for IQP coverage on all grid IQ types
-
Swap assert for return false in support check
-
Move IQP mul_mat_id test
Co-authored-by: Georgi Gerganov ggerganov@gmail.com
Website:
Attestations:
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
- iOS XCFramework
Linux:
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (ROCm 7.14)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
Android:
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) - CUDA 12.4 DLLs
- Windows x64 (CUDA 13) - CUDA 13.3 DLLs
- Windows arm64 (CUDA 13) (preview) - CUDA 13.4 DLLs
- Windows x64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (ROCm 7.14)
openEuler:
- DISABLED
- openEuler x86 (310p)
- openEuler x86 (910b, ACL Graph)
- openEuler aarch64 (310p)
- openEuler aarch64 (910b, ACL Graph)
UI: