github ggml-org/llama.cpp b11408

pre-release38 minutes ago
Details

ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (#28479)

  • ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60

On the SpacemiT X60, IME matrix acceleration only covered Q4_0/Q4_1/Q4_K.
Q8_0 had no IME1 kernel, and since the SpacemiT build sets
GGML_CPU_REPACK=OFF there was no repack path compiled in either, so Q8_0
had no accelerated path at all and ran roughly ten times slower than
Q4_0 for prefill on the same board.

  • add make_block_q8_0x16 and the Q8_0 repack entry: interleave the
    weights into the 16-column layout the IME1 vmadot sequence expects
  • add ime1::gemm_kernel_i8i8, an int8 x int8 IME1 kernel with a
    single-row and a 4-row A path; the 4-row path loads each B panel once
    and reuses it across 4 rows of A
  • add quantize_a_4row_i8 for the 4-row activation quantization
  • wire both into forward_mul_mat and the repack factory for Q8_0
  • docs: mark Q8_0 as supported on X60

Correctness was checked against a quant-exact integer reference for
K = 32 up to 4096, with a max relative error of about 1e-6, and by
checking that generation stays coherent across several prompts.

Tested on Milk-V Jupiter (SpacemiT X60), Bianbu 2.1.1, gcc 14.2, with
Qwen2.5-0.5B-Instruct Q8_0. llama-bench -t 4 under taskset -c 0-3, 5
repetitions on an idle board: pp128 goes from 10.70 to 93.87 t/s. Q4_0
is unchanged at 106.40 -> 107.51 t/s, as expected since this does not
touch that path.

  • ggml-cpu : move q8_0_16x32 decl to IME1 section

  • ggml-cpu : align q8_0 IME1 kernel assignments

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.