github ggml-org/llama.cpp b10208

latest releases: b10211, b10210, b10209...
4 hours ago
Details

SYCL: add oneMKL GEMM flash attention for XMX-accelerated prompt proc… (#25025)

  • SYCL: add oneMKL GEMM flash attention for XMX-accelerated prompt processing

  • fattn-mkl: fix interleaved dst layout in normalize kernel

  • Fix mkl_fa_normalize_head: use interleaved dst layout
    ((query * n_q_heads + head) * DV) matching TILE's
    flash_attn_combine_results. Previously used dense head-major
    layout which wrote head outputs to wrong addresses, corrupting
    attention for all models except Qwen3.6-27B (where GQA=6 heads
    were sparse enough to avoid visible overlap).

  • Remove 7 redundant stream->wait() calls — SYCL in-order queue
    already serializes pure SYCL kernel dependencies. Retain only
    the 4 MKL GEMM ↔ SYCL handshake barriers (oneMKL GEMM uses its
    own internal queue that does not respect SYCL in-order).

  • Remove unused dst_row_stride, diagnostic clutter, and dead
    K/V hex dump (fa_diag block in fattn-mkl.cpp).

  • Add MKL_FA_DISABLE=1 env var for A/B testing.

  • Add FA-DISP watchdog (MKL_FA_DEBUG=1) and FA-DIAG output
    fingerprint (MKL_FA_DIAG=1) in fattn.cpp.

Tested: Gemma-4-26B, Gemma-4-31B, Qwen3.6-27B, Qwen3.6-35B-A3B
Perf (B70/Battlemage, 32K, q8_0 KV):
Gemma-4-26B: 1473 t/s MKL vs 746 TILE (1.97x)
Qwen3.6-27B: 609 t/s MKL vs 330 TILE (1.85x)

Co-Authored-By: Claude Code on DeepSeek-v4-Pro

  • Thank you for the review feedback: rename env vars, use GGML_LOG_INFO, document in SYCL.md

Completed the following:

  • Rename MKL_FA_DISABLE → GGML_SYCL_ENABLE_MKL_FA (inverted: 0 to disable)
  • Rename MKL_FA_DEBUG → GGML_SYCL_MKL_FA_DEBUG
  • Rename MKL_FA_DIAG → GGML_SYCL_MKL_FA_DIAG
  • Replace fprintf(stderr, ...) / fflush(stderr) with GGML_LOG_INFO() macro
  • Document all three env vars in docs/backend/SYCL.md under Runtime
  • Add comment explaining MKL FA activation trigger (flash-attn + quantized
    KV cache + batch-size >= 1024 + n_kv >= 1024)

Resolves review feedback from arthw.
Again, thank you!!!

Co-Authored-By: Claude Code on DeepSeek-v4-Pro

  • Thank you for the review feedback round 2: use ggml_sycl_get_env, remove dup waits, gate perf macros
  • Replace raw getenv() with ggml_sycl_get_env() in all 4 env-var checks
    (fattn.cpp: GGML_SYCL_ENABLE_MKL_FA, GGML_SYCL_MKL_FA_DEBUG,
    GGML_SYCL_MKL_FA_DIAG; fattn-mkl.cpp: GGML_SYCL_MKL_FA_DEBUG)
  • Remove duplicated stream->wait() before ev.wait_and_throw() in GEMM
    KQ and GEMM VKQ — ev.wait_and_throw() already waits for completion
  • Gate MKL_ACCUM macro behind do_print so timing accumulators are
    no-ops in normal operation
  • Remove redundant MIT/Intel copyright header from fattn-mkl.cpp
  • Remove unused #include
  • Expand SYCL.md MKL FA docs with step-by-step activation trigger
    and example llama-cli command

Again, thank you!!!

Co-Authored-By: Claude Code on DeepSeek-v4-Pro

  • fattn-mkl: enable MKL FA for all KV cache types

Remove the quantized-only restriction on MKL activation — the MKL
kernel converts any non-F16 K/V to F16 via to_fp16_sycl before GEMM,
so F16 (default), BF16, and F32 caches all benefit from XMX hardware
acceleration. The type restriction was an unnecessary gate.

Before (F16/BF16 default cache + FA on at 32K prefill): ~356 t/s (TILE path)
After: ~670 t/s (MKL path, matching quantized-cache baseline)

Minimal change: two conditions removed, one comment updated in fattn.cpp.
No kernel or conversion code changes — the dequant pipeline already
covers all types.

  • fattn-mkl: rename mkl_disable -> mkl_enable for clarity

  • fattn-mkl: refine MKL FA dispatch gates

Three changes:

  1. Remove quantized-only restriction - MKL FA activates for all
    KV cache types (F16 default, BF16, F32, quantized). The MKL
    kernel converts non-F16 K/V via to_fp16_sycl before GEMM.
  2. Rename mkl_disable -> mkl_enable to match env var
    (GGML_SYCL_ENABLE_MKL_FA).
  3. Replace batch-size threshold with Q->ne[1] >= 32 gate.
    Keeps TG (Q=1) and MTP drafts (Q=3-8) on VEC path where
    fused kernel beats MKL launch overhead. Routes all
    multi-token prefill through XMX-accelerated GEMM.

Production data confirms Q patterns: 1-8 TG, 32-127 cache reuse,
128+ full reprocess. At 32K F16/BF16 FA-on: 356 -> 670 t/s.

  • ggml-sycl: fix F16 cache + MKL FA multi-turn corruption; add gate guards

Two changes:

  1. Always copy F16 K/V to dense row-major buffers before MKL GEMM.
    Previously F16 was read in-place with raw tensor strides. During
    multi-turn conversations, the accumulated KV cache had different
    stride properties than a fresh prefill, producing corrupted outputs.
    Now dense F16 gets a fast memcpy; interleaved (Gemma) gets a strided
    copy kernel. This matches what the quantized paths already did through
    to_fp16_sycl.

  2. Gate MKL FA on unsupported op params (max_bias, logit_softcap, batch
    dim mismatch) and pathological F16 strides (nb[1] not a multiple of
    ne[0]*2). These conditions would previously crash inside the MKL
    kernel. Pathological strides (test-only) and ALiBi/softcap fall
    through to TILE/VEC which handle them correctly.

The stride check uses modulo rather than equality, so both dense
(nb1 == ne02) and interleaved (nb1 == H * ne02) pass — all real
models use these layouts. Only test cases with overlapping rows
(nb1=32 or nb1=75 for ne0=40) are blocked.

Thanks to hmscider for the oneDNN FA PR (#25222) which surfaced the
same insight: always normalize inputs to contiguous F16 before GEMM.

Co-Authored-By: Claude Code using DeepSeek-V4-Pro noreply@anthropic.com

  • fattn-mkl: fix quant+GQA KV strides, tighten MKL gate, add K>=1024 tests

Adding K>=1024 flash-attn test cases surfaced several MKL bugs:

  • Quant K/V with a padded seq-view (real KV cache) used the wrong
    strides in the dequant path... only the true Gemma interleave
    layout should reconstruct strides. nb[2] vs ne[1]*nb[1]
  • Gate was firing on shapes the kernel doesn't handle: head_dim < 64
    or not a multiple of 64, MHA, attention sinks, and
    bf16 decode... fell through to vec which no bf16 case.

Gate MKL to the validated envelope: gqa>=2, head_dim 64 through 512
(has to be a multiple of 64) with matching K/V head size, mask,
no sinks/alibi/softcap... everything else falls back to tile.
Covers Qwen Dense/MoE and Gemma4 Dense/MoE

Ran test-backend-ops -o FLASH_ATTN_EXT: 3641/3641 pass.
Perplexity unchanged... 6.7267 MKL vs 6.7290 stock using
Qwen 27b q5_k_xl

  • Update ggml/src/ggml-sycl/fattn.cpp

Co-authored-by: Neo Zhang zhang.jianyu@outlook.com

  • Update ggml/src/ggml-sycl/fattn.cpp

Co-authored-by: Neo Zhang zhang.jianyu@outlook.com

  • Update ggml/src/ggml-sycl/fattn.cpp

Co-authored-by: Neo Zhang zhang.jianyu@outlook.com

  • fattn-mkl: bound attention scratch so it doesn't grow with batch or context... also dropped the bf16 comment in fattn.cpp per arthw review.

  • Update ggml/src/ggml-sycl/fattn-mkl.cpp

Co-authored-by: Neo Zhang zhang.jianyu@outlook.com

  • Update ggml/src/ggml-sycl/fattn-mkl.cpp

Co-authored-by: Neo Zhang zhang.jianyu@outlook.com

  • apply arthw suggestions: enum for dequant modes, macro for wg_size, env-var one-liners

Co-authored-by: Claude Code using DeepSeek-V4-Pro noreply@anthropic.com
Co-authored-by: Neo Zhang zhang.jianyu@outlook.com

Website:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.