ggml-org/llama.cpp b8417
on GitHub

latest releases: b9848, b9847, b9846...

3 months ago

Details

CANN: support flash attention for head dim not multiple of 16, fix ALiBi slope offset (#20031)

Allow FLASH_ATTN_EXT when head dimension D is not a multiple of 16 by
padding Q/K/V to D_padded = GGML_PAD(D, 16), running FusedInferAttentionScoreV2,
then slicing the output back to D (ggml-cann.cpp + aclnn_ops.cpp).
Fix aclnn_get_slope second-part offset: use ggml_type_size(dtype) instead of
sizeof(float) so ALiBi slopes are correct when dtype is F16 (e.g. GQA with
48 heads); fixes buffer overflow and large numerical errors in those cases.

macOS/iOS:

Linux:

Windows:

openEuler:

Check out latest releases or
releases around ggml-org/llama.cpp b8417

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.

Get notifications