github ggml-org/llama.cpp b10164

latest release: b10165
one hour ago
Details

ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration (#22675)

  • ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration

  • cuda: added SSD CICD fixes for CUDA / HIP / MUSA / MSVC.

  • ggml-cuda: review comments fixed.

  • ggml-cuda: Fuse M matrix materialization into pre_matmul kernel and enabled test.

  • ggml-cuda: test updates and fixes

  • ggml-cuda: test updates to remove hardcoding of tensor initialise data limits.

  • ggml-cuda: ssd minor review comment fixed.

  • ggml-cuda: ssd minor CICD fixed.

  • CUDA SSD: Fixes correctness by promoting s0_stride_seq to int64_t, improves memory coalescing in ssm_ssd_prepare_dt_kernel, and boosts efficiency by merging B_weighted and C_scaled; also addresses prior review comments.

  • cuda: fix sdata read-write race in prepare_dt fallback scan loop

Website:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.