github ggml-org/llama.cpp b11351

pre-releaseone hour ago
Details

ggml : add alloc_buffer_n to buffer type interface (#23671)

  • ggml : add alloc_buffer_n to buffer type interface

Add alloc_buffer_n method to ggml_backend_buffer_type_i
interface, with a public API ggml_backend_buft_alloc_buffer_n.

  • Default implementation in ggml-backend.cpp handles multi-buffer
    splitting and tensor allocation via ggml_tallocr
  • Meta buffer type provides custom implementation that creates
    per-device sub-contexts and delegates to simple buffer types
  • ggml_backend_alloc_ctx_tensors_from_buft now collects tensors
    into a list and delegates to the new API
  • Remove temporary ggml_backend_meta_alloc_ctx_tensors_from_buft
  • Add NULL alloc_buffer_n to all existing buffer type
    interfaces (cpu, metal, openvino, hexagon, webgpu, zdnn, virtgpu, repack)

Assisted-by: llama.cpp:local pi

  • cont : fix cur_buf_size init after flushing a buffer

  • ggml : add TODO tag for shared buffer split logic

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

  • tests : add alloc_buffer_n coverage

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

  • cont : fix compile warnings

  • tests : add descriptions for alloc_buffer_n tests

Assisted-by: pi:llama.cpp/Qwen3.8-27B

  • ggml : address review comments on alloc_buffer_n
  • restore GGML_LOG_ERROR on buffer alloc / tensor init failure in the
    default impl (name the failing tensor)
  • check the malloc result and drop the _impl indirection in
    ggml_backend_alloc_ctx_tensors_from_buft
  • remove comments that restate the code
  • fix the TAG_ALLOC_SHARED_BUFFER_SPLIT typo

Assisted-by: pi:llama.cpp/Qwen3.8-27B

  • ggml : add get_alloc_size_n to buffer type interface
  • Add ggml_backend_buft_get_alloc_size_n public API
  • Add optional get_alloc_size_n callback to ggml_backend_buffer_type_i
  • Share tensor->buffer planning between alloc_buffer_n default and get_alloc_size_n default
  • Replace unchecked realloc with std::vector in alloc_buffer_n default
  • Make ggml_backend_alloc_ctx_tensors_from_buft_size use the new API
  • Add test-alloc coverage for get_alloc_size_n

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

  • cont : report malloc failure

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.