github ggml-org/llama.cpp b11490

pre-releaseone hour ago
Details

hexagon: enable alloc_buffer_n (#30126)

  • hex-bufs: add support for alloc_buffer_n

  • hex-bufs: add support for splitting large tensors into separate buffers

  • hex-bufs: update GGML_HEXAGON_MBUF to accept three values dyn,static,total

  • hex-bufs: bump dyn. default to 512MB since 128MB causes perf regressions with big MOEs

  • hex-run: add --no-embd-offload option to simplify command lines on devices that need it

  • Update scripts/snapdragon/run.py

Co-authored-by: Jhen-Jie Hong iainst0409@gmail.com


Co-authored-by: Jhen-Jie Hong iainst0409@gmail.com

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.