github ggml-org/llama.cpp b11338

latest releases: b11342, b11339
pre-releaseone hour ago
Details

hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (#29685)

  • hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA

  • hex-cpy: various fixes on top of the concat optimizations

Removed CONCAT_DMA_MIN_ROW logic, it was broken with 64-bit DMA.
While it's kinda silly to use DMA for tiny stuff if that tensor gets mapped to an extended buffer the only way to read it is DMA.

Added missing dma_queue_flush() calls.

Added additional guards for conditions we don't support.


Co-authored-by: Max Krasnyansky maxk@qti.qualcomm.com

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.