github ggml-org/llama.cpp b11163

pre-releaseone hour ago
Details

llama: add llama_batch_ext (#24669)

  • (wip) add llama_batch_ext

  • wip

  • updated design

  • updated impl

  • change signature

  • unused var

  • demo common_prompt_batch_decode

  • fix pos

  • tmp disable test-batch-alloc

  • fix compat

  • nits: add const

  • no more pos_max

  • add comment about llama_batch_ext_set_embd_state

  • handle n_embd_out properly

  • rename api --> embd_token

  • llama_embd

  • stub llama_batch_ext_set_embd_state

  • support both token + embd + state in batch

  • llama_batch_ext_add_embd

  • upstream some changes

  • nits

  • fix test-batch-alloc

  • add test for compat

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.