github ggml-org/llama.cpp b11552

pre-release2 hours ago
Details

server: leave a busy slot untouched when a request pins it (#30295)

A request asking for a busy id_slot still ran the prompt cache update
on that slot before being deferred. When the RAM cache held a better
match, it was loaded into the slot while another request was still
generating there, and that generation continued on the wrong context.

The busy slot is now returned as is and the request waits for it.

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.