github ggml-org/llama.cpp b11307

latest releases: b11309, b11308
pre-releaseone hour ago
Details

llama : preserve original batch order for speculative decoding layer inputs (#29019)

  • llama: preserve original batch order for layer inputs

Assisted-by: Codex

  • tests: cover layer-input order across KV layouts

Assisted-by: Codex

  • tests: exercise layer-input ordering on CUDA devices

Assisted-by: Codex

  • llama: make layer input reordering compatible with tensor split

Copy each microbatch tensor from offset zero and restore original row order after synchronization. Extend the layer-input regression to cover tensor split and repeated reads and decodes.

Assisted-by: Codex

  • llama: restore token order for unmasked NextN embeddings

Use the original-token mapping for unmasked NextN rows, including when
layer-input capture is disabled. Keep masked NextN rows on the logits
output mapping and preserve offset-zero tensor copies.

Extend the existing regression to cover NextN alone, combined layer
capture, and masked outputs with repeated decodes and getters.

Validation: all 256 CPU/CUDA/tensor configurations pass. Qwen3.8-27B
Q4_K_M MTP completes MT-Bench at concurrency 16 before and after.

Assisted-by: Codex

  • ggml: fix WebGPU reservation and OpenVINO hidden-state capture

Reserve WebGPU vector attention scratch across batch sizes and refresh reservations when NextN capture settings change. Preserve requested OpenVINO outputs, dynamic shapes, sequence counts, and current graph bindings.

Extend existing WebGPU regression coverage and enable strict allocation checks.

Assisted-by: Codex

  • llama: defer regression test and backend fixes to follow-ups

Keep this PR focused on restoring token order for layer inputs and unmasked NextN embeddings. Remove the added regression test, OpenVINO and WebGPU changes, and the separate NextN reservation change.

Assisted-by: Codex

  • llama: keep n_embd declaration in its original position

Assisted-by: Codex

  • llama : pass token count to layer input extraction

Assisted-by: Codex

  • llama : name original batch indices batch_idxs

Assisted-by: Codex

  • llama : name extracted embedding indices embd_batch_idxs

Assisted-by: Codex

  • llama : tag target embedding reordering

Assisted-by: Codex

  • llama : tag extraction and name the index capture flag

Assisted-by: Codex

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.