Details
llama : fix unexpected graph reallocation in the k-pool models (#29958)
- llama : fix unexpected graph reallocation in the k-pool models
Both k-pool models built a graph shape that depends on state the
full-context reserve cannot know:
- qwen4exp branched on inp->cache_safe, which turns false as soon as
llama_memory_seq_cp shares cells (e.g. batched-bench -pps): the QSA
layers swapped scatter+gather for fill+concat and dropped the
new_pool_rep leaf, so the decode graph had 12 fewer nodes than the
reserved one - glm5-next branched on gather = n_tokens <= 16 && n_kv > n_sel, so the
TG decode built the gather shape (7564 nodes) while the last reserve,
the PP one, had the dense shape (7762 nodes)
Either mismatch forces a decode-time re-reserve that drops the
worst-case sizing and bakes in the current state, so the next state
growth (n_pool, n_kv, n_new) needs more room at an unchanged graph size
and aborts under GGML_SCHED_DEBUG_REALLOC=1. Reproduce with, e.g.:
GGML_SCHED_DEBUG_REALLOC=1 ./bin/llama-batched-bench
-hf ggml-org/GLM-5.3-Flash-GGUF:Q2_K -npp 2500 -ntg 32 -npl 1,2
-c 32768 -pps -kvu
Always scatter+gather the pooled keys, and pick gather from context
constants only: n_ubatch bounds every ubatch, top_k + kpool - 1 bounds
n_sel. Every graph of a context then shares one shape, which the
reserve covers, and the dense path measured faster than the gather path
at 2.5k and 16k context.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
- llama : drop the unused k-pool cache_safe graph API
The k-pool graphs no longer branch on cache_safe, so nothing reads
get_kpool_cache_safe() or the conditional new_pool_rep any more: both
models always pass the scatter target, which set_input_kpool now
requires instead of merely preferring.
Also drop the cache_safe copy in kpool_build_sizes(), a sizes-only
helper. The layout and state flag itself stays, it still decides which
pools a layout with shared cells must re-pool.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
- tests : add a shared-seq graph reserve regression test
Decode a prompt into seq 0, share its cells with seq 1 via
llama_memory_seq_cp (what llama-batched-bench does for -pps), then keep
decoding both sequences. For the k-pool models sharing clears
cache_safe, which changes the graph topology while the pools keep
growing, so a scheduler that re-reserves with the current state
instead of the worst-case one aborts under GGML_SCHED_DEBUG_REALLOC=1.
The test registration sets that flag, and the test aborts on both
k-pool models before 2220411.
kimi-linear and minimax-01 are skipped: they reserve the final pp graph
with n_seqs = 1 (see [TAG_RESERVE_DIAG_DECAY] in llama-context.cpp), so
every multi-seq graph has a different layout and re-reserves by design.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
-
cont : add TODOs
-
cont : fix comment
-
cuda: match the moe weighted reduction on empty ubatches
ggml_cuda_match_moe_weighted_reduction rejected tensors with zero
rows. A ubatch without outputs shrinks the last layer to zero rows
through inp_out_ids, so graph_optimize dropped its alloc dep there and
the scheduler graph lost one node compared to the reserved one. The
scheduler then re-reserved at the size of that ubatch, and the next
ubatch with the same node count but larger tensors aborted under
GGML_SCHED_DEBUG_REALLOC=1.
The compute loop already skips empty nodes before trying any fusion,
so the guard only made the alloc deps depend on the row count.
- tests: build the rollback test only where internal symbols link
The shared-seq case calls llm_arch_from_string, which libllama does
not export through LLAMA_API, so linking test-recurrent-state-rollback
fails on Windows with shared libraries. Its build now sits in the
NOT WIN32 OR NOT BUILD_SHARED_LIBS block, next to test-llama-archs and
the test registration it already lives under.
- tests: skip archs by name in the shared-seq reserve test
The skip of kimi-linear and minimax-01 went through llm_arch_from_string,
which libllama does not export through LLAMA_API, so the test could not
link on Windows with shared libraries. It now compares the
general.architecture string directly, and the test builds on every
platform again.
Co-authored-by: Pascal admin@serveurperso.com
Website:
Attestations:
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
- iOS XCFramework
Linux:
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
- Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
- Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
- Ubuntu x64 (ROCm 10.0)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
- Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) - CUDA 12.4 DLLs
- Windows x64 (CUDA 13) - CUDA 13.4 DLLs
- Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
- Windows x64 (Vulkan)
- Windows arm64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (ROCm 10.0)
openEuler:
- DISABLED
- openEuler x86 (310p)
- openEuler x86 (910b, ACL Graph)
- openEuler aarch64 (310p)
- openEuler aarch64 (910b, ACL Graph)
UI: