Details
context : do not re-reserve the scheduler when toggling causal_attn (#28751)
- context : do not re-reserve the scheduler when toggling causal_attn
llama_context::set_causal_attn() marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma models, resulting in two expensive sched_reserve() passes per image. This is especially slow for multi-image or video inputs.
The cost of a re-reserve scales with context and ubatch configurations, so larger settings pay more per image (see table below).
The re-reserve is unnecessary in this case because causal_attn only changes the values written to KQ mask, not tensor shapes or any other buffer sizes.
Note: causal_attn is a graph reuse key (llm_graph_params via cparams), so a new graph is built regardless of sched_need_reserve, so this doesn't change the graph rebuilding behaviour.
llama-server with gemma-4-26B-A4B Q4_0 + BF16 mmproj, 130-token images,
cache_prompt=false, prompt_ms median of 3 (before -> after):
| images | config | H200 before -> after | RTX 4090 before -> after |
|---|---|---|---|
| 1 | -c 8192 -ub 512
| 134 -> 105 ms (1.27×) | 201 -> 119 ms (1.69×) |
| 24 | -c 8192 -ub 512
| 2278 -> 1562 ms (1.46×) | 3559 -> 1748 ms (2.04×) |
| 24 | -c 32768 -ub 2048
| 5379 -> 1584 ms (3.40×) | 13377 -> 1759 ms (7.61×) |
Generated output remains identical before and after.
- qwen4exp : make the indexer bias shape independent of causal_attn
The block/cell bias path was selected on cparams.causal_attn, so the
causal and non-causal graphs differed in tensor shapes and ops. With the
re-reserve removed (previous commit), a runtime flip resulted in
reallocating the compute buffers, which would fail under
GGML_SCHED_NO_REALLOC.
This commit selects the block path from the mask shape only, independent
of causal_attn. causal_attn is instead passed to set_input_qsa.
causal_attn is fixed per graph as it's part of the reuse key. Causal
values are unchanged. Non-causal values now follow the reference rule,
where every visible block competes on score and only unpooled cells are
always selected.
-
context : state the causal_attn shape rule in the comment
-
cont : add TODOs
Co-authored-by: Georgi Gerganov ggerganov@gmail.com
Website:
Attestations:
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
- iOS XCFramework
Linux:
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (CUDA 12) - CUDA 12.8 libraries
- Ubuntu x64 (CUDA 13) - CUDA 13.4 libraries
- Ubuntu arm64 (CUDA 13) - CUDA 13.4 libraries
- Ubuntu x64 (ROCm 10.0)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
- Linux arm64 (Snapdragon: CPU, Adreno GPU, Hexagon NPU) - setup guide
Android:
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) - CUDA 12.4 DLLs
- Windows x64 (CUDA 13) - CUDA 13.4 DLLs
- Windows arm64 (CUDA 13) - CUDA 13.4 DLLs
- Windows x64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (ROCm 10.0)
openEuler:
- DISABLED
- openEuler x86 (310p)
- openEuler x86 (910b, ACL Graph)
- openEuler aarch64 (310p)
- openEuler aarch64 (910b, ACL Graph)
UI: