Overview
llama.cpp v0.6.0 introduces the new llama_batch_ext extended batch API (with llama_process) for mixed token/embedding inputs and MTP/deepstack state embeddings, adds support for the GLM-5.3-Flash (GLM5-Next) 320B hybrid model, the Clef decision model (text and vision) and MTP speculative decoding for Qwen4Exp, ships a new /v1/systemone server API for decision models (laya, julia-1, lev, openjev, kev, nimble), overhauls the Web UI with a Hugging Face Hub data layer and model download pipeline, adds a Metal tensor-API flash attention kernel for F16 KV, sparse flash attention for quantized K/V on Vulkan, and updates ggml to v0.26.0.
Highlights
- New
llama_batch_extextended batch API withllama_process(), supporting mixed token/embedding batches and per-token "state" embeddings for MTP and deepstack models #24669 - New models: GLM-5.3-Flash (GLM5-Next), a 320B text+vision hybrid model #27773, and the Clef decision model, fully supported with both text and vision #29831 #29969
- Qwen4Exp: high-quality support is now available, with MTP speculative decoding (~1.5x decode speedup on DGX Spark) and various correctness fixes #29761 #29751
- llama and server: new
/v1/systemoneAPI supporting five decision models - laya, julia-1, lev, openjev (+vision), kev #29818 - Metal: new tensor API flash attention kernel for F16 KV #29570
- Metal: new few-row MMA mat-mul kernels for speculative and batched decoding, up to ~3x faster mat-mul on Apple GPUs #29869
- New
llama_prefetch_rows()using MADVISE-based prefetching of PLE tensors in Qwen4Exp and Gemma4 #29599
API changes
include/llama.h: newllama_batch_extbatch API withllama_embd,llama_process()andllama_process_type#24669, newllama_get_causal_attn()#28876, session formats bumped toLLAMA_SESSION_VERSION11 andLLAMA_STATE_SEQ_VERSION4include/llama-cpp.h: addedllama_batch_ext_ptrand deleter for the new extended batch API #24669tools/mtmd/mtmd.h:mtmd_get_memory_usage()now returns anmtmd_memory_usagestruct withimage_max_tokensanduse_non_causal#29773tools/server: new/v1/systemoneendpoint for decision models #29818 and/v1/embeddingsnow accepts typed vision/audio/video content #29556
New models
- GLM-5.3-Flash (GLM5-Next): 320B KDA/DSA hybrid text+vision model with mHC and MoE #27773
- Clef decision model, fully supported with both text and vision #29831 #29969
- Ling 3.0 VL, folded into the BailingMoeV3 architecture #29151
- Nimble decision model #29844
- Registered
Lfm2BidirectionalForMaskedLMfor LFM2.5-Encoder-230M/350M #29862 - Added
classifier_poolingsupport for rerankers #29627
Core changes
- Migrated examples, speculative decoding, mtmd and server to the new
llama_batch_extAPI #29385 #29601; batches now accept both embd and raw tokens #29622 - Added
llama_prec_policyand a model-driven W4A4 (NVFP4/MXFP4) mul_mat path #24364 - Qwen4Exp: halved indexer score memory #29825, optimized mask constructions #29824, re-enabled the
-smtensor #28569; GLM5-Next: unique scatter rows for dead indexer slots #29745 - KV cache: fixed restoring mismatched KV cache rotation #28498, fixed K/V and recurrent state cleanup after failed restores #27530 and an invalid assert in recurrent memory #29799
- Speculative decoding: probabilistic sampling for simple draft and MTP #27694, fixed n-gram drafts rejected at temp > 0 after truncation #29924, preserved original batch order for layer inputs #29019, stop accepting draft tokens at EOG #29638
- k-pool models: fixed unexpected graph reallocation #29958 and clamped kpool re-pool bound to existing pools #29805
- Fixed tensor split for fused qkv with uneven K/V head sizes #29294, gather recurrent states once so the reserve covers every split #29856, and properly handle KV on training #28520
- Context: do not re-reserve the scheduler when toggling
causal_attn#28751 - DFlash drafts: write Gemma embedding scale during conversion #29802 and add dflash support for MiMo #29650
Multi-modality changes
- Cap
max_imageton_ubatchfor non-causal models #29773 - Fixed the mel preprocessor in LFM2 audio #29403
- Migrated input processing to
llama_batch_ext#29385
Server changes
- New
/v1/systemonedecision-model API with dedicated decision pipeline #29818, extended to the nimble decision model #29844 - Support vision input for Clef #29969
GET /v1/modelsandGET /modelsnow report model input/output modalities in a newarchitectureobject #29987/v1/embeddings: accept typed content (vision/audio/video) input #29556 and return HTTP 400 for invalid embedding requests #29060- Allow RANK pooling batch splitting for causal LLM rerankers (Qwen3, Qwen3-VL) #28876
- Reject partial media truncation #24076
- Logging: self-contained colors and split child commands from logs in router mode #29895, allow preset to set log file #29334, fixed dead
LLAMA_ARG_HF_REPO_FILEkey in preset allow-list #29938 - Fixed laya abort by limiting
n_batchton_ubatch#29903 - Remove the built-in UI's service worker when the UI is not served #29565
UI changes
- New model download pipeline #27959, Hugging Face Hub data layer #27947, model memory-fit estimation #27957 and model id grammar for sidecars, quants and capability parsing #27946
- Type-safe API types, fetch helpers and download-ready models store plumbing #29582
- Shared model display primitives #29644
- Fixed missing
svg useand animation elements in preview and download #28962 - Use
toLocaleString()formatting consistently across chat message statistics #27990
ggml changes
- ggml updated to v0.26.0: a new
alloc_buffer_n/get_alloc_size_nbuffer allocation API, sparse flash attention kernels on SYCL, Vulkan and Metal, and major lightning indexer improvements (halved score memory, tiling, MUSA support). The CPU backend gains BF16 ops and a tiled k-quant mul_mat, CUDA gains a model-driven W4A4 (NVFP4/MXFP4) mul_mat path plus MMVQ shared-expert fusion, and the Hexagon backend adds a sampler and more quant types. Model loading is faster with stricter GGUF size validation, Windows ARM64 MSVC builds are enabled, and WebGPU/OpenVINO/OpenCL/SYCL pick up numerous new kernels, ops and fixes.
Assets
Nightly build: b11429
More info
Changelog since v0.5.0
d812350 llama.cpp : bump version to 0.6.0 (#29997)
4d60b4d common, server : report model input/output modalities in GET /models (#29987)
c06f841 sync : ggml
f05c8b2 ggml : bump version to 0.26.0 (ggml/1652)
e117148 CUDA: make the alloc_deps check batch independent (#29986)
6c59c40 vulkan: fix Flash Attention shmem write out of bounds (#29988)
3c9e747 vulkan: revert mul_mat_id tile selection PR #29182 (#29936)
b809b88 cuda: use the vector lightning indexer kernel on MUSA (#29990)
994e8f2 ci : add "Require Docker" flag to make-release workflow (#29989)
9d853bb webui: Use toLocaleString() format consistently across chat message statistics (#27990)
8f9ae20 ci : disable failing test on virtual Metal device (#29993)
9871df5 server: support vision input for Clef (#29969)
8b2fbaf CUDA: Optimize accumulation in mmq for NVFP4 type (#29857)
2ed93db ci : disable unused qemu in docker build (#29984)
8e16421 server: reject partial media truncation (#24076)
806eee9 vulkan: fix stale prealloc_y reuse across flash attention and soft_max (#29591)
b3daa07 vulkan: sparse flash attention for quantized K/V (#29639)
c173a53 llama : fix unexpected graph reallocation in the k-pool models (#29958)
2107910 kv-cache: fix restoring mismatched KV cache rotation by saving exact rotation metadata (#28498)
e5983d6 ci : winget urls must be separate strings (#29978)
4ca6b76 ci : fix docker workflow permissions (#29979)
9f12cd4 ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (#28479)
ebe18be vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON (#29912)
8216c84 webgpu: add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 (#29483)
1b43d31 cuda: tile the lightning indexer over keys and tokens for 4 heads (#29901)
a3a1c47 metal : few-row MMA mat-mul (#29869)
9d3aba6 CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size (#29633)
d89651a CUDA: prefer whole-tile FlashAttention scheduling for efficient two-stage kernels (#29435)
a7fb71f log, server: self contained colors, split child commands from logs in router mode (#29895)
0bb496d llama: support both embd + raw tokens in batch (#29622)
2ca15f5 CUDA: refactor swizzling code (#29612)
a7b94df ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (#29806)
0eb6d9a cuda : move neu_padded to where it is used (#29940)
2e7c58c ci : windows llvm build requires ninja multi-config (#29959)
7f2dd88 ci : add windows arm64 vulkan release (#29954)
bf79dbb AGENTS.md : revamp (#29656)
dbe4c3e chat-peg-parser : clear current_tool when pending_tool_call is reset (#29942)
46847e6 ci : set default permissions (#29945)
2bc5635 cuda : move blocks_per_col to where it is used (#29939)
dd26678 CUDA: fix MMQ memory fault if n_expert >> n_ubatch (#29941)
16c163d vulkan: fix rdna4 mat_vec tuning (#29934)
0504396 imatrix: calculate activation-based statistics for new format (GGUF) imatrices (#14891)
8330e96 spec : fix n-gram drafts rejected at temp > 0 after truncation (#29924)
6716df6 common : prepare load_from_models_dir() for path conversion (#29674)
bf9a0cc server : fix dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list (#29938)
0faee50 ci : pushing tag needs deploy key (#29937)
f98b31c ci : improve release flow (#29913)
11fe021 webgpu: add f16 support to fill/set_rows (#29897)
836d571 mtmd : fix deprecated strdup warning on Windows (#29863)
eec18f5 vendor : update cpp-httplib to 0.59.0 (#29886)
1537a0a server : fix laya abort by limiting n_batch to n_ubatch (#29903)
edd6e2b common : add common_is_tty() helper and fix deprecated warnings on Windows (#29860)
9bf55f4 chat : honor json_schema in Ling 3.0 parser (#29813)
a55e952 ci: fix flaky ADD_ADD f16 by using the fused ADD tolerance (#29904)
436f6f8 graph: gather the recurrent states once so the reserve covers every split (#29856)
b92761a ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (#29852)
cb7934c model : Add LFM2.5-Encoder-350M and LFM2.5-Encoder-230M (#29862)
889edf4 qwen4exp : halve the indexer score memory (#29825)
99b9548 model: add support for clef decision model (text-only) (#29831)
bed0a85 CUDA: fuse shared experts into MMVQ (#29184)
4ebdf2c ci : use t4-medium for cuda jobs (#29842)
1fb7ef3 spec : add probabilistic sampling for simple draft and MTP (#27694)
134b2bb ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (#27663)
2923cf2 ggml-quants : avoid invalid rounding in qkx3 scale search (#29817)
dd4c286 ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (#27096)
46ca246 model: support nimble decision model (#29844)
d8fbd25 readme : add cmd install commands (#29850)
926862e metal : add tensor API flash attention kernel for F16 KV (#29570)
a4cb4c6 llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) (#29818)
70849ee common : remove fs_open_ifstream() by using u8path() (#29841)
8d81559 llama : silence unused-result warnings (#29839)
6805ae3 llama : use GGML_ABORT instead of throw (#29840)
a8c9a4e opencl: use sigmoid f16 for bf16 (#29787)
392ded6 SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (#29186)
9e258a6 vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (#28531)
b933289 sycl: large register file for D=512 FA vec kernels (#29062)
c328acc sycl : do not use slow oneDNN reference matmul and fattn (#28985)
4e2713c qwen4exp : optimize mask constructions (#29824)
631109b ggml : add alloc_buffer_n to buffer type interface (#23671)
254b177 ci : fix missing zdnn backend check (#29837)
fb4b273 vulkan: add logging to pipeline compile issues (#29794)
207bdab pyproject : add linux platform marker to uv torch source (#29177)
5fc4f3c hexagon: install rebuilt HTP skels (#29828)
159c651 qwen4exp: fix tests (#29819)
a868c3e hexagon: add q2_k and q3_k quant type support (#29717)
ec7630a CUDA: fix 2 broken Volta FA cases (#29803)
78e2964 llama: refer to segment documentation [no ci] (#29074)
f1cee99 common,rpc : fix cache dir creation through symlinks on buggy libstdc++ (#29816)
68e79bd skill: note about model-specific CLI arguments + testings (#29808)
e358d59 ci: fix Fusion / metal by updating the qwen4exp baseline (#29812)
81e39ad llama : clamp kpool re-pool bound to existing pools (#29805)
dcd387a hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (#29685)
d775ebf server: return HTTP 400 for invalid embedding requests (#29060)
2b36825 convert : write Gemma embedding scale for DFlash drafts (#29802)
42d9581 cuda : route sm70 to the Turing MMVQ nwarps table (#29753)
13b4d71 metal : release temporary private transfer buffers (#29777)
4b1622a webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 (#29358)
869034b llama : fix invalid assert in recurrent memory (#29799)
b56f34a CUDA: Handle compute type for NVFP4 on cublass path (#29173)
c061df1 Qwen4Exp: add MTP (#29761)
66e0c17 llama: fix qwen4exp (#29751)
7677678 CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (#29792)
552f18f mtmd: cap max_image to ubatch for non_causal models (#29773)
5503b04 meta: clear inactive AllReduce shards with FILL, not SCALE (#29793)
def4d40 jinja : skip copying loop scope unless a loop filter needs it (#29776)
32dd62e llama-mmap : avoid a second full-size copy of each tensor with direct-io (#29749)
f11d642 HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (#29572)
3aa0ce9 hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (#29785)
b0aca3c BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (#29640)
b8f96c3 common : add LLM-jp-4.1 Harmony dialect handler (#29681)
3ec4df4 opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (#29698)
db33d3c vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3 (#29734)
7dad6db llama-bench : fix verbosity filter to show GGML_LOG_ERROR (#28229)
2232bc8 metal : use bf16 math for mxfp4 mul-mat (#29770)
79625e0 llama-bench : fix docs (#29464)
66bcc27 docs : refresh CPU ops support matrix (#29666)
10f340d model : re-enable -sm tensor for qwen4exp (#28569)
0c1e570 webgpu: fix SSM_SCAN binding aliasing (#29750)
f7b384c ggml-opencl : replace alloca() with std::vector (#29765)
f872b59 cuda: guard the iq4_nl dequantize row kernel against short rows (#29683)
a4d880f Hexagon: optimize ALLREDUCE with support for safe scatter mode (#29757)
feb9a3d args: fix cli download mmproj arg (#28977)
4453b53 llama : preserve original batch order for speculative decoding layer inputs (#29019)
4f31296 test-llama-archs : toggle causal_attn to catch graph shape changes (#29724)
b016f46 convert : fix LoRA conversion crash for Qwen3.5 V-head reorder (#28324)
81ff93e llama: properly handle KV on training (#28520)
60e9cf7 batch: migrate the rest of examples to llama_batch_ext (#29601)
05af0d2 glm5-next: give dead indexer slots unique scatter rows (#29745)
2149c00 ggml/gguf : fix integer overflow (#29384)
876c75b codeowners : remove former ZenDNN owner (#29747)
b046420 cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast (#29722)
22bdcc4 mimo : support dflash (convert + feature extraction) (#29650)
ca2e203 jinja : support coerced array attributes (#29574)
bdeb855 ggml-et : remove useless alloca() (#29663)
3b3d022 ci : fix Models Backend Check by shortening the hrm_text fixture (#29744)
185103d llama: llama_prefetch_rows (#29599)
2090f60 ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (#29675)
90c908d cpu: accept BF16 in src1 of mul_mat (#28937)
8df332d model-conversion : add --add-bos to run org model script (#29558)
4a096b8 ui : shared model display primitives (#29644)
8664eae ui : model download pipeline (#27959)
4cfb6d1 ui : model memory-fit estimation (#27957)
9b43336 ui : Hugging Face Hub data layer (#27947)
f653250 ui : model id grammar for sidecars, quants and capability parsing (#27946)
fa2bde5 ui : type-safe API types, fetch helpers and download-ready models store plumbing (#29582)
25747b0 openvino: serve GET_ROWS on a weight view from the base Constant (#28381)
db00347 ci : fix Fusion / metal by adding glm5-next to MTL.csv (#29712)
272aad8 musa : define CUDA_ARCH for device passes (#29508)
72db1e0 ci : add models backend check (#29651)
2a53ace SYCL: reduce tensor allreduce sync with pinned host buffers (#29604)
649dcb1 add GLM-5.3-Flash (GLM5-Next) support (#27773)
931351e vendor: update BoringSSL to 0.20260929.0 (#29669)
eae11d2 ggml-zdnn: impl buffer reset, fix memory leaks (#29637)
19e28a2 Hexagon f16 activation ops (#29209)
a6ea155 gguf : reject tensor size that wraps after padding (#26979)
d3954b9 ggml : check row bounds in get_rows_back (#29575)
48de2a1 model : support classifier_pooling for rerankers (#29627)
7fee178 hexagon: optimize concat op (#29673)
6a2743f CUDA: bitonic argsort handles rows wider than one block (#28957)
748d422 ggml-cuda: HIP: optimize packed byte subtraction (__vsubss4 -> __vsub4) (#29478)
cee37ff ci: add zdnn backend build but not test (#29541)
6dbbac4 opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (#29555)
5c200e0 vulkan: Tune GDN kernel, fix Intel performance (#29476)
83dd71f vulkan : Load F32 A matrix 2 at a time when its 2-aligned (#29254)
94a0ae3 vulkan: MOE aware mat_mul_id tile selection (#29182)
da89bb3 ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (#29504)
a3f84fa vocab : keep NORMAL in PLaMo-2 and PLaMo-3 (#29580)
284153e ggml : accumulate f16 dot products in f32 on AVX512-FP16 (#29545)
b5cf8ce ggml : require input tensors to be GGML_OP_NONE (#29647)
e904318 hexagon: add FP32 GELU_ERF and GEGLU_ERF support (#29631)
d280808 common : stop accepting draft tokens at EOG (#29638)
ba0ba54 server : remove the built-in UI's service worker when the UI is not served (#29565)
00af635 common : use fs::path for config dir (#29649)
c85b92c tests : adjust server string regex to also match m2 utlra results (#29648)
31385c9 common : add fs_write_atomic() (#29642)
8019dc5 ggml : collect all input tensors into graph_inputs (#29634)
86ea01d ggml-zdnn: fix 0-row tensor crash (#29636)
18b74ff musa: build the docker image and CI container from the MUSA SDK images (#29624)
c13e04e ggml : speed up model loading (#29598)
c8cda8b ci: remove gpu-rocm keyed directory logs (#28940)
6d78fb0 llama : fix init in several tools/examples (#29632)
18bbc46 metal: FWHT perf optimizations (#29602)
0bc845d vulkan : reuse descriptor sets when bindings are constant (#29280)
139997d chat : fix Muse Glimmer ignoring response_format json_schema with --jinja (#29615)
76a5bc8 common : use fs::path for cache dirs (#29595)
46e17a6 tests : skip pytest workers when PYTEST_WORKERS=1 (#29610)
fc07d78 ci : update the oneAPI toolkit to 2026.1 (#29273)
526c43b mtmd: fix GCC 15 stringop-overflow in decode_embd_batch (#29607)
1c47294 hex-scripts: show trace events smaller than 100nsec in perfetto (#29614)
680a036 server : support typed content (vision/audio/video) input for /v1/embeddings endpoint (#29556)
66e665c vulkan: include functional header (#29597)
57b557c models: pad on the left with ggml_pad_ext (#29567)
14ebbd5 ggml-openvino: mark unaligned batch-stride views unsupported (#29603)
f1ea206 batch: migrate speculative, mtmd and server to batch_ext (#29385)
6c7a87f common : fix HF cache paths on Windows (#29475)
f00a64c webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor (#29471)
d77dd08 tests : refactor test-recurrent-state-rollback (#29426)
6f767fe ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 (#29423)
f916130 ci : ignore more vgpr spills in > 256 DQK fattn kernels (#29571)
03a667a vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain (#29520)
c2a9e16 HIP: fix template skip for DKQ > 256 mfma kernels (#29559)
4364bf7 metal: support left and circular padding in GGML_OP_PAD (#29561)
ed7ac35 context : do not re-reserve the scheduler when toggling causal_attn (#28751)
0c6a6a7 Enables Windows ARM64 build with MSVC cl.exe (#28362)
81ef10e tests : fix ggml init (#29554)
5262471 vulkan: fix wrong results when a mul_mat reads a slice of a larger cache (#28956)
4da6337 server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876)
a97cce8 common : avoid side effects around params parsing (#29537)
136887b common : make string_split throw on invalid values (#29518)
9adc7f4 convert : export YaRN scaling parameters for PLaMo-3 (#29528)
6fd50a4 ci : bump ty to 0.0.84 (#29529)
33c923d jinja : add support for dict builtin (#29477)
c9064dd opencl: refine bin kernel loading condition (#29503)
c829670 sycl: FWHT kernels for block widths above 512 (#29243)
36d7b08 CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 (#26289)
2ebd9ae HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes (#28907)
cea7462 vulkan: fix argsort kernel selection for Adreno (#29469)
da6c28e common : throw instead of abort on grammar without llguidance (#29516)
d7fb90e RPC: use RDMA completion channel to not spin (#29440)
7fb2b08 ci : enable GGML_SCHED_DEBUG_REALLOC=1 for ctest workflows (#29514)
187664b llama-bench : fix OOB access of hf_file (#29515)
85ca3b5 hrm : fix layer placement of z_l_init weight (#29512)
7ac59a6 hexagon: support tiled Q4_0 and Q8_0 GET_ROWS (#29511)
2b129cc hexagon: support for backend sampler (#29502)
9588757 cuda: support Nemotron 3 Puzzle state size 96 for ssm scan (#28717)
694ec23 musa: build the docker images from the PH1 MUSA SDK image (#29481)
6f856c7 cuda: add F16 input to the FWHT (#29096)
fcb3074 server : fix wake_fd warning on Windows (#29479)
2145525 Revert "Change max context length for auto-fitting with unified KV (#28849)" (#29437)
81bc6b8 jinja : implement sameas test (#29448)
86a24a1 jinja : fix compile error (#29468)
08618ff llama : fix K/V and recurrent state cleanup after failed restores (#27530)
a1de614 jinja : support noncall test statements with arg (#29443)
965f897 polished Readme and llama-bench (#28968)
d834d44 ggml-cpu: tiled mul_mat for k-quants (#27851)
9f70b2c opencl: add A8 Q8_0 non-MoE dp4a binary kernel (#29439)
4e74811 hexagon: find software divide calls using binary inspection tool (#29449)
171e884 vendor : update cpp-httplib to 0.58.0 (#29407)
4b1a27f common,rpc : simplify fs_create_directory_with_parents() (#29432)
fcc8915 mtmd: fix mel preprocessor in LFM2 audio (#29403)
a25c986 opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin, kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin (#29401)
e85e15c Fixing the vulkan build issue of legacy GLSLC version that has no cooperativeMatrix API support (#29373) (#29409)
b248f4a gguf-py : ByteLevel processing defaults bos/eos to False (#29422)
d81aef1 gguf-py : TemplateProcessing has final word on add_special_token (#29417)
27b20ba common : extract shared unicode path/string helpers (#29415)
e351231 metal: FWHT kernels for block widths above 512 (#29095)
5a75f14 metal : split fa kernels into per-dtype libraries (#29329)
e9f824d llama : add llama_prec_policy + model-driven W4A4 path (#24364)
d028c69 HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m3 support in 6.2 (#29231)
66963a8 rpc: include nb in the get_alloc_size cache key and floor the result at ggml_nbytes (#29283)
cd74ef6 [SYCL] support sparse FA (#28796)
f9af9be musa: fix PH1 (MTT S5000) operator failures and build issues (#29193)
1ab7e5a CUDA: fuse RMS_NORM + SCALE into one kernel (#29393)
f805c57 llama : fix tensor split for fused qkv with uneven K/V head sizes (#29294)
4de0926 hexagon: add q5_k quant type support (#29123)
ed319fe hexagon: use DMA for contiguous dim1 CONCAT (#29404)
84e76d8 metal : fix graph capture and handle empty graphs (#29390)
cdc0642 metal : optimize sparse FA + clean-up (#29377)
bced459 sync : ggml (#29396)
a02c7f5 hexagon: handle multi-sequence in concat_2d (#29344)
5cf3a35 llama-grammar: fix numeric truncation for token_id parsing (#29382)
07fc586 hexagon: dynamic quantizer improvements (#29395)
97a418b hexagon: support I32 CPY and CONT (#29379)
a72e04a cuda : add F16 kernel support for CONV_2D_DW (#29064)
8212c78 test: flush status (#28352)
945064f ui : fix missing svg use and animation elements in preview and download (#28962)
fc343a8 llama: add llama_batch_ext (#24669)
308883b server : change default pytest workers to 4 (#29376)
70596c4 ci : use hf-jobs-cpu-performance, disable pytest workers (#29369)
70c4e15 vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (#27952)
6b790a9 vulkan: handle misalignment in conv_2d and conv_3d (#29365)
3423f94 vulkan: tune KHR cooperative matrix support for Adreno GPUs (#29328)
53ed051 cuda : add conv3d with implicit GEMM (#29137)
f830688 model : add Ling 3.0 VL support (#29151)
2b70583 server,common : fix the GCC 12 stringop-overread false positive (again) (#29325)
4c5957c test-save-load-state : print a per-model results table in --models mode (#29316)
9710a32 hexagon: reject MUL_MAT_ID when src1 precision is F32 (#29348)
013b31c scripts : make-release-desc - link previous release in changelog title (#29336)
bd4f514 convert : allow vision target for DFlash/Dspark (#29339)
b9ae43a server: allow preset to set log file (#29334)
d2e5458 tests: add -b/--backend option to test-llama-archs for testing a specific backend (#27372)
6e60f35 ci : use hf-jobs-cpu-xl runner in server sanitize workflow (#29297)
fee39dd opencl: add A8 Q6_K non-MoE dp4a binary kernel (#29057)