github ggml-org/llama.cpp v0.6.0

one hour ago

Overview

llama.cpp v0.6.0 introduces the new llama_batch_ext extended batch API (with llama_process) for mixed token/embedding inputs and MTP/deepstack state embeddings, adds support for the GLM-5.3-Flash (GLM5-Next) 320B hybrid model, the Clef decision model (text and vision) and MTP speculative decoding for Qwen4Exp, ships a new /v1/systemone server API for decision models (laya, julia-1, lev, openjev, kev, nimble), overhauls the Web UI with a Hugging Face Hub data layer and model download pipeline, adds a Metal tensor-API flash attention kernel for F16 KV, sparse flash attention for quantized K/V on Vulkan, and updates ggml to v0.26.0.

Highlights

  • New llama_batch_ext extended batch API with llama_process(), supporting mixed token/embedding batches and per-token "state" embeddings for MTP and deepstack models #24669
  • New models: GLM-5.3-Flash (GLM5-Next), a 320B text+vision hybrid model #27773, and the Clef decision model, fully supported with both text and vision #29831 #29969
  • Qwen4Exp: high-quality support is now available, with MTP speculative decoding (~1.5x decode speedup on DGX Spark) and various correctness fixes #29761 #29751
  • llama and server: new /v1/systemone API supporting five decision models - laya, julia-1, lev, openjev (+vision), kev #29818
  • Metal: new tensor API flash attention kernel for F16 KV #29570
  • Metal: new few-row MMA mat-mul kernels for speculative and batched decoding, up to ~3x faster mat-mul on Apple GPUs #29869
  • New llama_prefetch_rows() using MADVISE-based prefetching of PLE tensors in Qwen4Exp and Gemma4 #29599

API changes

  • include/llama.h: new llama_batch_ext batch API with llama_embd, llama_process() and llama_process_type #24669, new llama_get_causal_attn() #28876, session formats bumped to LLAMA_SESSION_VERSION 11 and LLAMA_STATE_SEQ_VERSION 4
  • include/llama-cpp.h: added llama_batch_ext_ptr and deleter for the new extended batch API #24669
  • tools/mtmd/mtmd.h: mtmd_get_memory_usage() now returns an mtmd_memory_usage struct with image_max_tokens and use_non_causal #29773
  • tools/server: new /v1/systemone endpoint for decision models #29818 and /v1/embeddings now accepts typed vision/audio/video content #29556

New models

  • GLM-5.3-Flash (GLM5-Next): 320B KDA/DSA hybrid text+vision model with mHC and MoE #27773
  • Clef decision model, fully supported with both text and vision #29831 #29969
  • Ling 3.0 VL, folded into the BailingMoeV3 architecture #29151
  • Nimble decision model #29844
  • Registered Lfm2BidirectionalForMaskedLM for LFM2.5-Encoder-230M/350M #29862
  • Added classifier_pooling support for rerankers #29627

Core changes

  • Migrated examples, speculative decoding, mtmd and server to the new llama_batch_ext API #29385 #29601; batches now accept both embd and raw tokens #29622
  • Added llama_prec_policy and a model-driven W4A4 (NVFP4/MXFP4) mul_mat path #24364
  • Qwen4Exp: halved indexer score memory #29825, optimized mask constructions #29824, re-enabled the -sm tensor #28569; GLM5-Next: unique scatter rows for dead indexer slots #29745
  • KV cache: fixed restoring mismatched KV cache rotation #28498, fixed K/V and recurrent state cleanup after failed restores #27530 and an invalid assert in recurrent memory #29799
  • Speculative decoding: probabilistic sampling for simple draft and MTP #27694, fixed n-gram drafts rejected at temp > 0 after truncation #29924, preserved original batch order for layer inputs #29019, stop accepting draft tokens at EOG #29638
  • k-pool models: fixed unexpected graph reallocation #29958 and clamped kpool re-pool bound to existing pools #29805
  • Fixed tensor split for fused qkv with uneven K/V head sizes #29294, gather recurrent states once so the reserve covers every split #29856, and properly handle KV on training #28520
  • Context: do not re-reserve the scheduler when toggling causal_attn #28751
  • DFlash drafts: write Gemma embedding scale during conversion #29802 and add dflash support for MiMo #29650

Multi-modality changes

  • Cap max_image to n_ubatch for non-causal models #29773
  • Fixed the mel preprocessor in LFM2 audio #29403
  • Migrated input processing to llama_batch_ext #29385

Server changes

  • New /v1/systemone decision-model API with dedicated decision pipeline #29818, extended to the nimble decision model #29844
  • Support vision input for Clef #29969
  • GET /v1/models and GET /models now report model input/output modalities in a new architecture object #29987
  • /v1/embeddings: accept typed content (vision/audio/video) input #29556 and return HTTP 400 for invalid embedding requests #29060
  • Allow RANK pooling batch splitting for causal LLM rerankers (Qwen3, Qwen3-VL) #28876
  • Reject partial media truncation #24076
  • Logging: self-contained colors and split child commands from logs in router mode #29895, allow preset to set log file #29334, fixed dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list #29938
  • Fixed laya abort by limiting n_batch to n_ubatch #29903
  • Remove the built-in UI's service worker when the UI is not served #29565

UI changes

  • New model download pipeline #27959, Hugging Face Hub data layer #27947, model memory-fit estimation #27957 and model id grammar for sidecars, quants and capability parsing #27946
  • Type-safe API types, fetch helpers and download-ready models store plumbing #29582
  • Shared model display primitives #29644
  • Fixed missing svg use and animation elements in preview and download #28962
  • Use toLocaleString() formatting consistently across chat message statistics #27990

ggml changes

  • ggml updated to v0.26.0: a new alloc_buffer_n/get_alloc_size_n buffer allocation API, sparse flash attention kernels on SYCL, Vulkan and Metal, and major lightning indexer improvements (halved score memory, tiling, MUSA support). The CPU backend gains BF16 ops and a tiled k-quant mul_mat, CUDA gains a model-driven W4A4 (NVFP4/MXFP4) mul_mat path plus MMVQ shared-expert fusion, and the Hexagon backend adds a sampler and more quant types. Model loading is faster with stricter GGUF size validation, Windows ARM64 MSVC builds are enabled, and WebGPU/OpenVINO/OpenCL/SYCL pick up numerous new kernels, ops and fixes.

Assets

Nightly build: b11429

More info

Changelog since v0.5.0

d812350 llama.cpp : bump version to 0.6.0 (#29997)
4d60b4d common, server : report model input/output modalities in GET /models (#29987)
c06f841 sync : ggml
f05c8b2 ggml : bump version to 0.26.0 (ggml/1652)
e117148 CUDA: make the alloc_deps check batch independent (#29986)
6c59c40 vulkan: fix Flash Attention shmem write out of bounds (#29988)
3c9e747 vulkan: revert mul_mat_id tile selection PR #29182 (#29936)
b809b88 cuda: use the vector lightning indexer kernel on MUSA (#29990)
994e8f2 ci : add "Require Docker" flag to make-release workflow (#29989)
9d853bb webui: Use toLocaleString() format consistently across chat message statistics (#27990)
8f9ae20 ci : disable failing test on virtual Metal device (#29993)
9871df5 server: support vision input for Clef (#29969)
8b2fbaf CUDA: Optimize accumulation in mmq for NVFP4 type (#29857)
2ed93db ci : disable unused qemu in docker build (#29984)
8e16421 server: reject partial media truncation (#24076)
806eee9 vulkan: fix stale prealloc_y reuse across flash attention and soft_max (#29591)
b3daa07 vulkan: sparse flash attention for quantized K/V (#29639)
c173a53 llama : fix unexpected graph reallocation in the k-pool models (#29958)
2107910 kv-cache: fix restoring mismatched KV cache rotation by saving exact rotation metadata (#28498)
e5983d6 ci : winget urls must be separate strings (#29978)
4ca6b76 ci : fix docker workflow permissions (#29979)
9f12cd4 ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (#28479)
ebe18be vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON (#29912)
8216c84 webgpu: add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 (#29483)
1b43d31 cuda: tile the lightning indexer over keys and tokens for 4 heads (#29901)
a3a1c47 metal : few-row MMA mat-mul (#29869)
9d3aba6 CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size (#29633)
d89651a CUDA: prefer whole-tile FlashAttention scheduling for efficient two-stage kernels (#29435)
a7fb71f log, server: self contained colors, split child commands from logs in router mode (#29895)
0bb496d llama: support both embd + raw tokens in batch (#29622)
2ca15f5 CUDA: refactor swizzling code (#29612)
a7b94df ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (#29806)
0eb6d9a cuda : move neu_padded to where it is used (#29940)
2e7c58c ci : windows llvm build requires ninja multi-config (#29959)
7f2dd88 ci : add windows arm64 vulkan release (#29954)
bf79dbb AGENTS.md : revamp (#29656)
dbe4c3e chat-peg-parser : clear current_tool when pending_tool_call is reset (#29942)
46847e6 ci : set default permissions (#29945)
2bc5635 cuda : move blocks_per_col to where it is used (#29939)
dd26678 CUDA: fix MMQ memory fault if n_expert >> n_ubatch (#29941)
16c163d vulkan: fix rdna4 mat_vec tuning (#29934)
0504396 imatrix: calculate activation-based statistics for new format (GGUF) imatrices (#14891)
8330e96 spec : fix n-gram drafts rejected at temp > 0 after truncation (#29924)
6716df6 common : prepare load_from_models_dir() for path conversion (#29674)
bf9a0cc server : fix dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list (#29938)
0faee50 ci : pushing tag needs deploy key (#29937)
f98b31c ci : improve release flow (#29913)
11fe021 webgpu: add f16 support to fill/set_rows (#29897)
836d571 mtmd : fix deprecated strdup warning on Windows (#29863)
eec18f5 vendor : update cpp-httplib to 0.59.0 (#29886)
1537a0a server : fix laya abort by limiting n_batch to n_ubatch (#29903)
edd6e2b common : add common_is_tty() helper and fix deprecated warnings on Windows (#29860)
9bf55f4 chat : honor json_schema in Ling 3.0 parser (#29813)
a55e952 ci: fix flaky ADD_ADD f16 by using the fused ADD tolerance (#29904)
436f6f8 graph: gather the recurrent states once so the reserve covers every split (#29856)
b92761a ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (#29852)
cb7934c model : Add LFM2.5-Encoder-350M and LFM2.5-Encoder-230M (#29862)
889edf4 qwen4exp : halve the indexer score memory (#29825)
99b9548 model: add support for clef decision model (text-only) (#29831)
bed0a85 CUDA: fuse shared experts into MMVQ (#29184)
4ebdf2c ci : use t4-medium for cuda jobs (#29842)
1fb7ef3 spec : add probabilistic sampling for simple draft and MTP (#27694)
134b2bb ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (#27663)
2923cf2 ggml-quants : avoid invalid rounding in qkx3 scale search (#29817)
dd4c286 ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (#27096)
46ca246 model: support nimble decision model (#29844)
d8fbd25 readme : add cmd install commands (#29850)
926862e metal : add tensor API flash attention kernel for F16 KV (#29570)
a4cb4c6 llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) (#29818)
70849ee common : remove fs_open_ifstream() by using u8path() (#29841)
8d81559 llama : silence unused-result warnings (#29839)
6805ae3 llama : use GGML_ABORT instead of throw (#29840)
a8c9a4e opencl: use sigmoid f16 for bf16 (#29787)
392ded6 SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (#29186)
9e258a6 vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (#28531)
b933289 sycl: large register file for D=512 FA vec kernels (#29062)
c328acc sycl : do not use slow oneDNN reference matmul and fattn (#28985)
4e2713c qwen4exp : optimize mask constructions (#29824)
631109b ggml : add alloc_buffer_n to buffer type interface (#23671)
254b177 ci : fix missing zdnn backend check (#29837)
fb4b273 vulkan: add logging to pipeline compile issues (#29794)
207bdab pyproject : add linux platform marker to uv torch source (#29177)
5fc4f3c hexagon: install rebuilt HTP skels (#29828)
159c651 qwen4exp: fix tests (#29819)
a868c3e hexagon: add q2_k and q3_k quant type support (#29717)
ec7630a CUDA: fix 2 broken Volta FA cases (#29803)
78e2964 llama: refer to segment documentation [no ci] (#29074)
f1cee99 common,rpc : fix cache dir creation through symlinks on buggy libstdc++ (#29816)
68e79bd skill: note about model-specific CLI arguments + testings (#29808)
e358d59 ci: fix Fusion / metal by updating the qwen4exp baseline (#29812)
81e39ad llama : clamp kpool re-pool bound to existing pools (#29805)
dcd387a hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (#29685)
d775ebf server: return HTTP 400 for invalid embedding requests (#29060)
2b36825 convert : write Gemma embedding scale for DFlash drafts (#29802)
42d9581 cuda : route sm70 to the Turing MMVQ nwarps table (#29753)
13b4d71 metal : release temporary private transfer buffers (#29777)
4b1622a webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 (#29358)
869034b llama : fix invalid assert in recurrent memory (#29799)
b56f34a CUDA: Handle compute type for NVFP4 on cublass path (#29173)
c061df1 Qwen4Exp: add MTP (#29761)
66e0c17 llama: fix qwen4exp (#29751)
7677678 CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (#29792)
552f18f mtmd: cap max_image to ubatch for non_causal models (#29773)
5503b04 meta: clear inactive AllReduce shards with FILL, not SCALE (#29793)
def4d40 jinja : skip copying loop scope unless a loop filter needs it (#29776)
32dd62e llama-mmap : avoid a second full-size copy of each tensor with direct-io (#29749)
f11d642 HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (#29572)
3aa0ce9 hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (#29785)
b0aca3c BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (#29640)
b8f96c3 common : add LLM-jp-4.1 Harmony dialect handler (#29681)
3ec4df4 opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (#29698)
db33d3c vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3 (#29734)
7dad6db llama-bench : fix verbosity filter to show GGML_LOG_ERROR (#28229)
2232bc8 metal : use bf16 math for mxfp4 mul-mat (#29770)
79625e0 llama-bench : fix docs (#29464)
66bcc27 docs : refresh CPU ops support matrix (#29666)
10f340d model : re-enable -sm tensor for qwen4exp (#28569)
0c1e570 webgpu: fix SSM_SCAN binding aliasing (#29750)
f7b384c ggml-opencl : replace alloca() with std::vector (#29765)
f872b59 cuda: guard the iq4_nl dequantize row kernel against short rows (#29683)
a4d880f Hexagon: optimize ALLREDUCE with support for safe scatter mode (#29757)
feb9a3d args: fix cli download mmproj arg (#28977)
4453b53 llama : preserve original batch order for speculative decoding layer inputs (#29019)
4f31296 test-llama-archs : toggle causal_attn to catch graph shape changes (#29724)
b016f46 convert : fix LoRA conversion crash for Qwen3.5 V-head reorder (#28324)
81ff93e llama: properly handle KV on training (#28520)
60e9cf7 batch: migrate the rest of examples to llama_batch_ext (#29601)
05af0d2 glm5-next: give dead indexer slots unique scatter rows (#29745)
2149c00 ggml/gguf : fix integer overflow (#29384)
876c75b codeowners : remove former ZenDNN owner (#29747)
b046420 cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast (#29722)
22bdcc4 mimo : support dflash (convert + feature extraction) (#29650)
ca2e203 jinja : support coerced array attributes (#29574)
bdeb855 ggml-et : remove useless alloca() (#29663)
3b3d022 ci : fix Models Backend Check by shortening the hrm_text fixture (#29744)
185103d llama: llama_prefetch_rows (#29599)
2090f60 ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (#29675)
90c908d cpu: accept BF16 in src1 of mul_mat (#28937)
8df332d model-conversion : add --add-bos to run org model script (#29558)
4a096b8 ui : shared model display primitives (#29644)
8664eae ui : model download pipeline (#27959)
4cfb6d1 ui : model memory-fit estimation (#27957)
9b43336 ui : Hugging Face Hub data layer (#27947)
f653250 ui : model id grammar for sidecars, quants and capability parsing (#27946)
fa2bde5 ui : type-safe API types, fetch helpers and download-ready models store plumbing (#29582)
25747b0 openvino: serve GET_ROWS on a weight view from the base Constant (#28381)
db00347 ci : fix Fusion / metal by adding glm5-next to MTL.csv (#29712)
272aad8 musa : define CUDA_ARCH for device passes (#29508)
72db1e0 ci : add models backend check (#29651)
2a53ace SYCL: reduce tensor allreduce sync with pinned host buffers (#29604)
649dcb1 add GLM-5.3-Flash (GLM5-Next) support (#27773)
931351e vendor: update BoringSSL to 0.20260929.0 (#29669)
eae11d2 ggml-zdnn: impl buffer reset, fix memory leaks (#29637)
19e28a2 Hexagon f16 activation ops (#29209)
a6ea155 gguf : reject tensor size that wraps after padding (#26979)
d3954b9 ggml : check row bounds in get_rows_back (#29575)
48de2a1 model : support classifier_pooling for rerankers (#29627)
7fee178 hexagon: optimize concat op (#29673)
6a2743f CUDA: bitonic argsort handles rows wider than one block (#28957)
748d422 ggml-cuda: HIP: optimize packed byte subtraction (__vsubss4 -> __vsub4) (#29478)
cee37ff ci: add zdnn backend build but not test (#29541)
6dbbac4 opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (#29555)
5c200e0 vulkan: Tune GDN kernel, fix Intel performance (#29476)
83dd71f vulkan : Load F32 A matrix 2 at a time when its 2-aligned (#29254)
94a0ae3 vulkan: MOE aware mat_mul_id tile selection (#29182)
da89bb3 ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (#29504)
a3f84fa vocab : keep NORMAL in PLaMo-2 and PLaMo-3 (#29580)
284153e ggml : accumulate f16 dot products in f32 on AVX512-FP16 (#29545)
b5cf8ce ggml : require input tensors to be GGML_OP_NONE (#29647)
e904318 hexagon: add FP32 GELU_ERF and GEGLU_ERF support (#29631)
d280808 common : stop accepting draft tokens at EOG (#29638)
ba0ba54 server : remove the built-in UI's service worker when the UI is not served (#29565)
00af635 common : use fs::path for config dir (#29649)
c85b92c tests : adjust server string regex to also match m2 utlra results (#29648)
31385c9 common : add fs_write_atomic() (#29642)
8019dc5 ggml : collect all input tensors into graph_inputs (#29634)
86ea01d ggml-zdnn: fix 0-row tensor crash (#29636)
18b74ff musa: build the docker image and CI container from the MUSA SDK images (#29624)
c13e04e ggml : speed up model loading (#29598)
c8cda8b ci: remove gpu-rocm keyed directory logs (#28940)
6d78fb0 llama : fix init in several tools/examples (#29632)
18bbc46 metal: FWHT perf optimizations (#29602)
0bc845d vulkan : reuse descriptor sets when bindings are constant (#29280)
139997d chat : fix Muse Glimmer ignoring response_format json_schema with --jinja (#29615)
76a5bc8 common : use fs::path for cache dirs (#29595)
46e17a6 tests : skip pytest workers when PYTEST_WORKERS=1 (#29610)
fc07d78 ci : update the oneAPI toolkit to 2026.1 (#29273)
526c43b mtmd: fix GCC 15 stringop-overflow in decode_embd_batch (#29607)
1c47294 hex-scripts: show trace events smaller than 100nsec in perfetto (#29614)
680a036 server : support typed content (vision/audio/video) input for /v1/embeddings endpoint (#29556)
66e665c vulkan: include functional header (#29597)
57b557c models: pad on the left with ggml_pad_ext (#29567)
14ebbd5 ggml-openvino: mark unaligned batch-stride views unsupported (#29603)
f1ea206 batch: migrate speculative, mtmd and server to batch_ext (#29385)
6c7a87f common : fix HF cache paths on Windows (#29475)
f00a64c webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor (#29471)
d77dd08 tests : refactor test-recurrent-state-rollback (#29426)
6f767fe ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 (#29423)
f916130 ci : ignore more vgpr spills in > 256 DQK fattn kernels (#29571)
03a667a vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain (#29520)
c2a9e16 HIP: fix template skip for DKQ > 256 mfma kernels (#29559)
4364bf7 metal: support left and circular padding in GGML_OP_PAD (#29561)
ed7ac35 context : do not re-reserve the scheduler when toggling causal_attn (#28751)
0c6a6a7 Enables Windows ARM64 build with MSVC cl.exe (#28362)
81ef10e tests : fix ggml init (#29554)
5262471 vulkan: fix wrong results when a mul_mat reads a slice of a larger cache (#28956)
4da6337 server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876)
a97cce8 common : avoid side effects around params parsing (#29537)
136887b common : make string_split throw on invalid values (#29518)
9adc7f4 convert : export YaRN scaling parameters for PLaMo-3 (#29528)
6fd50a4 ci : bump ty to 0.0.84 (#29529)
33c923d jinja : add support for dict builtin (#29477)
c9064dd opencl: refine bin kernel loading condition (#29503)
c829670 sycl: FWHT kernels for block widths above 512 (#29243)
36d7b08 CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 (#26289)
2ebd9ae HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes (#28907)
cea7462 vulkan: fix argsort kernel selection for Adreno (#29469)
da6c28e common : throw instead of abort on grammar without llguidance (#29516)
d7fb90e RPC: use RDMA completion channel to not spin (#29440)
7fb2b08 ci : enable GGML_SCHED_DEBUG_REALLOC=1 for ctest workflows (#29514)
187664b llama-bench : fix OOB access of hf_file (#29515)
85ca3b5 hrm : fix layer placement of z_l_init weight (#29512)
7ac59a6 hexagon: support tiled Q4_0 and Q8_0 GET_ROWS (#29511)
2b129cc hexagon: support for backend sampler (#29502)
9588757 cuda: support Nemotron 3 Puzzle state size 96 for ssm scan (#28717)
694ec23 musa: build the docker images from the PH1 MUSA SDK image (#29481)
6f856c7 cuda: add F16 input to the FWHT (#29096)
fcb3074 server : fix wake_fd warning on Windows (#29479)
2145525 Revert "Change max context length for auto-fitting with unified KV (#28849)" (#29437)
81bc6b8 jinja : implement sameas test (#29448)
86a24a1 jinja : fix compile error (#29468)
08618ff llama : fix K/V and recurrent state cleanup after failed restores (#27530)
a1de614 jinja : support noncall test statements with arg (#29443)
965f897 polished Readme and llama-bench (#28968)
d834d44 ggml-cpu: tiled mul_mat for k-quants (#27851)
9f70b2c opencl: add A8 Q8_0 non-MoE dp4a binary kernel (#29439)
4e74811 hexagon: find software divide calls using binary inspection tool (#29449)
171e884 vendor : update cpp-httplib to 0.58.0 (#29407)
4b1a27f common,rpc : simplify fs_create_directory_with_parents() (#29432)
fcc8915 mtmd: fix mel preprocessor in LFM2 audio (#29403)
a25c986 opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin, kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin (#29401)
e85e15c Fixing the vulkan build issue of legacy GLSLC version that has no cooperativeMatrix API support (#29373) (#29409)
b248f4a gguf-py : ByteLevel processing defaults bos/eos to False (#29422)
d81aef1 gguf-py : TemplateProcessing has final word on add_special_token (#29417)
27b20ba common : extract shared unicode path/string helpers (#29415)
e351231 metal: FWHT kernels for block widths above 512 (#29095)
5a75f14 metal : split fa kernels into per-dtype libraries (#29329)
e9f824d llama : add llama_prec_policy + model-driven W4A4 path (#24364)
d028c69 HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m3 support in 6.2 (#29231)
66963a8 rpc: include nb in the get_alloc_size cache key and floor the result at ggml_nbytes (#29283)
cd74ef6 [SYCL] support sparse FA (#28796)
f9af9be musa: fix PH1 (MTT S5000) operator failures and build issues (#29193)
1ab7e5a CUDA: fuse RMS_NORM + SCALE into one kernel (#29393)
f805c57 llama : fix tensor split for fused qkv with uneven K/V head sizes (#29294)
4de0926 hexagon: add q5_k quant type support (#29123)
ed319fe hexagon: use DMA for contiguous dim1 CONCAT (#29404)
84e76d8 metal : fix graph capture and handle empty graphs (#29390)
cdc0642 metal : optimize sparse FA + clean-up (#29377)
bced459 sync : ggml (#29396)
a02c7f5 hexagon: handle multi-sequence in concat_2d (#29344)
5cf3a35 llama-grammar: fix numeric truncation for token_id parsing (#29382)
07fc586 hexagon: dynamic quantizer improvements (#29395)
97a418b hexagon: support I32 CPY and CONT (#29379)
a72e04a cuda : add F16 kernel support for CONV_2D_DW (#29064)
8212c78 test: flush status (#28352)
945064f ui : fix missing svg use and animation elements in preview and download (#28962)
fc343a8 llama: add llama_batch_ext (#24669)
308883b server : change default pytest workers to 4 (#29376)
70596c4 ci : use hf-jobs-cpu-performance, disable pytest workers (#29369)
70c4e15 vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (#27952)
6b790a9 vulkan: handle misalignment in conv_2d and conv_3d (#29365)
3423f94 vulkan: tune KHR cooperative matrix support for Adreno GPUs (#29328)
53ed051 cuda : add conv3d with implicit GEMM (#29137)
f830688 model : add Ling 3.0 VL support (#29151)
2b70583 server,common : fix the GCC 12 stringop-overread false positive (again) (#29325)
4c5957c test-save-load-state : print a per-model results table in --models mode (#29316)
9710a32 hexagon: reject MUL_MAT_ID when src1 precision is F32 (#29348)
013b31c scripts : make-release-desc - link previous release in changelog title (#29336)
bd4f514 convert : allow vision target for DFlash/Dspark (#29339)
b9ae43a server: allow preset to set log file (#29334)
d2e5458 tests: add -b/--backend option to test-llama-archs for testing a specific backend (#27372)
6e60f35 ci : use hf-jobs-cpu-xl runner in server sanitize workflow (#29297)
fee39dd opencl: add A8 Q6_K non-MoE dp4a binary kernel (#29057)

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.