github ggml-org/whisper.cpp v1.9.5

5 hours ago

Overview

New version has been released.

Nightly build: b5454
More info: dist : releases and versioning of ggml-org projects

Changelog since v1.9.4

d1be6fd whisper : bump version to 1.9.5 (#4102)
4afec37 talk-llama : update llama.cpp to v0.6.0
3d451ab sync : ggml
47f1c7d ggml : bump version to 0.26.0 (ggml/1652)
e71b784 CUDA: make the alloc_deps check batch independent (llama/29986)
cdc8730 vulkan: fix Flash Attention shmem write out of bounds (llama/29988)
563a9f7 vulkan: revert mul_mat_id tile selection PR #29182 (llama/29936)
ebb8a17 cuda: use the vector lightning indexer kernel on MUSA (llama/29990)
8c6ca4a CUDA: Optimize accumulation in mmq for NVFP4 type (llama/29857)
82fa9df vulkan: fix stale prealloc_y reuse across flash attention and soft_max (llama/29591)
f691542 vulkan: sparse flash attention for quantized K/V (llama/29639)
d477de3 llama : fix unexpected graph reallocation in the k-pool models (llama/29958)
95283ca ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (llama/28479)
155c5c2 vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON (llama/29912)
7091d4a webgpu: add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 (llama/29483)
4cc62b8 cuda: tile the lightning indexer over keys and tokens for 4 heads (llama/29901)
7e576cb metal : few-row MMA mat-mul (llama/29869)
91da470 CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size (llama/29633)
8bb85d9 CUDA: prefer whole-tile FlashAttention scheduling for efficient two-stage kernels (llama/29435)
ce6552e CUDA: refactor swizzling code (llama/29612)
bfe3dc6 ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (llama/29806)
25453c0 cuda : move neu_padded to where it is used (llama/29940)
7914776 cuda : move blocks_per_col to where it is used (llama/29939)
6421d2c CUDA: fix MMQ memory fault if n_expert >> n_ubatch (llama/29941)
05d4d89 vulkan: fix rdna4 mat_vec tuning (llama/29934)
51bfbd6 webgpu: add f16 support to fill/set_rows (llama/29897)
f2aa80a ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (llama/29852)
7f4a67a qwen4exp : halve the indexer score memory (llama/29825)
cf5d9e3 CUDA: fuse shared experts into MMVQ (llama/29184)
d7682b5 ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (llama/27663)
6b70471 ggml-quants : avoid invalid rounding in qkx3 scale search (llama/29817)
0b35d1f ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (llama/27096)
d0dc436 metal : add tensor API flash attention kernel for F16 KV (llama/29570)
cca8f72 opencl: use sigmoid f16 for bf16 (llama/29787)
cd1bee5 SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (llama/29186)
95a05ed vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (llama/28531)
c4051a3 sycl: large register file for D=512 FA vec kernels (llama/29062)
ae92605 sycl : do not use slow oneDNN reference matmul and fattn (llama/28985)
9c968f7 qwen4exp : optimize mask constructions (llama/29824)
4435763 ggml : add alloc_buffer_n to buffer type interface (llama/23671)
8291ab8 vulkan: add logging to pipeline compile issues (llama/29794)
eaadab3 hexagon: install rebuilt HTP skels (llama/29828)
0295ef6 hexagon: add q2_k and q3_k quant type support (llama/29717)
396f68d CUDA: fix 2 broken Volta FA cases (llama/29803)
3a3598c llama: refer to segment documentation [no ci] (llama/29074)
15229ac hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (llama/29685)
aa53800 cuda : route sm70 to the Turing MMVQ nwarps table (llama/29753)
81ca4f8 metal : release temporary private transfer buffers (llama/29777)
88948bd webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 (llama/29358)
10872be CUDA: Handle compute type for NVFP4 on cublass path (llama/29173)
3b68f90 CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (llama/29792)
56500f4 meta: clear inactive AllReduce shards with FILL, not SCALE (llama/29793)
817294a HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (llama/29572)
b4a085a hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (llama/29785)
0ecc317 BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (llama/29640)
67ad86d opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (llama/29698)
61ea502 metal : use bf16 math for mxfp4 mul-mat (llama/29770)
d371e37 webgpu: fix SSM_SCAN binding aliasing (llama/29750)
c1c1390 ggml-opencl : replace alloca() with std::vector (llama/29765)
bffb6e4 cuda: guard the iq4_nl dequantize row kernel against short rows (llama/29683)
6773ef8 Hexagon: optimize ALLREDUCE with support for safe scatter mode (llama/29757)
02ea0c2 ggml/gguf : fix integer overflow (llama/29384)
067bb06 ggml-et : remove useless alloca() (llama/29663)
fd66b6e ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (llama/29675)
c1b3fb1 cpu: accept BF16 in src1 of mul_mat (llama/28937)
d0af734 openvino: serve GET_ROWS on a weight view from the base Constant (llama/28381)
1cbf7e7 musa : define CUDA_ARCH for device passes (llama/29508)
d03bfd7 SYCL: reduce tensor allreduce sync with pinned host buffers (llama/29604)
66e8b95 ggml-zdnn: impl buffer reset, fix memory leaks (llama/29637)
b8051db Hexagon f16 activation ops (llama/29209)
6e48a36 gguf : reject tensor size that wraps after padding (llama/26979)
1442399 ggml : check row bounds in get_rows_back (llama/29575)
8d2a2cb hexagon: optimize concat op (llama/29673)
45093fc CUDA: bitonic argsort handles rows wider than one block (llama/28957)
482baac ggml-cuda: HIP: optimize packed byte subtraction (__vsubss4 -> __vsub4) (llama/29478)
3b86cab ci: add zdnn backend build but not test (llama/29541)
9f12c9f opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (llama/29555)
5439f3d vulkan: Tune GDN kernel, fix Intel performance (llama/29476)
6c07d82 vulkan : Load F32 A matrix 2 at a time when its 2-aligned (llama/29254)
8cc3ca4 vulkan: MOE aware mat_mul_id tile selection (llama/29182)
deae6f3 ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (llama/29504)
b97d419 ggml : accumulate f16 dot products in f32 on AVX512-FP16 (llama/29545)
4d4817e ggml : require input tensors to be GGML_OP_NONE (llama/29647)
5e1914a hexagon: add FP32 GELU_ERF and GEGLU_ERF support (llama/29631)
763b67f ggml : collect all input tensors into graph_inputs (llama/29634)
8478129 ggml-zdnn: fix 0-row tensor crash (llama/29636)
f449de9 ggml : speed up model loading (llama/29598)
b30ef05 metal: FWHT perf optimizations (llama/29602)
eecee67 vulkan : reuse descriptor sets when bindings are constant (llama/29280)
9769010 vulkan: include functional header (llama/29597)
830dc66 ggml-openvino: mark unaligned batch-stride views unsupported (llama/29603)
81b951e webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor (llama/29471)
ee86a15 tests : refactor test-recurrent-state-rollback (llama/29426)
83fd155 ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 (llama/29423)
65449db vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain (llama/29520)
0c17f7d HIP: fix template skip for DKQ > 256 mfma kernels (llama/29559)
0e53751 metal: support left and circular padding in GGML_OP_PAD (llama/29561)
b7b8495 Enables Windows ARM64 build with MSVC cl.exe (llama/28362)
85f6926 vulkan: fix wrong results when a mul_mat reads a slice of a larger cache (llama/28956)
51db475 opencl: refine bin kernel loading condition (llama/29503)
55a695d sycl: FWHT kernels for block widths above 512 (llama/29243)
846af52 CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 (llama/26289)
743f1ad HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes (llama/28907)
14d1aa7 vulkan: fix argsort kernel selection for Adreno (llama/29469)
24cf265 RPC: use RDMA completion channel to not spin (llama/29440)
7eea518 hexagon: support tiled Q4_0 and Q8_0 GET_ROWS (llama/29511)
fe06027 hexagon: support for backend sampler (llama/29502)
fc1ebfa cuda: support Nemotron 3 Puzzle state size 96 for ssm scan (llama/28717)
dfe8fbd cuda: add F16 input to the FWHT (llama/29096)
447a775 ggml-cpu: tiled mul_mat for k-quants (llama/27851)
bf1d787 opencl: add A8 Q8_0 non-MoE dp4a binary kernel (llama/29439)
9731f6f hexagon: find software divide calls using binary inspection tool (llama/29449)
2d91497 opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin, kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin (llama/29401)
669188e Fixing the vulkan build issue of legacy GLSLC version that has no cooperativeMatrix API support (ggml-org/llama.cpp#29373) (llama/29409)
9a91942 metal: FWHT kernels for block widths above 512 (llama/29095)
331d3c8 sync : ggml
4b0998a metal : split fa kernels into per-dtype libraries (llama/29329)
bad9b54 sync : ggml
afc10c2 llama : add llama_prec_policy + model-driven W4A4 path (llama/24364)
d7e83b1 HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m3 support in 6.2 (llama/29231)
26bfdc6 rpc: include nb in the get_alloc_size cache key and floor the result at ggml_nbytes (llama/29283)
93c4308 support sparse FA (llama/28796)
f7d5bf3 musa: fix PH1 (MTT S5000) operator failures and build issues (llama/29193)
40c872b CUDA: fuse RMS_NORM + SCALE into one kernel (llama/29393)
0458c9e hexagon: add q5_k quant type support (llama/29123)
eea9aa6 hexagon: use DMA for contiguous dim1 CONCAT (llama/29404)
4550b9d metal : fix graph capture and handle empty graphs (llama/29390)
e1b93a5 metal : optimize sparse FA + clean-up (llama/29377)
61032a2 hexagon: handle multi-sequence in concat_2d (llama/29344)
845d1c0 hexagon: dynamic quantizer improvements (llama/29395)
e5a8877 hexagon: support I32 CPY and CONT (llama/29379)
5a9a3c0 cuda : add F16 kernel support for CONV_2D_DW (llama/29064)
cf0852d vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (llama/27952)
6d0197c ggml : bump version to 0.25.3 (ggml/1645)
a2b9a22 ggml : fix ubsan error in ggml_graph_nbytes (ggml/1644)
7ab1529 ggml : bump version to 0.25.2 (ggml/1642)
f938f31 vulkan: handle misalignment in conv_2d and conv_3d (llama/29365)
bf8128e vulkan: tune KHR cooperative matrix support for Adreno GPUs (llama/29328)
3ef6a83 cuda : add conv3d with implicit GEMM (llama/29137)
6f2a84a hexagon: reject MUL_MAT_ID when src1 precision is F32 (llama/29348)
0adffd5 opencl: add A8 Q6_K non-MoE dp4a binary kernel (llama/29057)
60c0be6 whisper : add check for ttype in whisper and parakeet (#4091)
6e4ab85 cli : fix incorrect error code when files are failing (#4080)
d09f61a ci : cover GGML_BACKEND_DL in ubuntu-22-clang-arm64 (#4047)
a664346 sync : ggml
84f080e ggml : bump version to 0.25.1 (ggml/1637)
8f5ac2a CUDA: add a reserve to avoid spurious warning on older GCC builds (llama/29317)
8917ea0 metal: add the missing f32 x bf16 mul_mv variants (llama/28741)
f48aebe CUDA: enable sparse-fa for dsv4 prefill (again) (llama/29298)
6bc51cc metal : key the fa-vec tuned table by family instead of SKU (llama/29075)
1852132 vulkan: add IQ4_XS MMQ/MMV matmul kernels (llama/28415)
ed1339a sync : ggml
3e7723f ggml : bump version to 0.25.0 (ggml/1635)
711ef84 common : fix for two functions when top_k exceeds the vocabulary size. (ggml/1633)
f6b039f sycl : fix compile warnings
431ecf5 ggml-meta: resolve multi buffer views (llama/29266)
e0ca36d cuda: top-k MoE should always fire (llama/28432)
d6075ff sycl : support new UT case for mul_mat_hadamard fp16 (llama/29218)
6570b79 sycl: extend MMVQ GLU fusion, add rms_norm+scale and ssm_conv+silu fusions (llama/28931)
a1c7f97 sycl : support op get_rows_back, only support fp32/fp16 (llama/25266)
fbcb94d vulkan: hide internal symbols to prevent duplicate-dlopen state destruction (llama/29139)
b18bad0 hex-dma: introduce direct-mapped DMA cache that is better suited for HVX FA mask handling (llama/29282)
e59366b HIP : optimize IQ2/IQ3 (__vsub4 __vcmpne4) using SWAR (llama/27962)
736efcf opencl: add bin kernel kernel_gemm_noshuffle_q4_k_q8_1_dp4a_ila_a8_bin (llama/29056)
e900c88 vulkan: add Intel Xe flash attention optimization kernels (2/3, Xe-LPG Plus/Xe2/Xe3) (llama/24406)
ad0058b metal : gate mul_mm_id src1 rescale behind ggml_prec (llama/29029)
ff565cf ggml : IQ1_M build prefix sums once per block (llama/28706)
0409ed2 Performance tune for gemma4-26b-a4b flash attention shape. (llama/28450)
dd67678 opencl: add A8 Q4_0 non-MoE dp4a binary kernel (llama/29055)
898392b hexagon: new HMX-optimized GATED_DELTA_NET (llama/29199)
63412d3 metal : fix mask bounds in flash attention block pre-pass (llama/29220)
0e640a1 cuda: fix sm_70 tile compilation error (llama/29224)
9a7d43d ggml-cuda : convert contiguous tensors four elements at a time (llama/29155)
86ef1b1 cuda : accelerate conv2d with implicit GEMM (llama/29135)
d18a422 ggml : fix dimension and stride truncation in ggml_permute (llama/29227)
eb279cf sycl : support gated DSV4_HC_PRE and optional HC_POST comb matrix (llama/29132)
0ecf57d sycl : pinned memory use right device context instead of 0 (llama/28895)
34a4c12 CUDA: Follow up of #25635, refactoring FA shared smem swizzle (llama/28536)
b7b1fe4 ggml-metal : simplify fusion pattern op list declaration (llama/29206)
42c873a sycl : coalesce MKL-FA softmax loads instead of one work-item per row (llama/28918)
b0c5149 ggml-cpu: ARM Repack kernels for Q1_0 (llama/23492)
b50d5c3 hexagon: overhaul of buffer and DMA handling to support 64bit mappings + improvements (llama/29197)
76d02a1 cuda : tune MMVQ to MMQ crossover for SM70 (Volta) (llama/28912)
87b6876 metal : fix deprecation warnings from macOS 27 SDK (llama/29136)
29e710c webgpu : add fused gdn + cpy (llama/28976)
4ef9fd7 CUDA: tune FA for Gemma 4 on Ampere or newer (llama/29152)
984e400 metal : support arbitrary hc in dsv4_hc_pre (llama/29169)
3d949a3 CUDA: enable sparse fa for qwen4 (llama/28770)
c77f6bb metal: add F16 input to the FWHT (llama/29094)
7f5ac73 hexagon: enable I32 GET_ROWS (llama/29116)
232d718 hexagon: add support for GEGLU_QUICK (llama/29114)
099090c hexagon: enable support for TOP_K op (llama/29113)
a00fa32 metal : add MoE and SSM_CONV fusion optimizations (llama/28948)
47456f6 metal : fix FA support checks (llama/29122)
4f1cce1 metal : support qwen4exp hc ops (llama/29000)
58844b0 cuda : fix CUB argsort corruption caused by in-place keys (llama/28389)
97dc017 opencl: add support for bin kernel flash_attn_f32_f16_bin (llama/29046)
b29b439 hexagon: add ROLL op support (llama/29105)
e2358df hexagon: im2col update (llama/29103)
c01abce hexagon: HMX flash-attention head_dim padding (support DK=DV=72) (llama/26539)
1b6c6a9 opencl: add bin kernel kernel_gemm_noshuffle_q6_k_f32_32b_trans_ila_a8_bin (llama/28678)
8f38603 ggml-cpu: add F16 input to the FWHT (llama/27779)
bdf289e ggml-webgpu: fix supports_op condition for GET_ROWS (llama/28978)
26d6dcf ggml : handle graph buffer reservation failure (llama/26070)
8e336cd vulkan: add IQ3_S MMQ matmul kernels (llama/28822)
6b752e5 vulkan: raise the hoisted row-id limit for mul_mat_id from 256 to 512 experts (llama/28501)
f455712 openvino : Update OpenVINO to 2026.4;fix clangd,MSVC warnings; (llama/29009)
fbdbbc7 vulkan: split buffers and debug code into separate files, add shared headers (llama/28732)
7826422 gguf : align the data section relative to the GGUF start, not the file (llama/28993)
b9e5f3a sycl : fix the B70 mem allocate error when >19.3GB (llama/28953)
84a020e vulkan: skip unneeded MoE work in mul_mm coopmat1 path (llama/25483)
cb44896 sycl: ssm_conv: fuse the SiLU epilogue into the ssm_conv kernel (llama/28929)
d4b9101 opencl: fix various warnings (llama/28984)
d375e3c vulkan: fix buffer_reference alignment in im2col shaders (llama/28996)
31c972a vulkan: support qwen4exp hc ops (llama/28988)
d673fba Fix function signature for ggml_backend_sycl_split_buffer_type (llama/28981)
67cf515 vulkan: work around NV bug with argsort_large.comp (llama/28975)
6e220ab hexagon: Support for K-Quants Q4_K and Q6_K (llama/28994)
6ca20bb hexagon: accept the zeroed rope probe in supports_op (llama/28995)
94d6e24 CUDA/HIP: improve access patterns in im2col (llama/28013)
af5153c spacemit : fix wrong transpose function for int16 data (llama/25161)
dce52b1 rpc : invalidate cached compute graph when a referenced buffer is freed (llama/24292)
fa2c801 qwen4exp: add hc ops (llama/28901)
f9ad986 HIP: broaden MoE ncols_opt tile heuristic on RDNA3.5 architecture (llama/28935)
0b9fb0f vulkan: make MUL_MAT_ID BN/2 tail unconditional (llama/28923)
edbb13e metal: fix NaN in mul_mm_id when activations exceed f16 range (llama/26223)
7dd0ce9 hexagon: add back missing contiguous fast-path and hvx_copy_uu for each run (llama/28886)
9e6e308 hex-cpy: use dma if src and dst are contiguous (llama/28906)
43f7501 HIP: Enable AllReduce for ROCm (llama/27825)
bae0f97 opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (llama/27637)
1fc756a rpc : hash-cache only weights (llama/28789)
6189e6f cuda: support row-contiguous SUM_ROWS (llama/26308)
f69f590 vulkan: support sparse Flash Attention (llama/28105)
4c341e2 OpenVINO: optimize stateful decode and GPU MoE inference (llama/28638)
67630d0 opencl: add generic ssm_scan (llama/28881)
4be99fe metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (llama/28599)
b1fd0ca cuda : enable i16 and i32 for DUP (llama/28897)
50e4eed HIP: fattn-mma: use fp32 accumulation on MFMA devices (llama/28576)
398997e whisper : add abort_callback on lang detection (#4077)
a44e078 vad : reject n_encoder_layers other than 4 in model load (#4064)
307869a devops : reduce Vulkan Docker image to 39.4% of its original size (now 692 MB) (#4038)
5670d5c fix(yt-wsp): Resolve script path without GNU realpath (#4072)
b27fbff cli : load backends after validating input files (#4069)
fd7d8ab ci : update android-actions to v4.0.4 (#4074)
d5d6e59 docs : clarify VAD mode timestamps and CWD model path errors (#4019)
7a2ceef readme : document the ANEForge encoder backend (#4073)
4afa009 whisper : optional ANEForge encoder backend (Apple Neural Engine) (#3905)
da54572 whisper : fix int overflow in whisper_full_parallel chunk offsets (#4044)
1d549b3 sync : ggml
ce5c557 ggml : bump version to 0.24.0 (ggml/1627)
950a4a4 tests(s390x): add non-vxe build to tests (llama/28776)
0a86673 sycl: rfc: Use radix select for top_k (llama/28670)
a3d2600 ggml-cpu : disable PCH and fix CACHE_LINE_SIZE ambiguity to fix heap corruption (llama/28882)
0bea289 sycl : fix oneDNN scratchpad breaking the pool free order (llama/28704)
ac78ae9 ggml-cuda: fallback to F32 on device without BF16 hardware acceleration (llama/28846)
8f4cd19 ggml-cpu(s390x): guard VXE-only repack helpers (llama/28775)
51ee927 sycl : Fix get mem error (llama/28227)
f59047c vulkan: workaround NV queuesubmit driver bug (llama/28830)
17e6392 opencl: apply the noshuffle row-alignment rule to q4_K, q5_K and q8_0, not just q6_K (llama/28575)
8751eae ggml-cuda: hip add specific config table for AMD GCN (llama/27841)
2403583 syscl : Handle (fail gracefully) unsupported tq1_0 quants (llama/28681)
df31856 rpc : fix linking when compiling with BUILD_SHARED_LIBS=OFF (llama/28492)
07825dc opencl: fix several bugs where the backend aborts (llama/27630)
ec82a96 opencl: add bin kernel kernel_gemm_noshuffle_q4_k_f32_32b_trans_ila_a8_bin (llama/28677)
b0e076d webgpu: align tensor bindings to the type block size (llama/28382)
ea3ef8f hexagon: support for multi-device model split (aka row-split) (llama/28589)
b177941 ggml-webgpu: Update to a recent version of Dawn (llama/28683)
2c1b525 ggml: skip 0-sized ids tensor when offloading selected experts (llama/28739)
5f310fa metal : skip the empty half of the mul_mm_id token tile (llama/28301)
76da352 cmake : add PCH and unity build to improve build times (llama/28091)
9d70ac2 metal : single-source fusion table + fusion debug rework (llama/28164)
59cca2c metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024 (llama/28692)
61b318f CUDA/HIP: Flash Attention tuning (gfx1201) (llama/28102)
468710c vulkan: fix data race and OOB access in argsort(large) (llama/28705)
e5369a6 opencl: add A8 Q4_0 mm binary kernel support (llama/28268)
14e868e vulkan: use CPU writes in ggml_backend_vk_cpy_tensor_async if the context is idle (llama/28618)
fb8f427 vulkan: use add_alloc_dep to enable topk_moe fusion for prefill (llama/28422)
e955658 vulkan: small M matrix optimizations for qwen (llama/28457)
d374147 vulkan: fall back to shared-memory reduction for dmmv on PowerVR (llama/28341)
a6e85dd vulkan : add command-buffer debug labels for GPU profilers (llama/28101)
9f7331f ggml-cpu(s390x): add repack support for q4_0 (llama/28667)
d3945bf ggml-cpu(s390x): add Q1_0 vector intrinsic support (llama/28606)
3a54d53 vulkan: use spec constant for matrix matrix multiplication A-type (llama/25773)
bafeaca hexagon: rope updates (llama/28628)
84b8db9 vulkan: Convert FILL to distribute workgroups in 2D to avoid exceeding maxComputeWorkGroupCount (llama/28592)
f55b67f CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (llama/28552)
4d506f5 CUDA: replace GGML_FA_ALL_QUANTS with GGML_FA_QUANTS, more control over what is compiled (llama/28079)
dca2df4 vulkan: add dedicated iq4_xs mat-vec shader (llama/28426)
355d90b vulkan: add f16 B-type matmul pipelines and warp tile size tuning for Intel coopmat1 (llama/27471)
facf4b5 Add IQ type handling for MoE (llama/28476)
32a749d ggml : fix msvc+clang ggml_vld1q_u32 (llama/28284)
006e53e Revert "ggml-cuda : restore prop.integrated on HIP builds (llama/24233)" (llama/28604)
5b98494 metal : fix idle threads in mul_mv_iq3_xxs for ne00 < 1024 (llama/28086)
8bae082 llama : add missing headers (llama/28566)
b127543 vulkan : fuse UNARY(GELU|SIGMOID|SILU|SOFTPLUS) + MUL (llama/27220)
69fcec3 Fix Vulkan-Hpp handle usage on 32-bit targets. (llama/22892)
37b210d opencl: properly handle non-contiguous inputs to conv2d (llama/28503)
33cadde ggml : update ggml_prec specification (llama/26675)
707c3ee hexagon: add RELU and LEAKY_RELU ops (llama/28585)
34f5336 Revert "CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (#24546)" (llama/28551)
0f9591a webgpu: format the GET_ROWS case block (llama/28542)
b61186d sycl: add a batched L2_NORM kernel (llama/28222)
8ce432f vulkan: add DeepSeek-V4 hyper-connection fused ops (DSV4_HC_COMB/PRE/POST) (llama/26578)
4c5a9a0 CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (llama/24546)
f15e1a6 ggml: add gfx90c HIP support (llama/26454)
a11d16a CUDA: branchless Q4_K/Q5_K unpack to speed up mmvq, L2 prefetch on DGX Spark (llama/26705)
60e475f vulkan: support type-aligned GET_ROWS (llama/28253)
c6135ac ggml-cuda: fix divergent barrier in f16 flash attention (llama/27870)
1da558c ggml: allow backend inputs to not create another split (llama/28387)
8d3ed20 vulkan: rms_norm fusion opportunities (llama/28024)
088b3b4 vulkan: add TQ1_0 support (mm, mat-vec, mat-vec-id, dequant, get_rows) (llama/27765)
ad344d1 opencl: properly choose weights pack for q4_K, q5_K mul_mat (llama/28402)
a0e75c8 cuda: fixes races in mmid and mmf (llama/28475)
1299eb6 metal : add remaining fa-vec tunings for M2 Max (llama/28458)
614c74a metal : fix memory leak in early return (llama/28399)
0006e3e sycl : fix test-backend-ops CI break && restore Kronecker product FWHT support (#28016) (llama/28254)
0779610 sycl: attribute device allocations by site (GGML_SYCL_MEMTRACE) (llama/27631)
ad218df metal : add remaining fa-vec tunings for M3 (llama/28396)
25c6ab5 opencl: extend the elementwise and data‐movement op coverage (#27633)
7fb2daf opencl: add Adreno xmem SDPA path (llama/26331)
70598ee scripts : fix sync (#0)
f133970 ci : use devlab-dispatch for npu-amd-windows (#4060)
1da4dc8 ci : rename cublas to cuda in release.yml (#4057)
0261298 ci : use devlab-dispatch for npu-amd-linux (#4056)

Don't miss a new whisper.cpp release

NewReleases is sending notifications on new releases.