Warning
Rolling preview build from v0.4.4 at 0b035b3a26f1. Assets and moving Docker tags are replaced by newer successful branch builds. Last updated: 13/08/2026 12:25.
Changelog
- Updated the llama.cpp base through upstream commit
84e908c6. Notable inherited changes include Granite-Switch and Muse Glimmer model support, MTP support for Nemotron and DFlash support for Nemotron 3.5, Pocket TTS audio generation, multi-output backend sampling for speculative decoding, media-aware server slot save/restore, the Web UIread_mediatool, and expanded tool isolation through SSH and rootless Podman. The merge also adds the defaultload-mode autopolicy that avoids memory mapping on integrated GPUs, Vulkan TQ2_0 support, a warp-per-row CUDA WKV7 kernel for single-token decode, narrower CUDA-graph synchronization, hardened GGUF loading, semantic versioning, and version-aware CMake package metadata.
macOS:
Linux:
- Ubuntu x64 CPU
- Ubuntu arm64 CPU
- Ubuntu x64 CUDA 12.4
- Ubuntu x64 CUDA 13.1
- Ubuntu x64 Vulkan
- Ubuntu x64 ROCm 7.2
- Ubuntu x64 SYCL
Windows:
- Windows x64 CPU
- Windows x64 Vulkan
- Windows x64 SYCL
- Windows x64 CUDA 12.4 - DLLs
- Windows x64 CUDA 13.1 - DLLs
- Windows x64 HIP
Docker:
- CPU:
docker pull ghcr.io/anbeeld/beellama.cpp:server-cpu-preview-v0.4.4 - CUDA:
docker pull ghcr.io/anbeeld/beellama.cpp:server-cuda-preview-v0.4.4 - CUDA 12:
docker pull ghcr.io/anbeeld/beellama.cpp:server-cuda12-preview-v0.4.4 - CUDA 13:
docker pull ghcr.io/anbeeld/beellama.cpp:server-cuda13-preview-v0.4.4 - ROCm:
docker pull ghcr.io/anbeeld/beellama.cpp:server-rocm-preview-v0.4.4 - Vulkan:
docker pull ghcr.io/anbeeld/beellama.cpp:server-vulkan-preview-v0.4.4 - SYCL:
docker pull ghcr.io/anbeeld/beellama.cpp:server-sycl-preview-v0.4.4