github Anbeeld/beellama.cpp preview-v0.4.4
v0.4.4 Preview

pre-release4 hours ago

Warning

Rolling preview build from v0.4.4 at 0b035b3a26f1. Assets and moving Docker tags are replaced by newer successful branch builds. Last updated: 13/08/2026 12:25.

Changelog
  • Updated the llama.cpp base through upstream commit 84e908c6. Notable inherited changes include Granite-Switch and Muse Glimmer model support, MTP support for Nemotron and DFlash support for Nemotron 3.5, Pocket TTS audio generation, multi-output backend sampling for speculative decoding, media-aware server slot save/restore, the Web UI read_media tool, and expanded tool isolation through SSH and rootless Podman. The merge also adds the default load-mode auto policy that avoids memory mapping on integrated GPUs, Vulkan TQ2_0 support, a warp-per-row CUDA WKV7 kernel for single-token decode, narrower CUDA-graph synchronization, hardened GGUF loading, semantic versioning, and version-aware CMake package metadata.

macOS:

Linux:

Windows:

Docker:

  • CPU: docker pull ghcr.io/anbeeld/beellama.cpp:server-cpu-preview-v0.4.4
  • CUDA: docker pull ghcr.io/anbeeld/beellama.cpp:server-cuda-preview-v0.4.4
  • CUDA 12: docker pull ghcr.io/anbeeld/beellama.cpp:server-cuda12-preview-v0.4.4
  • CUDA 13: docker pull ghcr.io/anbeeld/beellama.cpp:server-cuda13-preview-v0.4.4
  • ROCm: docker pull ghcr.io/anbeeld/beellama.cpp:server-rocm-preview-v0.4.4
  • Vulkan: docker pull ghcr.io/anbeeld/beellama.cpp:server-vulkan-preview-v0.4.4
  • SYCL: docker pull ghcr.io/anbeeld/beellama.cpp:server-sycl-preview-v0.4.4

Browse all container images

Don't miss a new beellama.cpp release

NewReleases is sending notifications on new releases.