Warning
Rolling preview build from v0.4.3 at 9d59db2e3e99. Assets and moving Docker tags are replaced by newer successful branch builds. Last updated: 29/07/2026 16:00.
Changelog
- Fixed HIP/ROCm KVarN caches incorrectly falling back to full F16 materialization when precision tails were enabled. KVarN now honors the backend's dedicated native-tail capability before consulting generic segmented-tail support, avoiding context-sized fallback allocations on supported AMD GPUs while leaving ordinary HIP segmented-tail routing unchanged.
macOS:
Linux:
- Ubuntu x64 CPU
- Ubuntu arm64 CPU
- Ubuntu x64 CUDA 12.4
- Ubuntu x64 CUDA 13.1
- Ubuntu x64 Vulkan
- Ubuntu x64 ROCm 7.2
- Ubuntu x64 SYCL
Windows:
- Windows x64 CPU
- Windows x64 Vulkan
- Windows x64 SYCL
- Windows x64 CUDA 12.4 - DLLs
- Windows x64 CUDA 13.1 - DLLs
- Windows x64 HIP
Docker:
- CPU:
docker pull ghcr.io/anbeeld/beellama.cpp:server-cpu-preview-v0.4.3 - CUDA:
docker pull ghcr.io/anbeeld/beellama.cpp:server-cuda-preview-v0.4.3 - CUDA 12:
docker pull ghcr.io/anbeeld/beellama.cpp:server-cuda12-preview-v0.4.3 - CUDA 13:
docker pull ghcr.io/anbeeld/beellama.cpp:server-cuda13-preview-v0.4.3 - ROCm:
docker pull ghcr.io/anbeeld/beellama.cpp:server-rocm-preview-v0.4.3 - Vulkan:
docker pull ghcr.io/anbeeld/beellama.cpp:server-vulkan-preview-v0.4.3 - SYCL:
docker pull ghcr.io/anbeeld/beellama.cpp:server-sycl-preview-v0.4.3