PrismML llama.cpp
PrismML's llama.cpp fork as a Harbor backend, so Ternary Bonsai 2 PTQ1_0/PQ2_0 GGUFs run on CPU, NVIDIA and ROCm with the same model workflow, harbor prismml CLI and cross-service integrations as llamacpp. On a 16 GB RTX 4090 Laptop, Bonsai 2 27B scores within a few points of a Q3 Qwen3.8-27B on GSM8K/HumanEval/MATH-500 while using 8.3 GB instead of 14.4 GB and generating 25% faster; numbers are in the docs.
harbor up prismmlMisc
harbor models pull --source prismmldownloads Bonsai GGUFs through a runtime that understands them.- Every llama.cpp cross-service integration has a PrismML twin, so
harbor up prismml webuiand friends just work.
Full Changelog: v0.5.11...v0.5.12
