v0.10.9 — Gemma-4-12B now fits 16 GB cards (with inference)
Fix: repacking-architecture models (Gemma) on 16 GB GPUs. llama.cpp repacks tied-embedding architectures on load (Gemma materialises a separate output.weight), so its resident tensor layout differed from the canonical GGUF the possession walk pins. The previous build treated any layout mismatch by unloading the inference engine and walking a full second raw copy of the weights — which dropped inference and ran a 16 GB card out of VRAM, so Gemma got demoted to a smaller tier there (e.g. a 5070 Ti / 5080 served GLM instead of Gemma).
v0.10.9 gathers the walk per tensor: every tensor whose resident copy matches the canonical GGUF (unique name + exact size + on-device) is walked zero-dup in place, and only the handful llama repacked are re-uploaded from the possession index (falling back to a raw copy only if >25 % of the blob would need it). Gemma now stays resident as a single ~12 GB copy on a 16 GB card — walk and inference — instead of OOMing on a second copy.
Validated on the fleet (pool mining): Gemma-4-12B mining + serving on RTX 5070 Ti and RTX 5080 (16 GB); GLM-4-9B and Qwen3.5-9B unchanged.
Also:
- VRAM-fit tier auto-select tuned (per-tier floor ladder).
- New
--low-ram/--save-ram: brings models up one card at a time to cut the system-RAM spike when loading many GPUs.
Assets: linux-x86_64 (dynamic), HiveOS, and MMPOS tarballs, each in modern (newer GPUs, 575-class driver) and legacy (-legacy, 535-class driver) lines. SHA256SUMS-0.10.9.txt for verification. AMD and Windows builds are attached separately.