keryx-miner-supr 0.14.6
AMD release: a card that hosts the inference model now mines as well, from the same single copy of the model in VRAM. Recommended for every AMD rig; nothing changes for NVIDIA.
AMD: inference and mining from one model copy
- The inference card mines again on 12 GB (and other small) AMD cards. On AMD, inference runs in llama.cpp (Vulkan) and the PoM walk in OpenCL, so a card doing both used to need the model twice in VRAM. When model + walk did not fit (for example Qwen3.5-9B on a 12 GB RX 6700 XT: 6.2 GB × 2 + headroom), the miner reserved that card for inference only — it answered requests and showed
INFat 0 H/s, costing a 5-card rig one fifth of its hashrate. The miner now carries a Vulkan PoM v4 walk kernel that reads the 1 KB tiles directly from the inference engine's resident weight tensors (buffer device addresses through a per-tensor table), so the reserved card mines over the copy it already holds. Before the first share the miner runs a byte gate: it reads sampled chunks back through the exact gather path the walk uses and compares them with the model file; any difference keeps the card inference-only. Every winner is still rebuilt and verified on the CPU before it is submitted. - Dashboard: such a card is shown
MININGwith the[INF]badge and its hashrate; it still pauses briefly while it answers a request, like every AMD card. - Policy: cards where model and walk fit together keep the OpenCL walk (its RDNA3 WMMA / dp4a kernels remain the fastest).
KERYX_ZERO_DUP=forcemakes every card hosting the model use the Vulkan walk,KERYX_ZERO_DUP=offrestores the inference-only reservation. RDNA1 (RX 5600/5700) never uses it (known driver hang on buffer-address reads). Multi-GPU rigs: only the one card hosting the model changes; the others mine exactly as before. - Speed: the Vulkan walk matches the OpenCL kernels. A chase pass resolves each nonce's 256 tile addresses up front (the walk itself does no table lookups and prefetches two tiles ahead), the matrix cores are used on RDNA3 and newer (
VK_KHR_cooperative_matrix, int8, exact) and packed int8 dot products elsewhere. Measured on the pool models: RX 7600 XT 0.58 MH/s (OpenCL kernel 0.59), RX 7900 XTX 1.44 MH/s (OpenCL 1.51), MI50 0.96 MH/s (OpenCL 0.52). The two-phase scratch takes 67 MB of VRAM; without it the miner falls back to a single-phase walk. - Tested byte-exact against the CPU reference walk for all three seed eras (pre-H10, H10, H14) on RX 7900 XTX (gfx1100), RX 7600 XT (gfx1102) and MI50 (gfx906), with both kernel variants (packed int8 dot product and the scalar fallback); live on the pool with accepted shares and answered inference requests.
Fixed
- AMD: an inference answer cut inside a multi-byte character was rejected as a failed generation. When the token budget ended an answer in the middle of a UTF-8 character (byte-level BPE), the Vulkan engine discarded the whole answer, the pool request got no output, and after the usual failures the model was withdrawn and mining suspended ("no models ready — mining suspended"). The answer is now trimmed to its last complete character, as the NVIDIA engine already did.
Which download
| GPU | Download |
|---|---|
| NVIDIA Turing (RTX 20-series) and newer, Linux, driver 575+ | modern line (standalone, HiveOS, mmpOS, SMOS, Docker)
|
| NVIDIA Tesla V100 / Titan V / Quadro GV100 (Volta), GTX 10-series/Pascal, CMP 100–210, or driver 550–574 | legacy line (standalone, HiveOS, mmpOS, SMOS)
|
| Windows NVIDIA, Turing and newer | keryx-miner-supr-windows-nvidia-pom.zip
|
| AMD (Linux Vulkan / Windows) | -amd packages
|
No consensus, protocol or model changes; 0.14.x miners can run side by side. NVIDIA packages are rebuilt from the same source and are functionally identical to 0.14.5.