keryx-miner-supr v0.11.6 — solver tuning: +18% NVIDIA (on top of v0.11.5), RDNA3 WMMA default on AMD
NVIDIA (sm_80+): three dispatch-level improvements to the v4 tensor-core solver (kernels unchanged, still byte-exact — pool-validated with zero rejects):
- The tile-offsets buffer is now cached per device (the old path paid a cudaMalloc + 16 MB memset + free every batch).
- Default nonce batch raised 16K → 64K (
KERYX_POM_V4_BATCHstill overrides); a 64K batch is ~25 ms on a mid Blackwell card, well under the 100 ms block time. - The offset-chase now overlaps the walk on a secondary stream (
KERYX_POM_V4_OVERLAP=0disables).
Measured live (5070 Ti, defaults vs defaults): 2.38 → 2.82 Mh/s (+18.5%). Combined with v0.11.5's tensor-core solver that is ~+55% over v0.11.4 on Ampere+.
AMD (RDNA3): the v4 walk now uses int8 WMMA (V_WMMA_I32_16X16X16_IU8) by default — ~+11–18% on smaller RDNA3, +2.5% on the 7900 XTX.
The v0.11.5 VRAM-cooling note still applies — higher hashrate = proportionally more VRAM bandwidth/heat; see that release's notes if clocks drop.
Assets
NVIDIA modern (CUDA 12.9, driver 575+) + legacy (CUDA 12.4) × hiveos/mmpos/smos + SHA256SUMS-nvidia.txt; Windows + AMD attached separately.