keryx-miner-supr v0.6.8.0
Minor release: H200 / Hopper PoM optimization + NVIDIA CMP 100-210 (Volta) support.
What's new
PoM walk — 128-bit vector loads (ulonglong2). Each 32-byte weight chunk is now read as 2× ulonglong2 instead of 4× u64. The XOR fold is mathematically identical (a.x^a.y^c.x^c.y == the four u64), so the winning proof is byte-for-byte the same (consensus-safe, verified live post-H3-gate). Fewer load instructions + better coalescing → +1.8 % on H200/Hopper, neutral on Blackwell/Ampere. Validated 0-rej on a 5090 + CMP 170HX and on an H200.
CMP 100-210 / Volta (sm_70) support (legacy line). The PoM walk PTX was emitted at compute_75, which cannot JIT down to Volta sm_70. The legacy build now emits the walk at compute_70 (the walk is pure u64 + gather — no dp4a/fp16/tensor — so it compiles at sm_70 and forward-JITs to sm_75…sm_120 with identical memory-bound performance). This adds the NVIDIA CMP 100-210 (GV100 mining card) and Tesla V100. Use the legacy package on those cards.
H200 performance
Fastest PoM card measured: ~166 MH/s live @ ~628 W, 98 % util (HBM3e ~4.8 TB/s, mem clock 3201 MHz) — +32 % over an H100 (~123 MH/s). PoM is memory-bound → cap power for free efficiency; no core-clock gain. See BENCHMARKS.md.
Which package do I download?
| Card family | Package |
|---|---|
| Blackwell 50-series / datacenter (sm_120/sm_100), R575+ | modern |
| Turing / Ampere / Ada / Volta (V100, CMP 100-210) / datacenter, R550+ | legacy |
| Pascal (GTX 10-series, 1080 Ti, P100 — sm_60/61) | pascal |
Each line bundles its own CUDA runtime (driver floor). HiveOS/SMOS/mmpOS/generic-Linux tarballs for all three lines below; Windows build attached by CI.