v0.6.9.5 — AMD: streamed PoM tier load + TDR-safe dispatch
Optimization pass over the AMD/OpenCL PoM driver. The walk kernel itself is unchanged and byte-identical — it already sits at the memory-bandwidth ceiling of the hardware; this release cuts the miner's memory footprint and hardens dispatch instead.
Changes
- Streamed tier load (no host-RAM copy): the PoM tier blob is now streamed straight from the GGUF on disk into each GPU's buffer through a bounded 256 MiB window, using a new bulk chunk reader (one read per tensor instead of one per 32-byte chunk). This removes the permanent full-blob copy in system RAM (~2.5 GB for Gemma; ~28 GB it would have cost for the 70B tier) and cuts tier load time to seconds.
- TDR-safe sub-dispatch: the nonce search now runs in 2^18-nonce kernel launches (ascending order, early exit on the first winning sub-batch). Every launch stays far below the Windows TDR watchdog even on slow cards, and winning shares are submitted sooner. Result-identical lowest-nonce semantics; hashrate unchanged.
Consensus safety
pom.rs folds/layout untouched (the only addition is a read-only bulk accessor). Verified byte-exact: H3 salt test, pinned Gemma root R_T 846caa40…, and the real-tier GPU end-to-end test (walk through the streamed upload finds the pinned nonce) all pass.
Validation (live pool, 3-GPU rig: RX 7600 XT + 2× MI50/MI60)
- 29.4 Mh/s total — parity with v0.6.9.3/4
- 30+ shares accepted, 0 rejected (tier 1, H3 active)
- Miner RSS: 3.2 GB → 0.7 GB
- All cards PoM-resident within ~30 s of index load; GPU inference unaffected
AMD assets
| file | sha256 |
|---|---|
keryx-miner-supr-amd-0.6.9.5.tar.gz (HiveOS)
| e6edc1975cb32fe7393a5b940596156f3fbff052a62c23a1548ede86928f1d79
|
keryx-miner-supr-amd-mmpos_0.6.9.5.tar.gz (mmpOS)
| 06fd496d9868868d84becb7d9ad8318d91a0c6663628608e9c3c193723d38777
|
keryx-miner-supr-amd-0.6.9.5.zip (Linux runnable)
| 26248bfccdbee5aeed1ed5bc15b946ff1011a07e9a8a104755aa762fb9791ccb
|
NVIDIA/CUDA path is untouched by this release (the CUDA driver keeps its own load path). Windows AMD build follows via CI.
NVIDIA Linux assets (modern / legacy / pascal)
The 12 NVIDIA Linux tarballs + per-line SHA256SUMS are attached (walk PTX verified per line: modern sm_75, legacy sm_70, pascal sm_60; banner 0.6.9, BUILD 23). NVIDIA-side changes since v0.6.9.3:
- The new bulk chunk reader above is inert on CUDA — the CUDA walk already streams model tensors straight to VRAM (zero-dup with inference, no host blob). Byte-exact re-verified on the shared code: full PoM test suite + the pinned Gemma root (
R_T 846caa40…) against the real GGUF. - Clear arch-mismatch errors: if the walk kernel fails to load on an older GPU, the miner now names the build's PTX architecture and the correct build line (e.g. Tesla V100/Volta → legacy) instead of a bare
CUDA_ERROR_INVALID_PTX. - README build docs corrected: CUDA 13.x cannot compile for Volta/Pascal — source builds for those need CUDA 12.x (or use the prebuilt legacy/pascal lines).
Note: the keryx-miner-supr-macos-arm64-0.6.9.3.tar.gz asset is the v0.6.9.5 build — the filename carries a stale version label from CI naming; it will be corrected next release.
Windows assets
Attached by CI (click "Show all assets" below if the list is collapsed):
keryx-miner-supr-windows-nvidia-pom.zip— NVIDIA miner (exe + bundled CUDA DLLs + start.bat)keryx-miner-supr-windows-amd.zip— AMD/OpenCL miner (includes the v0.6.9.4 Adrenalin asm-fallback fix)keryxcuda-windows.zip— CUDA plugin DLL only