Hotfix for 0.1.40.3: decode on Pascal cards (Tesla P40, GTX 1080, GTX 1070) is back to the 0.1.40 speed. Default answers are unchanged. Update with UPDATE.bat (Linux: ./update.sh); setup replaces the engine with 0.1.40.4.
Pascal decode (#1469, #1470)
- Decode on sm_61 had halved since 0.1.40.2: the small-batch matrix kernels lost their
__restrict__hints and gained a prefetch that Pascal cannot overlap, so every decode step ran roughly twice as long. The reporters' own version of this fix took a P40 from 17.7 back to 36.7 tok/s; we have no sm_61 card, so a confirmation on 0.1.40.4 in #1469 helps. - The fix is limited to cards below sm_70. For Volta and newer (RTX 20, 30, 40, 50) the compiled code is byte-identical to 0.1.40.3, so nothing changes there. A Tesla P100 (sm_60) measured +2% (five A/B pairs, +1.7 to +2.4%, same answers).
- The Windows CUDA 12 zip (the one for Pascal and Volta) carries the fix; so does the Linux build from source.
Thanks to paulhothersall and lineape for the bisect that found the commit.
Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going: