A preprint reports lower HPC job-runtime error and a 17-fold speedup over AVX; its end-to-end result comes from a tested ALCF benchmark configuration.