A preprint reports higher failure reproduction and repair rates with selective replay than full reruns, while noting that benchmark results may not generalize.