A preprint audit finds AI self-training gains can reflect measurement artifacts, while external distillation reaches more low-base problems than tested self-training