A preprint reports higher code benchmark scores after faulty-code test synthesis and dense rewards, while an audit found synthetic tests remained invalid.