A preprint reports SRPO benchmark scores for math and agents alongside substantially lower training compute than GRPO and scaled SFT.