AI method reports benchmark gains over GRPO with far less compute
A preprint reports SRPO benchmark scores for math and agents alongside substantially lower training compute than GRPO and scaled SFT.
Developing Light ยท https://developinglight.com/editorial/developing-light