Mixed training scores higher than next-chunk reasoning in AI tests
A preprint finds mixed SFT scored higher than next-chunk reasoning after RLVR, which used over 60 times more GPU hours in the study.
Developing Light ยท https://developinglight.com/editorial/developing-light