Overlapping Tests Can Shift Physical-AI Model Rankings
A preprint finds benchmark overlap can shift physical-AI model rankings, while four tests retain 78.5% of the full suite’s measured utility.
Developing Light · https://developinglight.com/editorial/developing-light