A preprint finds benchmark overlap can shift physical-AI model rankings, while four tests retain 78.5% of the full suite’s measured utility.