A 44-sample preprint benchmark found higher PDE-term recovery scores from physical summaries than raw slices, using noise-free simulations.