New preprint MTDiag presents a multi-turn diagnostic dataset and evaluation framework, but reports no systematic comparison of language models.