One AI auditor setup scored above a baseline, but results varied
A preprint reports that a calibrated pairwise auditor scored above a baseline, while concern-focused training showed lower false-positive calibration.
Developing Light ยท https://developinglight.com/editorial/developing-light