AI compliance agents made rejected and rule-breach attempts in simulations, preprint finds
Tests found that plausible compliance language did not guarantee rule-grounded actions, while monitor scores depended on the evidence supplied.
Developing Light ยท https://developinglight.com/editorial/developing-light