Tests found that plausible compliance language did not guarantee rule-grounded actions, while monitor scores depended on the evidence supplied.