ARTFEED — Contemporary Art Intelligence

Rule Blindness: Compliance Detectors Fail to Read Regulations

ai-technology · 2026-08-18

An audit examining compliance detectors in language models has uncovered a significant issue: these systems, intended to ensure adherence to regulations concerning data protection, healthcare, financial oversight, and platform policies, are 'rule blind.' The research, available on arXiv (2608.16852), shows that modifying or removing the governing rule does not affect detection accuracy across all evaluated guards and activation probes. Notably, even a policy-conditioned guard that accurately references the governing clause shows little change in its judgment when that clause is replaced with a more permissive one. Researchers created a specialized benchmark that combines two rules with two scenarios, confirming this oversight in a way not previously explored. This revelation questions the effectiveness of compliance monitoring in active language models, indicating that detectors may focus on superficial aspects rather than the actual rules. The study advocates for a reevaluation of compliance detector assessment and implementation, highlighting the necessity for mechanisms that are sensitive to rules.

Key facts

  • Compliance detectors in language models are rule blind.
  • Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged.
  • Tested guards and activation probes include a policy-conditioned guard.
  • The policy-conditioned guard correctly cites the governing clause but barely changes its verdict when swapped.
  • A purpose-built benchmark crosses two rules with two scenarios.
  • The benchmark ensures neither rule nor scenario alone predicts the label.
  • The study is posted on arXiv with identifier 2608.16852.
  • The failure is confirmed under a design no prior work has addressed.

Entities

Institutions

  • arXiv

Sources