Principle-Bench: A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation
A new benchmark called Principle-Bench has been introduced by researchers to assess the reliability of LLM-as-judge systems within principle-based regulation. This benchmark includes 168 scenarios related to cryptoasset financial promotions, aligned with two principles from the UK Financial Conduct Authority (FCA). It features variations such as paraphrases, adversarial keyword manipulation, and boundary cases created following a pre-registered framework. The authors emphasize that evaluating any LLM judge should involve four criteria: accuracy, robustness against paraphrasing, adversarial resilience, and calibration. They also present Ceca (Calibrated Exemplar-Cluster Assessment), an auditable tool offering precise counterfactual attributions. A comparison of various methods reveals that no single approach excels in all areas. The research is accessible on arXiv with the identifier 2608.14329.
Key facts
- Principle-Bench includes 168 cryptoasset financial-promotion scenarios.
- Scenarios are mapped to two UK FCA principles.
- Perturbations include paraphrase, adversarial keyword-stuffing, and boundary cases.
- The benchmark covers four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration.
- Ceca (Calibrated Exemplar-Cluster Assessment) is introduced as a calibrated, auditable assessor.
- Ceca emits exact per-exemplar counterfactual attributions.
- Methods compared: keyword counting, three sentence-transformer embedders, an open-weight LLM judge, and a calibrated cascade.
- No method dominates across all axes.
- The paper is available on arXiv (2608.14329).
Entities
Institutions
- UK Financial Conduct Authority (FCA)
- arXiv