LLMs Fail at Many Simultaneous Constraints: New Benchmark Reveals Phase Transition
A new study from arXiv (2608.12426) introduces Constraint Saturation Evaluation (CSE), a benchmark to test large language models' ability to follow multiple simultaneous instructions. The research reveals a phase transition: while models handle individual constraints well, performance collapses when many must be satisfied at once. The study tested 15 models across 36 constraint types, totaling 369,753 checks at k=1-12 constraints. Key findings include a gradual decay in per-constraint pass rate but a sharp collapse in the probability of satisfying all constraints. For instance, a model passing individual constraints at ~41% at k=8 shows near-zero joint success. The benchmark uses deterministic, rule-based verification, avoiding LLM judges. This work is significant for AI safety and reliability, as LLMs are increasingly deployed in settings requiring adherence to multiple explicit constraints, such as reasoning structure, safety boundaries, and output schemas. The study also explores factors governing degradation and potential mitigation strategies. The paper is available at https://arxiv.org/abs/2608.12426.
Key facts
- Introduces Constraint Saturation Evaluation (CSE) benchmark
- Tests 15 models, 36 constraint types, 369,753 checks
- Constraint counts range from k=1 to k=12
- Uses deterministic, rule-based verification, no LLM judges
- Per-constraint pass rate decays gradually
- Joint satisfaction of all constraints collapses sharply
- Example: ~41% per-constraint pass rate at k=8 leads to near-zero joint success
- Aims to characterize compositional constraint satisfaction in LLMs
Entities
Institutions
- arXiv