Exception Chain Collapse in Frontier LLM Rule Evaluation
A research paper highlights a failure mode in advanced large language models known as exception chain collapse, which occurs during eligibility evaluations involving nested conditional rules. Initially reproducible, this failure exhibits an unstable empirical surface: between March and April 2026, several failure cells were quietly resolved under the same model alias without any version update (for instance, GPT-5.4 in construction insurance improved from 96.6% to 100% using the same prompt and harness). In regulated environments, the accuracy of frontier models represents a fluctuating compliance threshold that can change unexpectedly. The paper introduces the Aethis Eligibility Module, a neuro-symbolic framework where LLMs generate rules from reliable sources, and an SMT-based layer executes them deterministically, adhering to the original specifications despite model drift, reasoning defaults, or prompt variations.
Key facts
- Exception chain collapse is a failure class in frontier LLMs
- Failure observed in eligibility evaluation under nested conditional rules
- Between March and April 2026 failure cells closed silently with no version bump
- GPT-5.4 on construction insurance moved from 96.6% to 100% accuracy
- Frontier-model accuracy is a moving compliance boundary
- Aethis Eligibility Module is a neuro-symbolic architecture
- LLMs author rules from authoritative sources
- SMT-based layer executes rules deterministically
Entities
—