ARTFEED — Contemporary Art Intelligence

LLMs Know Constraints but Fail to Use Them: Activation Bottlenecks in Pragmatic Reasoning

ai-technology · 2026-08-15

A recent study published on arXiv (submission 2608.12321) examines the reasons behind the failures of large language models (LLMs) when a prominent surface cue contradicts an implicit feasibility constraint. The researchers define this issue as 'conditional constraint activation,' differentiating between the model's internal understanding of the constraint and its application in decision-making. They introduce a four-part diagnostic that assesses knowledge (whether the constraint is encoded), symmetry (its equal encoding in both constraint-present and -absent prompts), routing (the use of the constraint in final decisions), and repair (the ability to rectify the decision through donor activation). Testing 14 models reveals two distinct failure modes. Probes on two open-weight models achieve over 88% accuracy in decoding the constraint, with activation patching improving one model (+6.4 nats) while harming the other (-0.07 nats). The study also investigates mitigation strategies, concluding that no prompted intervention successfully reaches the repair corner, as all interventions increase conservative bias through prerequisite mention. The authors assert that hidden-constraint failure is primarily a routing issue rather than a knowledge issue. This research falls under Computer Science > Computation and Language and includes references, citations, and related code and data.

Key facts

  • The paper is titled 'LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning'.
  • It is available on arXiv with ID 2608.12321.
  • The study introduces the concept of 'conditional constraint activation'.
  • A quartet diagnostic measures knowledge, symmetry, routing, and repair.
  • 14 models were tested in the study.
  • Probes on two open-weight models decode the constraint with over 88% accuracy.
  • Activation patching repairs one model (+6.4 nats) but not the other (-0.07 nats).
  • No prompted intervention reaches the repair corner; all inflate conservative bias via prerequisite mention.
  • The paper concludes that hidden-constraint failure is a routing problem, not a knowledge problem.

Entities

Institutions

  • arXiv

Sources