ARTFEED — Contemporary Art Intelligence

CROWN-QA: Evaluating Completeness-Sensitive Negative Reasoning in LLMs

ai-technology · 2026-08-06

A recent research paper presents CROWN-QA, a benchmark designed to assess completeness-sensitive negative reasoning in large language models (LLMs). Published on arXiv (2608.04591), this study tackles a significant issue: LLMs frequently respond with 'no' when inquiring about the absence of items from records, lists, or contexts. However, such negative responses are only appropriate if the evidence fully encompasses the query; otherwise, 'unknown' should be the answer. CROWN-QA consists of two parts: CROWN-Synth, which fixes the question and observed facts while varying query-relative coverage, and CROWN-Real, which evaluates real documents with controlled coverage variations. Testing across three LLM families reveals inconsistent closure judgments and considerable over-closure, indicating models struggle to differentiate between a justified negative answer (Certified-Negative) and insufficient evidence (Unknown). The primary issue in CROWN-Synth is asymmetric, as models often acknowledge implicitly complete evidence but mistakenly classify implicitly partial evidence as covering the query. This paper underscores the necessity for enhanced reasoning capabilities in LLMs regarding absence queries, impacting AI reliability in information retrieval and question-answering systems.

Key facts

  • CROWN-QA benchmark introduced for completeness-sensitive negative reasoning
  • Comprises CROWN-Synth and CROWN-Real evaluation sets
  • Tests across three LLM families
  • Models show unstable closure judgments and over-closure
  • Failure mode: asymmetric recognition of implicit evidence
  • Paper available on arXiv with ID 2608.04591
  • Focus on distinguishing Certified-Negative from Unknown
  • Implications for AI reliability in absence queries

Entities

Institutions

  • arXiv

Sources