ARTFEED — Contemporary Art Intelligence

Research on LLMs Questions Self-Consistency of Preference Judgments

ai-technology · 2026-08-19

A recent study published on arXiv (ID 2608.17644) investigates the self-consistency of numerical preference judgments generated by large language models (LLMs). The researchers note that agents often interpret individuals' natural-language preferences by consulting an LLM for cardinal evaluations, such as determining how much someone might pay for a product. These evaluations help in estimating a utility function that influences decision-making. This approach assumes that a single utility function can accurately reflect the judgments. To evaluate this assumption, the study introduces statistical tests and interpretable metrics to measure deviations from the optimal self-consistent utility function. A significant example demonstrates that the difference in willingness-to-pay between two products should match the payment that would leave a person indifferent between them. The experiments encompass flight, apartment, and hotel selection scenarios across six distinct LLMs. The paper concludes that LLM-generated preference judgments lack self-consistency, highlighting a critical issue in using these outputs for preference modeling and decision-making.

Key facts

  • Paper posted on arXiv with ID 2608.17644 (announcement type: new)
  • Studies whether cardinal LLM preference judgments are self-consistent
  • Focuses on judgments like willingness-to-pay elicited via natural language
  • Common pipeline estimates a utility function from these judgments
  • Assumption tested: a single utility function can reproduce the judgments
  • Develops statistical tests and interpretable measures of departure from self-consistency
  • Experiments conducted on flight, apartment, and hotel examples
  • Six different LLMs were evaluated

Entities

Sources