ARTFEED — Contemporary Art Intelligence

LLM Moral Judgments Diverge from Human Rationales Despite Label Agreement

ai-technology · 2026-08-15

A new preprint on arXiv (2608.12368) challenges the common approach of evaluating large language models (LLMs) based on how well they align with human judgments. The authors argue that simply reaching the same conclusions doesn’t mean that both LLMs and humans share the same ethical beliefs, as differing contexts or interpretations can lead to similar outcomes. To explore this idea, they created a benchmark with 500 items from the ETHICS dataset across five moral areas and obtained annotations from both humans and LLMs, including final labels and justifications. While there was often strong agreement on labels, the reasoning behind those labels revealed significant differences in moral reasoning between humans and models. This highlights the need for better evaluation methods that go beyond just label matching.

Key facts

  • Preprint arXiv:2608.12368
  • Title: 'Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments'
  • Uses a curated 500-item ETHICS-derived benchmark
  • Covers five domains of moral judgment
  • Includes new human annotator and LLM annotations of final labels and rationales
  • Finds high agreement with human majority labels across frontier and open model families
  • Reveals systematic divergence in moral grounds at rationale level
  • Models redistribute attention across categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance

Entities

Institutions

  • arXiv

Sources