ARTFEED — Contemporary Art Intelligence

Study Finds AI Alignment Research Lacks Clear Definition of Human Values

ai-technology · 2026-08-13

A new paper on arXiv (2608.10327) examines the concept of human values in AI alignment research, finding that most studies do not define values and instead rely on preferences as a proxy. The authors annotated 94 value alignment research papers to discern the implicit theory of values in AI. They found that the majority of papers do not define values, with many using preferences as a stand-in. The paper raises questions about how the field of AI value alignment conceives of human values, how these conceptions are technically operationalized and evaluated, and what the emergent theory of value signifies for the future of AI. The study highlights the rise of harms from large language models (LLMs) and multi-modal foundation models, including toxic speech, hallucinations, and unauthorized actions by AI agents. These issues are often framed as the alignment problem, or models being misaligned with human values. The paper calls for more explicit definitions of values in AI alignment research.

Key facts

  • Paper on arXiv: 2608.10327
  • Annotated 94 value alignment research papers
  • Majority of papers do not define values
  • Preferences used as a stand-in for values
  • Harms include toxic speech, hallucinations, unauthorized actions
  • LLMs and multi-modal foundation models are the focus
  • Field of AI safety frames these as alignment problem
  • Authors question operationalization and evaluation of values

Entities

Institutions

  • arXiv

Sources