LLMs and Humans Diverge on Contextual Creativity Standards
A recent investigation published on arXiv (2607.22218) explores the conditions under which large language models (LLMs) either align with or differ from human assessments of creativity. Researchers conducted three studies utilizing six popular LLMs and discovered that these models depend on a more limited range of human creativity evaluation criteria. The alignment with human standards is most pronounced in terms of novelty, whereas the greatest discrepancies occur in the contextual aspect, which encompasses social, market, and reputational factors. Each LLM displays unique evaluation criteria that differ significantly in scope. In Study 2, which analyzed 1,103 ideas, these variations in standards were evident in actual creativity assessments, underscoring the challenges of using LLMs for creativity evaluation without recognizing their inherent biases.
Key facts
- Study published on arXiv with ID 2607.22218
- Three studies conducted using six widely used LLMs
- LLMs rely on narrower subset of human creativity standards
- Strongest convergence with humans in novelty dimension
- Clearest divergence in contextual dimension (social, market, reputational)
- Each LLM has distinct, model-specific standards
- Study 2 involved 1,103 ideas
- Differences in standards reflected in actual creativity judgments
Entities
Institutions
- arXiv