LLMs Reproduce Racial Stereotypes in Text Annotation
A recent study published on arXiv (2603.13891) indicates that large language models (LLMs) consistently mirror racial stereotypes during automated text annotation. Analyzing 19 LLMs across two experiments with over 4 million annotation assessments, the research found that subtle identity indicators influenced results. In a name-based experiment involving 39 annotation tasks, texts linked to Black names were perceived as more aggressive by 18 out of 19 models and more gossipy by the same number. For Asian names, a 'bamboo-ceiling' effect emerged: 17 models viewed individuals as more intelligent, while 18 assessed them as less confident and sociable. Arab names prompted cognitive elevation but interpersonal devaluation, with all four minority groups rated as less self-disciplined. Additionally, a matched dialect experiment revealed that the same sentence received different judgments based on dialect. This study emphasizes the dangers of employing LLMs in content moderation, hiring, and academic research, highlighting the necessity for thorough auditing of AI tools in sensitive contexts.
Key facts
- Study on arXiv:2603.13891
- 19 LLMs tested
- Over 4 million annotation judgments
- Names-based experiment with 39 annotation tasks
- Black-associated names rated more aggressive by 18/19 models
- Black-associated names rated more gossipy by 18/19 models
- Asian names rated more intelligent by 17/19 models
- Asian names rated less confident and less sociable by 18/19 models
- Arab names elicited cognitive elevation and interpersonal devaluation
- All minority groups rated less self-disciplined
- Matched dialect experiment showed same sentence judged differently
Entities
Institutions
- arXiv