Implicit Demographic Cues Alter LLM Outputs, Internal Signal Identified
A recent study published on arXiv (2608.11735) indicates that large language models (LLMs) adjust their responses based on implicit demographic indicators, even when users do not explicitly identify their demographic backgrounds. The research identifies a localized internal activation signal that monitors these shifts in recommendations, showing correlations as high as r=0.87 across five LLMs. When multiple demographic cues are present, their internal signals tend to merge, but the resulting output changes do not merely accumulate. Additionally, the study reveals that eliminating the internal signal linked to a specific cue can more effectively mitigate its impact than instructing the model to disregard demographics, while maintaining overall benchmark performance. Nonetheless, the capacity to selectively eliminate the influence of one dimension while keeping others intact is constrained. These findings enhance the understanding of implicit personalization in LLMs, raising important considerations for fairness and control in AI systems.
Key facts
- Study on arXiv:2608.11735
- Five LLMs tested
- Correlations up to r=0.87
- Internal signal tracks output changes
- Multiple cues combine internally but not additively in output
- Removing internal signal suppresses cue influence
- More effective than prompting to ignore demographics
- Benchmark performance largely preserved
Entities
Institutions
- arXiv