LLMs Reproduce Human Social Meaning Structure but Not Magnitude
A recent study on arXiv investigates the ability of large language models (LLMs) to mimic human social meaning. Researchers developed two innovative metrics, Effect Size Ratio (ESR) and Calibration Deviation Score (CDS), to evaluate how these models reflect social structures and calibrate their output. The research centered on the premise that social meaning is derived from language interpretation and speaker intent. Analyzing three advanced LLMs, the study found that although all models effectively represent qualitative social reasoning, they vary significantly in their accuracy regarding magnitude calibration. This work suggests that integrating insights from pragmatic theory could enhance LLM alignment with human thought processes.
Key facts
- Paper on arXiv:2604.02512v2
- Introduces Effect Size Ratio (ESR) and Calibration Deviation Score (CDS)
- Case study on numerical (im)precision
- Three frontier LLMs tested
- Models reproduce qualitative structure of human social inferences
- Magnitude calibration differs across models
- Prompting strategies derived from pragmatic theory
- Study addresses quantitative approximation of social meaning
Entities
Institutions
- arXiv