Psychometric Evaluation of AI with Situational Judgment Tests
A recently introduced framework suggests evaluating the consistent behavioral patterns of large language models (LLMs) through situational judgment tests (SJTs) and multidimensional item response theory (MIRT). The research indicates that behaviors influenced by persona remain stable across different runs, and latent trait scores can forecast external benchmarks such as TruthfulQA and EmoBench. MIRT uncovers a reliable latent structure. Validation of these findings comes from human annotation, benchmark assessments, and analyses of internal consistency. The identified traits are understood as stable behavioral patterns rather than reflections of human personality.
Key facts
- Framework uses SJTs and MIRT to measure LLM behavioral tendencies
- Persona-conditioned behaviors are stable across runs
- Latent trait scores predict TruthfulQA and EmoBench benchmarks
- MIRT reveals consistent latent structure
- Results validated via human annotation and internal consistency analyses
- Traits are stable behavioral tendencies, not human personality
- Study uses large-scale SJT and persona datasets
- arXiv paper 2510.22170
Entities
Institutions
- arXiv