Persona Conditioning Exposes LLM Assessor Sensitivity in IR Evaluation
A recent preprint on arXiv (2608.10385) explores the dependability of large language models (LLMs) in assessing relevance for information retrieval (IR) evaluations. The research presents persona conditioning as a method to reveal the sensitivity of assessors. Five unique assessor roles were created, focusing on intent interpretation, domain knowledge, contrastive judgment, evidence verification, and overall search quality assessment, utilizing personas from PersonaHub and NVIDIA Nemotron-Personas-USA. These roles were benchmarked against a standard UMBRELA baseline across six LLM architectures using TREC DL20 and RAG24 datasets. Results indicate that while judgments typically align with the baseline, variations occur in strictness, evidential thresholds, or emphasis on interpretation, rather than causing extensive relevance shifts. This study informs the design and understanding of LLM-based evaluations in IR, highlighting how assessor framing can systematically affect results. The paper is accessible on arXiv.
Key facts
- arXiv preprint 2608.10385
- LLMs used as relevance assessors in IR evaluation
- Persona conditioning as diagnostic mechanism
- Five assessor roles: intent interpretation, domain expertise, contrastive judgment, evidence verification, global search-quality assessment
- Personas from PersonaHub and NVIDIA Nemotron-Personas-USA
- Six LLM backbones tested
- Datasets: TREC DL20 and RAG24
- Sensitivity is structured, not uniform
Entities
Institutions
- arXiv
- PersonaHub
- NVIDIA Nemotron-Personas-USA
- TREC DL20
- RAG24