LLM Reliability as Surrogate Experts in Security Surveys: New arXiv Study
A newly published framework on arXiv (identifier 2608.16893) evaluates large language models (LLMs) as alternative experts in security research surveys. It tackles the difficulty of engaging domain specialists in Security Operations Centres (SOCs), which often face high workloads and confidentiality issues. This research suggests using LLMs to create synthetic responses and assesses their reliability in comparison to those provided by SOC professionals. The study analyzes both persona-based and aggregate LLM responses, focusing on their consistency and alignment with human input. Findings indicate that while LLMs yield stable answers, they differ from human experts, prompting concerns about their credibility. This framework seeks to define when LLM-generated information can be deemed trustworthy, influencing social science research with limited expert involvement.
Key facts
- Paper released on arXiv under identifier 2608.16893.
- Announcement type is marked as 'cross'.
- Focuses on using LLMs as surrogate experts in security surveys.
- Highlights difficulty of recruiting experts in Security Operations Centres (SOCs).
- Proposes a framework for evaluating LLMs as substitutes or supplements to expert survey respondents.
- Compares persona-based and aggregate LLM-generated answers across multiple models and prompting settings.
- Measures stability, inter-model agreement, and alignment with human responses.
- Finds LLMs produce internally consistent answers but systematically diverge from expert responses.
Entities
—