AI Benchmarks Miss Modality, Search, and Citations: Implications for Safety Evaluations
A recent preprint on arXiv (2608.06202) examines the foundational assumptions of benchmark evaluations for large language models (LLMs), which are often used to validate claims regarding model safety, reliability, and readiness for deployment. Researchers analyzed two access methods—ChatGPT's chat interface and OpenAI's API—with and without web search capabilities, utilizing a stratified sample of 401 prompts from the widely used benchmarks, BBQ and SafetyBench. They gathered a total of 4,812 responses over three iterations for each prompt. The study assesses not only conventional performance metrics but also response consistency, similarity of response text, citation grounding, and abstention behavior. Findings indicate that existing benchmarks, which typically utilize a single access method and conduct one run per prompt while emphasizing accuracy, overlook factors like web search that could influence model performance in real-world applications. This research underscores the necessity for more thorough evaluation frameworks that incorporate these elements to improve assessments of model safety and reliability.
Key facts
- Preprint on arXiv: 2608.06202
- Compares ChatGPT chat UI and OpenAI API
- With and without web search enabled
- 401 prompts from BBQ and SafetyBench
- 4,812 total responses collected
- Three repeated runs per prompt
- Evaluates response consistency, text similarity, citation grounding, and abstention
- Current benchmarks rely on single modality, single run, and accuracy
Entities
Institutions
- arXiv
- OpenAI