LLM Test Oracles: Authority Taxonomy Systematic Review
A literature review conducted in accordance with PRISMA 2020 guidelines investigates the application of large language models (LLMs) as test oracles in software testing. Initially, 2,436 records were evaluated, leading to a selection of 54 studies, which was further expanded through citation searching (snowballing) to a total of 83. The research classifies oracles by their authoritative source, form, and adjudication mechanism. Interestingly, over half of the studies arrive at conclusions without any explicit criteria, depending solely on the model's training. The review emphasizes that two seemingly identical oracles can be based on different foundations: one may embody a written specification, while the other is merely a reflection of the model's training data. Previous secondary studies categorized oracles by form or technique, but this review centers on the trustworthiness of the verdict based on its authoritative origin. These insights are significant for both software testing and AI reliability, especially as LLMs increasingly determine software correctness.
Key facts
- Systematic literature review following PRISMA 2020 guidelines
- Screened 2,436 records to 54 included studies
- Extended via citation searching (snowballing) to 83 total
- Categorizes oracles by source of authority, form, and adjudication mechanism
- Just over half of the corpus reaches a verdict with no specification at all
- LLMs increasingly write or act as test oracles
- Two oracles can look identical but rest on different grounds
- Prior secondary studies sorted oracles by form or technique
Entities
—