ARTFEED — Contemporary Art Intelligence

LLM Screening in Systematic Reviews: Batch Effects Outweigh Prevalence Metadata

ai-technology · 2026-08-18

A new study from arXiv examines the performance of large language models (LLMs) in imbalanced binary classification tasks, specifically focusing on study screening for systematic reviews. The research, conducted across five reviews, compared individual and batch processing methods, with and without the inclusion of prevalence metadata. The findings indicate that prevalence metadata has limited influence and does not improve performance. In contrast, batch processing led to larger behavioral changes that varied depending on class prevalence. Notably, aggregate and item-level analyses did not always align, suggesting that batch processing should be evaluated not only for cost efficiency but also for its impact on decision-making behavior. The study contributes to the growing body of research on LLM applications in academic and medical contexts, highlighting the need for careful consideration of processing methods in systematic review workflows.

Key facts

  • Study analyzes LLMs in imbalanced binary classification using systematic review screening.
  • Experiment conducted in five reviews.
  • Compared individual and batch processing, with and without prevalence metadata.
  • Prevalence metadata had limited influence and did not improve performance.
  • Batch processing produced larger behavioral changes varying by class prevalence.
  • Aggregate and item-level analyses did not always coincide.
  • Batch processing should be evaluated for effects on decision-making behavior, not just cost.
  • Study published on arXiv with ID 2608.14737.

Entities

Institutions

  • arXiv

Sources