ARTFEED — Contemporary Art Intelligence

LLM-Based Fusion Model Predicts Item Acceptance in Standardized Testing

ai-technology · 2026-08-10

A recent study presents an automated item evaluation (AIE) model designed to forecast the acceptance or rejection of test items for operational purposes, utilizing critiques generated by large language models (LLMs). This research is outlined in a preprint available on arXiv (2608.06609) and is based on historical rejection data from a significant standardized testing initiative. The dataset included 52,759 items in English language arts (ELA) and mathematics, with 34% being permanently rejected. Reasons for rejection encompassed inadequate psychometric properties, content-related issues, bias and sensitivity concerns, as well as non-content factors. The researchers optimized a DeBERTaV3-large classifier on the raw text of items, another DeBERTa classifier on critiques generated by Qwen3, and a fusion model that integrated both representations. The fusion model demonstrated the best overall performance, achieving an accuracy of .75, an F1 score of .64, an AUC of .80, sensitivity of .64, and specificity of .80. This method seeks to lessen dependence on manual expert evaluations and field testing, potentially enhancing item development in educational assessments. The study underscores the increasing role of AI in educational measurement, although the authors emphasize the necessity for additional validation across various item types and testing programs.

Key facts

  • The study was published on arXiv with identifier 2608.06609.
  • The dataset contained 52,759 English language arts (ELA) and mathematics items.
  • 34% of items were permanently rejected from future operational use.
  • Rejection reasons included poor psychometric properties, content issues, bias and sensitivity concerns, and non-content issues.
  • The model used DeBERTaV3-large and DeBERTa classifiers, along with Qwen3-generated critiques.
  • The fusion model achieved Accuracy = .75, F1 = .64, AUC = .80, Sensitivity = .64, Specificity = .80.
  • The research aims to automate item evaluation without manual expert review or field testing.

Entities

Institutions

  • arXiv

Sources