Supervisor-Derived Weights vs. Defaults in AI Thesis Assessment
A new study on arXiv (2608.00717) examines how thesis supervisors prioritize evaluation criteria and how their input affects AI-based assessment systems. The research surveyed 84 supervisors across four academic disciplines, collecting weighting data for 35 criteria. Comparing these with the default weights of the AI system RubiSCoT revealed substantial divergences. The supervisor-derived weights were then integrated into calibration configurations and tested on a corpus of 80 German-language theses. The study highlights the gap between expert judgment and actual practice, suggesting that AI assessment tools may need recalibration to reflect real-world priorities. The findings have implications for the design of rubric-based AI systems in education, emphasizing the need for empirical grounding of criterion weights. The study is part of ongoing efforts to improve the transparency and fairness of automated assessment.
Key facts
- Study on arXiv: 2608.00717
- Surveyed 84 thesis supervisors
- Four academic disciplines covered
- 35 thesis assessment criteria used
- Compared with RubiSCoT default weights
- Substantial divergences found
- Tested on 80 German-language theses
- Implications for AI-based assessment
Entities
Institutions
- arXiv
- RubiSCoT