MMGrader: AI Tool Assesses Mental Models from Multimodal Student Answers
Researchers have introduced a novel AI system named MMGrader, designed to evaluate the quality of students' mental models through their multimodal responses. This methodology, outlined in a paper available on arXiv (2603.00056v2), employs concept graphs to analyze and gauge the depth of conceptual understanding. The evaluation involved 9 publicly accessible models, including vision-language models, revealing that the top-performing models reached only around 40% accuracy, with a prediction error of 1.1 units. While the scoring distribution showed some alignment with human evaluations, it did not meet human-level performance. The paper emphasizes the difficulty of deducing mental models from student responses, highlighting both the promise and limitations of AI in educational assessment.
Key facts
- MMGrader is an approach that infers the quality of students' mental models from multimodal responses.
- It uses concept graphs as an analytical framework.
- The evaluation involved 9 openly available models.
- The best-performing models achieved approximately 40% accuracy.
- The prediction error was 1.1 units.
- The scoring distribution was fairly aligned with human judgments.
- The models fell short of human-level performance.
- The paper is available on arXiv with ID 2603.00056v2.
Entities
Institutions
- arXiv