EduClaw-Bench: Benchmarking LLM Agents in 30-Day Simulated Tutoring
A new benchmark called EduClaw-Bench has been introduced in an arXiv paper (2608.03206) to assess large language model (LLM) agents acting as tutors within a long-term educational framework. Unlike current solutions that focus on isolated tasks, EduClaw-Bench simulates a 30-day continuous interaction between a tutor agent and a learner. The learner's responses are informed by a knowledge tracing (KT) model, which utilizes real-student data to evaluate learning gains across 55 different scenarios. Agents are evaluated based on three main criteria: learning gain, responsiveness, and helpfulness, along with two curriculum design criteria rooted in Gagné and Rosenshine principles. This benchmark fills a gap in evaluating tutor agents over extended periods, as traditional tutoring typically spans days or weeks. The paper also serves as a cross-type announcement on arXiv, suggesting it may have been presented or published in other venues. This work is crucial for advancing AI in education by offering a standardized approach to measure the effectiveness of LLM tutoring agents over time.
Key facts
- EduClaw-Bench is a benchmark for evaluating LLM agents as tutors.
- It simulates a 30-day continuous relationship with a simulated learner.
- The learner is grounded in knowledge tracing (KT) trained on real-student data.
- The benchmark includes 55 scenarios for evaluating learning gain.
- Agents are scored on learning gain, responsiveness, and helpfulness.
- Two curriculum-design axes are based on Gagné and Rosenshine principles.
- The paper is an arXiv cross-type announcement (2608.03206).
- It addresses the lack of long-horizon evaluation in educational AI.
Entities
Institutions
- arXiv