ARTFEED — Contemporary Art Intelligence

EduClaw-Bench: Benchmarking LLM Agents in 30-Day Simulated Tutoring

ai-technology · 2026-08-06

A new benchmark called EduClaw-Bench has been introduced in an arXiv paper (2608.03206) to assess large language model (LLM) agents acting as tutors within a long-term educational framework. Unlike current solutions that focus on isolated tasks, EduClaw-Bench simulates a 30-day continuous interaction between a tutor agent and a learner. The learner's responses are informed by a knowledge tracing (KT) model, which utilizes real-student data to evaluate learning gains across 55 different scenarios. Agents are evaluated based on three main criteria: learning gain, responsiveness, and helpfulness, along with two curriculum design criteria rooted in Gagné and Rosenshine principles. This benchmark fills a gap in evaluating tutor agents over extended periods, as traditional tutoring typically spans days or weeks. The paper also serves as a cross-type announcement on arXiv, suggesting it may have been presented or published in other venues. This work is crucial for advancing AI in education by offering a standardized approach to measure the effectiveness of LLM tutoring agents over time.

Key facts

  • EduClaw-Bench is a benchmark for evaluating LLM agents as tutors.
  • It simulates a 30-day continuous relationship with a simulated learner.
  • The learner is grounded in knowledge tracing (KT) trained on real-student data.
  • The benchmark includes 55 scenarios for evaluating learning gain.
  • Agents are scored on learning gain, responsiveness, and helpfulness.
  • Two curriculum-design axes are based on Gagné and Rosenshine principles.
  • The paper is an arXiv cross-type announcement (2608.03206).
  • It addresses the lack of long-horizon evaluation in educational AI.

Entities

Institutions

  • arXiv

Sources