ARTFEED — Contemporary Art Intelligence

Pedagogical Suitability Index (PSI) Introduced to Evaluate LLM-Based AI Tutors

ai-technology · 2026-08-07

A new research paper on arXiv (2608.05411) introduces the Pedagogical Suitability Index (PSI), a composite metric designed to evaluate how well large language model (LLM) responses align with learner readiness and curricular progression in AI tutoring contexts. The study addresses a gap in existing evaluations, which focus primarily on answer correctness rather than instructional fit. PSI comprises six theory-informed sub-scores that assess aspects such as alignment with the learner's current foundation, course sequence, and timing of concept introduction. The researchers evaluated four LLM tutors—ChatGPT, Gemini, Gemma4, and Qwen3—across 240 scenario-based evaluations using paired standard and defective prompts. They then applied a PSI-guided regeneration protocol to 62 weak-performing responses to improve their pedagogical suitability. The paper is authored by researchers and was announced as a new submission on arXiv. The work highlights the growing importance of pedagogical considerations in AI-based education tools.

Key facts

  • The paper is available on arXiv under identifier 2608.05411.
  • The Pedagogical Suitability Index (PSI) is a composite metric with six sub-scores.
  • PSI evaluates alignment with learner readiness and curricular progression.
  • Four LLM tutors were evaluated: ChatGPT, Gemini, Gemma4, and Qwen3.
  • The study used 240 scenario-based evaluations with paired standard and defective prompts.
  • A PSI-guided regeneration protocol was applied to 62 weak-performing responses.
  • Existing evaluations focus mainly on answer quality, not instructional fit.
  • The research addresses the gap in measuring pedagogical appropriateness of AI tutor responses.

Entities

Institutions

  • arXiv

Sources