ARTFEED — Contemporary Art Intelligence

Khan Academy's Khanmigo: AI Tutoring Quality Metrics and Experiments

ai-technology · 2026-08-13

Khan Academy, a pioneer in AI-powered tutoring for K-12 education, launched Khanmigo in 2023, an AI tutor leveraging large language models (LLMs). The organization has developed a comprehensive framework for evaluating and improving the quality of AI tutoring, emphasizing robust evaluation and live experimentation due to the opaque nature of LLMs. They describe the metrics used to measure AI tutoring quality and student engagement, along with various experiments conducted. Key improvements that have positively impacted their metrics include changes in models, prompting, personalization, and the use of agents. The work is detailed in a paper titled 'Methodologies for Improving the Quality of AI Tutoring in K-12 Education,' available on arXiv under the Computer Science > Computers and Society category. The paper outlines the challenges of working with black-box LLMs and the necessity of continuous measurement and iteration. Khan Academy's approach involves tracking specific metrics to gauge the effectiveness of tutoring interactions and student engagement, and they have run experiments to test different strategies. The findings highlight the importance of model selection, prompt engineering, personalized learning paths, and agent-based interactions in enhancing the AI tutoring experience. This research contributes to the broader field of AI in education, offering insights for developers and educators aiming to implement AI tutors in K-12 settings.

Key facts

  • Khan Academy launched Khanmigo, an AI tutor for K-12 education, in 2023.
  • Khanmigo leverages large language models (LLMs).
  • The paper describes metrics for measuring AI tutoring quality and student engagement.
  • Experiments were run to test changes in models, prompting, personalization, and agents.
  • The paper is available on arXiv under Computer Science > Computers and Society.
  • The work emphasizes robust evaluation and live experimentation due to LLMs being opaque black boxes.
  • Changes in models, prompting, personalization, and agents have moved their metrics.
  • The paper is titled 'Methodologies for Improving the Quality of AI Tutoring in K-12 Education'.

Entities

Institutions

  • Khan Academy
  • arXiv

Sources