ARTFEED — Contemporary Art Intelligence

AIriskEval-edu Demo Audits Pedagogical Risks in Educational Explanations

ai-technology · 2026-07-29

The AIriskEval-edu Demo is an innovative tool designed to evaluate the quality of teaching explanations and provide clear audit results. It assesses these explanations through a framework that looks at five key areas of pedagogical risk: accuracy, thoroughness, relevance, suitability for students, and potential bias. Each area yields a yes or no result along with a confidence rating. When risks are detected, the tool offers a straightforward explanation, and for all but the thoroughness category, it includes specific supporting evidence. It operates using GPT-5.5 via an external API and a local Llama 3.1 8B evaluator that works with standard GPUs. This local evaluator is tailored using the AIriskEval-edu dataset, which features K-12 explanations marked for risk and clarity. The system supports two modes, where both evaluators analyze stored explanations from six different simulated scenarios, helping educators and content creators identify and mitigate risks in AI-generated teaching materials.

Key facts

  • AIriskEval-edu Demo audits pedagogical quality of instructional explanations.
  • Evaluates against five risk dimensions: factual accuracy, depth and completeness, focus and relevance, student-level appropriateness, ideological bias.
  • Returns binary decision and confidence score per dimension.
  • Detected risks include natural-language rationale and localized evidence span.
  • Integrates GPT-5.5 via external API and self-hosted Llama 3.1 8B evaluator on consumer-grade GPUs.
  • Local evaluator fine-tuned on AIriskEval-edu dataset of K-12 explanations.
  • Platform operates in two modes, including AI mode with both evaluators.
  • Designed for educators and content creators to mitigate risks in AI-generated content.

Entities

Sources