Adversarial AI Could Shield Educational Exercises from Cheating
A new study on arXiv (2608.01112) proposes using adversarial machine learning to protect educational tests from AI-driven cheating. It introduces a method involving multimodal multiple-choice questions that have slight visual tweaks designed to mislead AI into selecting certain wrong answers. If students try to copy AI responses, they would show these altered patterns, which would help identify them as non-genuine. The research tests this method in real-world black-box situations with three cutting-edge multimodal language models: Anthropic’s Claude, Google’s Gemini, and OpenAI’s ChatGPT. This work is a reaction to the rise of generative AI, which enables students to avoid engaging in critical thinking and independent thought. The paper was recently published on arXiv.
Key facts
- The paper is available on arXiv with identifier 2608.01112.
- The approach uses multimodal multiple-choice questions with visual perturbations.
- Perturbations steer AI solvers toward designated incorrect answers.
- Students who copy AI responses reproduce the induced answer pattern.
- The method is tested with Anthropic's Claude, Google's Gemini, and OpenAI's ChatGPT.
- The research aims to protect educational exercises against AI cheating.
- The paper is announced as new on arXiv.
- The study assumes realistic black-box assistant conditions.
Entities
Institutions
- Anthropic
- OpenAI
- arXiv