HeuristicEdu Aligns LLMs as Socratic Tutors via Reinforcement Learning
HeuristicEdu, a two-phase pipeline, has been created by researchers to steer the Qwen2.5-7B large language model towards Socratic tutoring rather than simply providing answers. This approach employs supervised warm-up and Group Relative Policy Optimization (GRPO) utilizing SocraticEdu, a collection of 797 multi-turn dialogues in Chinese focused on children's science from a live platform. A heuristic reward function measures cognitive depth, curiosity engagement, and directness, incorporating a K_query adjustment for terms introduced by students. New metrics, Scaffolding Effectiveness (SE) and Conversation Depth (CD), evaluate results beyond mere fluency. The top GRPO variant enhanced SE from 30.0% to 63.3% and minimized keyword leakage from 30.0% to 13.3% on 30 withheld questions. This research was published on arXiv (2607.22996v1).
Key facts
- HeuristicEdu pipeline aligns Qwen2.5-7B toward Socratic tutoring
- Uses supervised warm-up and Group Relative Policy Optimization (GRPO)
- Trained on SocraticEdu dataset of 797 multi-turn Chinese children's science dialogues
- Heuristic reward over cognitive depth (R_cog), curiosity engagement (R_eng), and directness (R_dir)
- K_query correction for student-introduced terms
- New metrics: Scaffolding Effectiveness (SE) and Conversation Depth (CD)
- Best GRPO variant improved SE from 30.0% to 63.3% on 30 held-out questions
- Keyword leakage reduced from 30.0% to 13.3%
- Published on arXiv with ID 2607.22996v1
Entities
Institutions
- arXiv