Search papers, labs, and topics across Lattice.
This paper introduces StudentSim, a novel training framework that creates individualized student simulators by leveraging pooled training data followed by per-student specialization. By addressing the limitations of existing models, StudentSim achieves higher behavioral fidelity and guidance responsiveness compared to GPT-5.4 across three domains: chess, second-language English writing, and mathematics. The framework not only enhances the accuracy of AI tutors but also demonstrates improved performance in reinforcement learning applications, as evidenced by expert evaluations of a chess tutor based on StudentSim.
StudentSim outperforms existing models, achieving a remarkable behavioral fidelity of 0.51 and guidance responsiveness of 0.91 in chess, setting a new standard for AI tutors.
AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but evidence about which guidance works for which student is sparse, slow, and costly to collect from real learners. Student simulators can provide this signal as a proxy, yet existing approaches are limited: state-tracking models fit student behavior but struggle to process explanations or corrections, while LLM role-play follows guidance fluently but does not reliably match the competence of the student being imitated. We present StudentSim, a training framework that turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization. The resulting simulators both mirror a student's own responses and update them under tutor guidance. We also introduce StudentSimEval, a standardized protocol covering 60 students across chess, second-language English writing, and mathematics, using public learner datasets with de-identified records shared for research. StudentSimEval measures behavioral fidelity (F), or how well a simulator matches a student's responses, and guidance responsiveness (R), or how readily it updates under tutor guidance, with all methods fit and evaluated on the same records. Across all three domains, StudentSim outperforms GPT-5.4 on both metrics. In chess, StudentSim reaches F=0.51 and R=0.91, compared with 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. As a proof of concept, using StudentSim as a reward model for tutor reinforcement learning produces a chess tutor that expert humans rate as more accurate, better-guided, and more personalized than a no-RL baseline and a tutor trained against a GPT-5.4 simulator reward. Code is available at https://github.com/microsoft/StudentSim.