Search papers, labs, and topics across Lattice.
EduPanel is a novel three-agent LLM judge designed to evaluate teaching videos by leveraging multimodal evidence and learner-specific contexts. The system achieves reliability comparable to human experts, with significant improvements in scoring accuracy (MAE reduced from 0.87 to 0.73) while maintaining human oversight through the ability to detect unreliable outputs (AUC = 0.77). This work highlights the potential of AI to enhance educational evaluation without fully replacing human evaluators, addressing a critical gap in scalable assessment methodologies.
EduPanel achieves human-level reliability in evaluating teaching videos while enhancing scoring accuracy and preserving expert oversight.
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality. Across expert studies, architecture ablations, and learner-persona analyses, EduPanel achieves reliability comparable to a median human expert. In expert evaluation, its feedback improves scoring accuracy (MAE 0.87 to 0.73), while experts remain able to detect unreliable outputs (AUC = 0.77) instead of accepting them blindly. These results suggest that EduPanel can serve as effective assistants for educational evaluation rather than replacements for human experts.