Search papers, labs, and topics across Lattice.
This study extends the Learning Engagement Assistant (LEA) from a single STEM course simulation to its first real-world classroom deployment across three courses and two academic levels. The findings indicate that while Answer Relevancy and Context Precision remain stable across different courses, Faithfulness declines significantly as the curriculum distance increases from the original course. This divergence from simulation predictions underscores the limitations of synthetic evaluations in anticipating real-world performance of AI tutoring systems.
Real-world deployment of the LEA reveals that while it performs well across courses, its ability to maintain content fidelity diminishes with curriculum distance.
This paper is an extension of a paper presented at the ICAART 2026 conference, which introduced LEA (Learning Engagement Assistant), an adaptive AI tutoring agent combining course-specific Retrieval-Augmented Generation (RAG) with structured Knowledge Component (KC) models across integrated Chat, Tutor, and Quiz modes. That prior work validated LEA on a single STEM course (CMP511) exclusively through simulation, using synthetic learner agents. This paper extends that work by reporting the first classroom deployment of LEA with real students (n = 8, CMP511) and the first empirical test of its cross-course scalability, deploying the system across three courses spanning two academic levels and two disciplinary domains. The study reveals a divergence from simulation predictions across modes, showing that synthetic evaluation alone cannot anticipate all aspects of real deployment. A RAGAS-based cross-course scalability evaluation (660 questions) finds Answer Relevancy and Context Precision broadly stable across courses (0.88-0.94 and 0.88-0.90 respectively), while Faithfulness declines with curriculum distance from the system's original course (0.69 to 0.50), a preliminary finding that may reflect generation logic tuned to the system's original subject rather than a scalability limitation. These findings suggest that while the orchestration layer requires no modification, full course-agnosticism of all downstream components requires further investigation.