Search papers, labs, and topics across Lattice.
This study investigates the effectiveness of AI-generated written corrective feedback (WCF) in English as a Foreign Language (EFL) contexts by evaluating over 20,000 essay drafts from nearly 2,000 students. The research reveals a significant disconnect between expert evaluations by experienced teachers and the feedback preferences of students, indicating that traditional assessment methods may overlook critical aspects of usability and helpfulness. These findings underscore the necessity for learner-centered evaluation frameworks to better align AI feedback with pedagogical best practices in language education.
AI-generated feedback may not meet learner needs, as it often misaligns with expert evaluations, revealing a critical gap in language education tools.
This study examines feedback in English as a Foreign Language (EFL) writing contexts, focusing on written corrective feedback (WCF). Large language models (LLMs) can provide WCF at scale, but aligning them with pedagogical best practices remains an ongoing challenge. WCF meeting criteria like factuality or relevance may still be unsuitable for learning contexts, highlighting the need for extrinsic evaluation based on the learner's perspective. We deployed WCF systems in a university-level EFL class with nearly 2,000 students, collecting over 20,000 drafts. We evaluated the generated WCF from two perspectives: intrinsic evaluation by experienced English teachers using a rubric, and extrinsic evaluation via student feedback and engagement metrics. Results revealed low alignment between teacher expert ratings and student feedback. These findings suggest that traditional expert evaluation alone may not fully capture WCF's usability or helpfulness from the learner's perspective, highlighting the importance of learner-centered evaluation frameworks for AI-based applications in language education.