Search papers, labs, and topics across Lattice.
This paper introduces EduZone, a comprehensive evaluation framework designed to assess the safety of large language models (LLMs) specifically within K-12 educational contexts. By systematically integrating various usage scenarios, curriculum concepts, and a detailed categorization of risks, the framework generates adversarial interactions to evaluate LLMs across multiple safety levels. The findings highlight significant vulnerabilities in LLMs, particularly in dynamic multi-turn conversations, indicating that current safety measures are insufficient for addressing education-specific risks.
Existing safety evaluations miss critical vulnerabilities in LLMs used in K-12 education, revealing that dynamic interactions pose unique risks that current guardrails fail to mitigate.
Large language models (LLMs) are increasingly used across diverse tasks in K-12 education, yet existing safety evaluations rarely examine how harmful or inappropriate content appears in interactions between LLMs and students or teachers. To address this, we present EduZone, an evaluation framework for LLM safety across diverse educational scenarios. Our framework systematically combines (1) student- and teacher-facing LLM usage contexts, (2) fine-grained curriculum concepts, and (3) 6 risk categories and 28 subcategories spanning both conventional and education-specific harms to generate contextually grounded adversarial interactions. We construct these interactions in three settings: single-turn requests, static multi-turn conversations, and dynamic multi-turn conversations. Using these interactions, we evaluate ten LLMs using four safety levels: refusal, safe assistance, risky assistance with safety guidance, and fully risky assistance. Our results reveal greater vulnerability to education-specific risks and dynamic multi-turn interactions, while existing safety guardrails fail to adequately address these risks. EduZone advances LLM safety in education by providing an automated, scalable evaluation framework that supports the development and deployment of safer LLMs in K-12 education.