Search papers, labs, and topics across Lattice.
This study evaluates the performance of a LoRA-adapted large language model, Gemma-3-27B-it, in automated essay scoring across various first-language backgrounds using the TOEFL11 corpus. The model demonstrated strong cross-prompt generalization with an overall band agreement of 77.79% and a quadratic weighted kappa of 0.702, indicating consistent scoring accuracy across unseen prompts. However, a notable first-language bias was identified, with essays from European-language backgrounds receiving higher scores than those from East-Asian-language backgrounds, highlighting potential fairness issues in automated scoring systems.
A systematic first-language bias in automated essay scoring reveals that European-language essays consistently outscore their East-Asian counterparts, raising critical fairness concerns.
This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and inference configuration reported in"AiAWE: An Open-Source LLM Automated Writing Evaluation System Using LoRA-Adapted Instruction-Tuned Models"(Gayed, 2026), which was fine-tuned on 480 argumentative essays from two prompts, we evaluate scoring accuracy on the full TOEFL11 corpus: 12,100 essays written by test-takers from 11 first-language backgrounds across eight prompts, none of which were seen during training. The model's raw scores (0.5-5.0) are mapped to the same three proficiency bands (low, medium, high) used by ETS, enabling direct comparison. The model achieved an overall band agreement of 77.79% and a quadratic weighted kappa of 0.702, with adjacent-band agreement of 99.98%. Accuracy was stable across all eight unseen prompts, with no advantage for prompts thematically related to the training data, indicating robust cross-prompt generalization. However, the model exhibited a systematic, L1-linked scoring offset. Within every proficiency band, essays from European-language backgrounds received consistently higher scores than essays from East-Asian-language backgrounds, a pattern not attributable to the composition of the fine-tuning data. This is the first large-scale L1 fairness analysis of a fine-tuned open-weight LLM for automated essay scoring.