Search papers, labs, and topics across Lattice.
This study investigates the effectiveness of LoRA-based supervised fine-tuning on large language models (LLMs) for coding student-generated mathematics metaphors, which are complex and time-consuming to analyze manually. By utilizing a human-coded dataset of 2,265 responses, the researchers evaluated the performance of both proprietary and open-weight models in valence-intensity and thematic coding tasks. The results demonstrated that fine-tuning significantly enhanced the performance and reliability of open-weight models, allowing them to compete with and often surpass proprietary models, thus enabling scalable and privacy-conscious analysis in mathematics education.
Fine-tuning compact open-weight LLMs can outperform proprietary models in coding complex student metaphor responses, revolutionizing educational assessment.
Student-generated metaphors about mathematics can reveal students'attitudes, beliefs, identities, and experiences, but human expert coding of these thematically and semantically complex open-ended responses is time-intensive and difficult to scale. This study examines whether LoRA-based supervised fine-tuning of large language models (LLMs) can improve their performance on codebook-guided coding tasks for student mathematics metaphors. We used a human-coded corpus of 2,265 Grade 6-8 responses to food- and animal-based metaphor prompts and instructed the LLMs to perform two coding tasks: valence-intensity coding to capture the direction and strength of students'affective orientations toward mathematics, and thematic coding to capture students'framings of mathematics as expressed through their metaphors. We compared two proprietary models, GPT-4o mini and GPT-5 mini, under prompt-only conditions with two open-weight models, DeepSeek-R1 1.5B and Mistral 7B, evaluated before and after fine-tuning. Results show that fine-tuning substantially improved the performance and run-to-run reliability of the open-weight models across both tasks relative to their base versions. The fine-tuned compact open-weight models became competitive with, and often outperformed, the proprietary prompt-only models. These findings suggest that compact open-weight LLMs can support scalable, locally controllable, and privacy-conscious AI-assisted measurement of students'metaphor responses in mathematics education.