Search papers, labs, and topics across Lattice.
This study investigates the representation of visual evidence in predicting item difficulty, comparing traditional text-based methods with two novel approaches: visual textualization and image-native modeling. By training large language models (LLMs) and vision-language models (VLMs) on Eedi items, the research finds that both visual interfaces yield lower root mean square error (RMSE) estimates compared to text-only methods, although they exhibit complementary errors and varying effectiveness based on model adaptation. The findings suggest that image-native modeling is a viable alternative to textualization, challenging the notion that textual representation is the sole effective approach for integrating visual components in item difficulty prediction.
Visual textualization and image-native modeling outperform traditional text-only methods in predicting item difficulty, revealing the critical role of visual evidence in assessment design.
Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are available. Existing approaches typically represent the question stem and answer choices as text. When mathematics items contain visual components, a common pipeline first textualizes that evidence and then applies a text predictor. We ask: how should visual evidence be represented for item difficulty prediction? We compare question text alone, visual textualization, which expresses visual evidence in language, and image-native modeling, which retains the original image. Using Eedi items with difficulty calibrated from student responses, we train large language models (LLMs) and vision-language models (VLMs) directly for difficulty regression. Both visual interfaces achieve the lowest point estimates, although the leading systems cannot be reliably ordered. Open-VLM textualization yields lower RMSE point estimates for all evaluated LLMs, while broader adaptation does so for all image-native VLMs. Test-time interventions show dependence on the paired full-item image, but do not isolate the additional visual component. The two visual interfaces also make partially complementary item-level errors and differ substantially in computational workflow. Thus, textualization should not be treated as the only practical interface: image-native modeling is a competitive alternative whose effectiveness depends on how the VLM is adapted.