Search papers, labs, and topics across Lattice.
This study investigates fairness-aware multimodal temporal models for estimating student attention using a comprehensive dataset that integrates facial images, wearable sensor data, and demographic metadata. The proposed multimodal model outperforms traditional baselines in mean test performance and exhibits the lowest worst-group error, although the improvements over simpler models are modest. Notably, while targeted regularization can reduce demographic disparities in validation data, these benefits do not consistently extend to new subjects, highlighting the need for careful evaluation of fairness in predictive models.
Fairness-aware modeling can improve student attention estimation but may not generalize across diverse demographic groups, revealing critical gaps in current evaluation practices.
Automated student-attention estimation can support learning analytics, but aggregate predictive metrics can conceal demographic disparities. This study evaluates fairness-aware multimodal temporal models on DIPSER, a naturalistic classroom dataset combining facial images, wearable-sensor measurements, attention annotations, and automatically inferred demographic metadata. Three baselines are compared across 10 training seeds: a Visual GRU, a Sensor GRU, and a Residual Fusion Transformer. The multimodal model achieves the best mean test performance (MAE 0.283, RMSE 0.363) and the lowest worst-group error among the evaluated baselines, although its gain over the Visual GRU is modest. Gender- and age-targeted MAE-gap regularization reduces disparities on validation data, but these gains do not consistently transfer to held-out subjects or repeated subject-level splits. On an NVIDIA A100-SXM4-40GB GPU, the warm end-to-end pipeline averages 50.65 ms per prediction window at a one-second stride, while the temporal model itself requires 1.02 ms. The findings show that multimodal fusion can modestly improve prediction and worst-group performance, but validation-level fairness gains should not be assumed to generalize. Robust fairness assessment therefore requires subgroup-aware evaluation, repeated subject-level validation, and larger, better balanced demographic samples.