Search papers, labs, and topics across Lattice.
Addressing the semantic limitations of traditional 2D point-of-gaze (PoG) estimation, this work formulates driver attention as an end-to-end gaze-object prediction task supported by the newly collected Urban Driving-Face Scene Gaze (UD-FSG) benchmark. Directly mapping facial and ocular cues to detected scene objects bypasses noisy geometric post-processing, providing a more robust visual awareness signal in complex driving environments. Evaluated on real-world driving data, the framework achieves 60% gaze-object prediction accuracy鈥攕ignificantly outperforming conventional two-stage PoG association (51%) while cutting object-background confusion by 49.7%.
Directly mapping driver facial cues to traffic objects via cross-attention outperforms intermediate point-of-gaze pipelines, slashing background-object misclassification errors by 49.7%.
Driver gaze provides information regarding driver visual attention and situational awareness to the surrounding traffic. Existing driver gaze estimation studies represent gaze in terms of gaze zone or gaze vector/point-of-gaze (PoG). However, object-level gaze information provides a more semantically meaningful representation of visual attention by identifying attended objects, such as vehicles, pedestrians, or traffic signals. In this study, we propose an end-to-end driver gaze object prediction framework, TransGaze-Object, Transformer-based Gaze Object prediction model. The proposed framework first extracts facial features, including face and iris-weighted eye features, along with trafficobject spatial features. A transformer based cross-attention mechanism is then used to compute similarity scores and attention weights for predicting the drivers gaze object. To train this model, we propose a benchmark driver gaze dataset, Urban Driving-Face Scene Gaze (UD-FSG), comprising synchronized driver-face and traffic-scene images, scene objects bounding boxes, and gaze labels in terms of 2D gaze coordinate and gaze object. The TransGaze-Object model achieves an overall accuracy of 60% for gaze-object prediction, compared to 51% accuracy obtained from associating the estimated Point-of-Gaze to traffic objects. The error analysis reveals that TransGaze-Object reduces confusion between traffic objects (predicted) and the background (ground-truth), achieving an error rate of 11.68%, a 49.7% relative reduction compared with 23.21% error obtained from PoG-based gaze-object association. Overall, the results demonstrate the effectiveness of directly predicting gaze objects from driver-face and traffic-scene information, rather than estimating an intermediate Point-of-Gaze and subsequently associating it with traffic objects.