Search papers, labs, and topics across Lattice.
This paper introduces G3Ego, a novel graph-based framework that leverages gaze as a structural cue to enhance egocentric action understanding by focusing on relevant hand-object interactions. By constructing action scene graphs from vision-language descriptions and grounding objects, G3Ego prunes irrelevant entities based on the wearer's gaze, leading to more efficient and interpretable representations. Experiments reveal that G3Ego outperforms traditional video-based methods in class-imbalanced scenarios, demonstrating the potential of gaze-guided graph representations in this domain.
G3Ego reveals that integrating gaze into graph construction can significantly enhance the efficiency and accuracy of egocentric action recognition.
Egocentric action understanding is often addressed using large video models pretrained on extensive exocentric datasets. However, many first-person actions depend on a small number of hand-object interactions involving only a few relevant entities. We propose G3Ego, a graph-based framework for egocentric action understanding that uses gaze as a structural cue to identify action-relevant entities in the scene. From sparsely sampled frames, G3Ego constructs action scene graphs from vision-language descriptions, grounded objects, and hand cues, and then prunes irrelevant entities using the camera wearer's gaze. The resulting graph embeddings are temporally aggregated for action recognition and anticipation. Unlike prior work that uses gaze primarily as an auxiliary modality or attention signal, G3Ego incorporates gaze directly into graph construction, producing efficient and interpretable representations focused on action-relevant interactions. Experiments on EGTEA Gaze+ and MECCANO show that G3Ego achieves competitive performance compared with video-based approaches and consistently improves Macro-F1 under class-imbalanced evaluation, while avoiding reliance on computationally expensive video pretraining. These results demonstrate the effectiveness of gaze-guided graph representations for egocentric action understanding.