Search papers, labs, and topics across Lattice.
This paper introduces Gazette, a novel framework for decoding human gaze into natural language descriptions of goals, moving beyond traditional categorical approaches. By framing gaze decoding as a generative learning problem, Gazette leverages multimodal large language models to produce nuanced, free-form descriptions that capture the complexities of human intentions. The method outperforms existing techniques across various tasks, showcasing its ability to generalize and effectively interpret gaze as a non-intrusive indicator of human goals and intentions.
Gazette transforms gaze data into rich, natural language descriptions of human intentions, revealing the depth of human attention beyond simple categories.
We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, which frames gaze decoding as a discriminative task over predefined categories, we formulate it as a generative learning problem: training a model to produce free-form descriptions that capture the rich nuances and open-ended nature of human intentions beyond fixed labels. To this end, we introduce Gazette, the first gaze-to-text decoding framework. Based on multimodal large language models (MLLMs), Gazette learns to decode gaze scanpaths into natural language for goals that may extend beyond categorical labels and require articulation in natural language. To help Gazette filter out individual differences in gaze behavior and learn the goal-specific spatiotemporal dynamics crucial for generating accurate natural language goal descriptions, we propose a novel strategy that leverages the encyclopedic knowledge and reasoning abilities of a large language model to synthesize natural language explanations of goal-directed attentional behavior called think-aloud transcripts. Instruction tuning on these synthetic narratives allows Gazette to achieve state-of-the-art performance in gaze decoding across multiple tasks, demonstrating its generalizability and versatility, thereby enabling gaze to serve as a powerful, non-intrusive cue for inferring human goals and intentions in diverse scenarios.