Search papers, labs, and topics across Lattice.
To overcome the limitations of clip-level supervision that discards continuous temporal variation, the authors curate Kairos, a video-language dataset composed of 10-to-30-minute videos with dense, time-resolved annotations. The dataset systematically tracks ongoing actions, entity attributes, interactions, and contextual evolution along continuous temporal timelines. This provides a rigorous benchmark and training bed for long-range spatial-temporal reasoning, dynamic state tracking, and fine-grained video instruction tuning.
Clip-level captions compress away the continuous visual dynamics needed for true long-horizon reasoning; Kairos restores this signal with dense, time-resolved annotations tracking actions, entities, and attributes across 10-to-30-minute videos.
Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most existing video datasets rely on coarse or sparsely aligned supervision, which compresses temporal variation and limits the ability of models to learn reusable representations of continuous visual dynamics. We introduce Kairos, a video dataset for video-language modeling with time-resolved annotations. Kairos consists of long-duration videos, ranging from ten minutes to half an hour, annotated with fine-grained temporal alignment. The annotations capture ongoing actions, entity appearances and attributes, interactions, and evolving contextual cues along the video timeline. This time-resolved structure supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation. Kairos provides a general-purpose foundation for modeling visual experiences over time.