Search papers, labs, and topics across Lattice.
Xiamen University, Kling Team, Kuaishou Technology
3
0
5
TimePLE redefines video temporal grounding by predicting valid intervals directly, leading to significant performance gains over traditional endpoint-based methods.
CineCap achieves a new state of the art in cinematographic video captioning by effectively balancing descriptive completeness with factual accuracy through innovative structured reasoning techniques.
Current Omni-modal LLMs can ace perception tasks but still fail at basic social interactions like knowing when and how to jump into a conversation.