Search papers, labs, and topics across Lattice.
3
9
6
5
VideoChat3 achieves unprecedented generalization in video understanding while maintaining high efficiency, outperforming larger models with just 4 billion parameters.
Achieve real-time, proactive video understanding with StreamOV, which uses bounded memory and a novel response trigger to overcome the limitations of offline methods.
Current LLMs and VLMs struggle with multi-step reasoning in long videos, often failing to maintain temporal coherence and procedural validity, as revealed by a new benchmark of hour-long narratives.