Search papers, labs, and topics across Lattice.
4
910
7
28
VideoChat3 achieves unprecedented generalization in video understanding while maintaining high efficiency, outperforming larger models with just 4 billion parameters.
Efficient context handling in video tasks can elevate multimodal models to new heights of agency and reasoning capability.
Current LLMs and VLMs struggle with multi-step reasoning in long videos, often failing to maintain temporal coherence and procedural validity, as revealed by a new benchmark of hour-long narratives.
Open-source multimodal models just leveled up: InternVL3 rivals closed-source titans like GPT-4o by pre-training vision and language together from the start.