Search papers, labs, and topics across Lattice.
3
0
5
0
TimeLens2 outperforms models with up to 397B parameters by effectively grounding temporal evidence in videos, redefining expectations for multimodal LLM capabilities.
Current MLLMs fail to provide adequate support for visually impaired individuals, particularly in anticipating navigation-critical events in real-time.
VideoChat3 achieves unprecedented generalization in video understanding while maintaining high efficiency, outperforming larger models with just 4 billion parameters.