Search papers, labs, and topics across Lattice.
7
0
6
10
TimeLens2 outperforms models with up to 397B parameters by effectively grounding temporal evidence in videos, redefining expectations for multimodal LLM capabilities.
VideoChat3 achieves unprecedented generalization in video understanding while maintaining high efficiency, outperforming larger models with just 4 billion parameters.
Current MLLMs fail to provide adequate support for visually impaired individuals, particularly in anticipating navigation-critical events in real-time.
MIMFlow achieves a 32.8% performance boost over standard Normalizing Flows by cleverly decoupling low-frequency and high-frequency image generation tasks.
UniDDT achieves a groundbreaking balance between multimodal understanding and generation, outperforming existing models in both tasks with enhanced semantic coherence.
Current multimodal LLMs choke on long-form video understanding, either forgetting details or getting lost in the timeline, but a new agentic architecture with dynamic memory management offers a promising fix.
Current video LLMs falter when faced with the demands of real-time interaction, a gap RIVER Bench directly addresses by providing a challenging new evaluation framework.