Search papers, labs, and topics across Lattice.
5
0
7
5
TimeLens2 outperforms models with up to 397B parameters by effectively grounding temporal evidence in videos, redefining expectations for multimodal LLM capabilities.
SER's innovative approach to grounding video reasoning demonstrates a 3.0-point leap in accuracy by integrating semantic verification into the reward structure.
Efficient context handling in video tasks can elevate multimodal models to new heights of agency and reasoning capability.
Future-L1 shows that preserving visual semantics in latent space can dramatically enhance video event prediction accuracy, outperforming previous models by substantial margins.
Current video LLMs falter when faced with the demands of real-time interaction, a gap RIVER Bench directly addresses by providing a challenging new evaluation framework.