Search papers, labs, and topics across Lattice.
2
0
4
2
VideoMM is introduced, which marks a paradigm shift from model-centric downsizing to adaptive perceptual granularity and accelerates inference by 2.73 times over current leading methods, establishing a highly scalable paradigm for long-video understanding.
Chemical reaction diagram parsing, a notoriously difficult task for vision-language models, sees a significant leap in performance thanks to a new multi-agent framework that enforces chemical consistency.