Search papers, labs, and topics across Lattice.
3
0
5
GMM-EVA slashes visual token budgets in long video understanding by intelligently prioritizing keyframes based on event-level structures.
Visual reranking and active rejection in MMAgent-R$^2$ significantly boost retrieval accuracy in challenging KB-VQA tasks, outperforming traditional methods.
MLLMs can be tricked into missing 90% of harmful content simply by encoding it in images that humans can easily read.