Search papers, labs, and topics across Lattice.
Xiamen University
2
0
3
AdaQ enables MLLMs to achieve superior long video understanding with just 64 frames, outperforming state-of-the-art methods by a striking margin.
Forget bolting vision onto language models – truly powerful multimodal AI demands rethinking architectures from the ground up.