Search papers, labs, and topics across Lattice.
University of Science and Technology of China
2
0
4
0
Pruning 75% of visual tokens without sacrificing performance could redefine efficiency benchmarks for VideoLLMs.
Leveraging resolution differences can yield significant performance gains in multimodal large language models without the need for external supervision or annotations.