Search papers, labs, and topics across Lattice.
2
0
3
Pruning 80.2% of vision tokens without sacrificing accuracy could revolutionize the efficiency of multimodal large language models.
Leveraging resolution differences can yield significant performance gains in multimodal large language models without the need for external supervision or annotations.