Search papers, labs, and topics across Lattice.
3
0
4
0
Achieving up to 11脳 reduction in storage for dynamic scene streaming without sacrificing quality or speed could revolutionize online video applications.
Pruning 80.2% of vision tokens without sacrificing accuracy could revolutionize the efficiency of multimodal large language models.
Leveraging resolution differences can yield significant performance gains in multimodal large language models without the need for external supervision or annotations.