Search papers, labs, and topics across Lattice.
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
2
0
4
Pruning 77.8% of visual tokens without losing performance could revolutionize the efficiency of multimodal large language models.
Autonomous driving gets a boost: EvoDriveVLA's collaborative perception-planning distillation framework significantly enhances VLA model performance by tackling perception degradation and planning instability.