Search papers, labs, and topics across Lattice.
Nankai University
3
0
6
Ultra-low token pruning can achieve 92.1% of peak performance with only 16 visual tokens, thanks to a novel approach that preserves distribution consistency.
The field of video understanding is rapidly shifting from isolated pipelines to unified models capable of adapting to diverse downstream tasks, demanding a re-evaluation of current approaches.
MLLMs can "hear" a little, but EgoSound reveals they're still largely deaf to the nuances of sound in egocentric video, especially when it comes to spatial and causal reasoning.