Search papers, labs, and topics across Lattice.
University of Science and Technology of China
2
0
2
Training with local visual cues can dramatically enhance MLLMs' ability to extract fine-grained visual details without altering their inference interface.
Current video MLLMs struggle to grasp fleeting visual events, with top models barely surpassing 39% accuracy on critical momentary tasks.