Search papers, labs, and topics across Lattice.
Wuhan University
2
0
4
4
PixelEyes achieves precise visual localization by separating reasoning from perception, drastically reducing the redundancy in multi-turn visual searches.
Achieving a 43.65% Effective Temporal F1 score, this work reveals that MLLMs can be effectively adapted for complex One-to-Many Temporal Grounding tasks, challenging the limitations of previous models.