Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
0
A unified framework that effectively fuses RGB and event data can drastically improve person re-identification accuracy across different camera views.
Augmenting images from a model's own failures can lead to substantial performance boosts in multimodal tasks, outperforming conventional augmentation techniques.
ERA's innovative approach to visual token pruning preserves attention integrity, enabling efficient MLLMs without sacrificing performance.