Search papers, labs, and topics across Lattice.
Wuhan University
3
0
6
16
VAORA reduces hallucinated reasoning in VLMs by aligning visual context with action outcomes, leading to better generalization in unseen tasks.
PixelEyes achieves precise visual localization by separating reasoning from perception, drastically reducing the redundancy in multi-turn visual searches.
Cosmos 3 sets a new benchmark for omnimodal models, outperforming existing state-of-the-art in both Text-to-Image and Image-to-Video tasks.