Search papers, labs, and topics across Lattice.
Affiliation:
2
0
5
17
GATO-Vid achieves superior spatial localization in text-to-video generation without the computational costs of traditional gradient-based methods.
LVLMs are often tripped up not by faulty vision, but by over-trusting the textual prompt, leading to surprisingly easy-to-fix hallucinations.