Search papers, labs, and topics across Lattice.
3
0
5
3
Decoupling perception from reasoning in visual tasks leads to a remarkable 93.2% accuracy on V-Star, showcasing a new paradigm for fine-grained visual reasoning.
Multimodal LLMs still struggle to faithfully recreate webpages from videos, particularly in capturing fine-grained style and motion, despite advances in other areas.
Forget fuzzy language – CoCo uses executable code as Chain-of-Thought to generate images with unprecedented control and precision, blowing away existing methods on complex scenes.