Search papers, labs, and topics across Lattice.
2
0
5
2
Current video generation models struggle with visual reasoning, achieving only 51% accuracy on a new benchmark designed to probe their capabilities.
Today's best AI agents can only complete 33% of common online tasks like booking appointments or filling out job applications, revealing a significant gap between current capabilities and real-world utility.