Search papers, labs, and topics across Lattice.
3
1
4
6
Current video generation models struggle with visual reasoning, with the best achieving only 51% accuracy on a new benchmark designed to probe their capabilities.
RTSKG reveals that integrating urban entities into a cohesive knowledge graph can dramatically improve the accuracy of ridership predictions and related urban analyses.
Current video understanding benchmarks and post-training datasets are riddled with linguistic biases, meaning VLMs might be acing tests without actually "watching" the video.