Search papers, labs, and topics across Lattice.
4
0
7
5
Models trained on LAION-BVD achieve state-of-the-art performance in multimodal tasks, showcasing the dataset's potential to redefine video understanding.
LITTLELEARNER reveals that even a well-defined knowledge scope can yield a competent language model, but it won't expand its capabilities beyond its educational boundaries.
Visual prompt engineering can boost video model reasoning performance beyond traditional text-based methods, revealing a new frontier in model optimization.
A 1000x larger video reasoning dataset reveals early signs of emergent generalization, offering a new foundation for training and evaluating spatiotemporal AI.