Search papers, labs, and topics across Lattice.
Affiliation:
4
0
6
31
Models trained on VBVR-Pro not only excel in native visual reasoning tasks but also reveal critical insights into the effectiveness of different generative modalities.
Current video generation models struggle with law-grounded reasoning, with the best achieving only 47% on the new Apple-PI benchmark.
A single unified model can outperform specialized systems across various computer vision tasks, all without the need for custom architectures.
Achieve 9x lower trajectory error and 3x better FID in motion generation by using a diffusion-based discrete motion tokenizer that elegantly handles both semantic and kinematic constraints.