Search papers, labs, and topics across Lattice.
2
0
3
11
Video generation models could be the key to unlocking general-purpose vision intelligence, outperforming specialized models with far less training data.
DINOv2's impressive unimodal performance doesn't translate to cross-modal understanding, but a simple training tweak can align embeddings across RGB, depth, and segmentation without sacrificing feature quality.