Search papers, labs, and topics across Lattice.
3
0
4
52
Frozen video diffusion models can effectively serve as competitive encoders for a wide range of tasks, merging generation and understanding seamlessly.
Text-to-3D models can get stuck in "sink traps" where they ignore your prompts, but unlocking their unconditional generation power lets you edit shapes they otherwise couldn't.
DINOv2's impressive unimodal performance doesn't translate to cross-modal understanding, but a simple training tweak can align embeddings across RGB, depth, and segmentation without sacrificing feature quality.