Search papers, labs, and topics across Lattice.
3
0
4
8
V-RAE achieves a remarkable 2.13 rFVD on K600, outperforming traditional video VAEs by retaining significantly more semantic information in its latent representations.
Predicting human decisions requires understanding not just the physical world, but also the mental states that drive behavior鈥擬WM makes this explicit.
Achieve SOTA joint audio-video generation with JavisDiT++ using just 1M public training examples, rivaling performance of models trained on proprietary datasets.