Search papers, labs, and topics across Lattice.
2
0
4
8
V-RAE achieves a remarkable 2.13 rFVD on K600, outperforming traditional video VAEs by retaining significantly more semantic information in its latent representations.
Forget unimodal tasks鈥擴niM throws down the gauntlet for truly unified multimodal AI, demanding models juggle any combination of text, image, audio, video, code, documents, and 3D inputs and outputs in a single, interleaved stream.