Search papers, labs, and topics across Lattice.
The Chinese University of Hong Kong, Shenzhen
1
0
2
6
Ex-Omni-2D generates visually coherent dialogue responses that seamlessly integrate text, speech, and video, all while avoiding the need for extensive multi-modal training data.