Search papers, labs, and topics across Lattice.
Affiliation:
1
0
3
This paper proposes MOCO, a novel diffusion-based framework capable of processing multiple simultaneous inputs, including speech audio, text descriptions, and trajectory data, to generate coherent and lifelike motions without requiring aligned multimodal data.