Search papers, labs, and topics across Lattice.
The paper introduces AnyMo, a unified multimodal framework for conditional human motion generation that leverages a new large-scale dataset, OmniHuMo, containing over 5,000 hours of motion data with aligned text, speech, music, and trajectory annotations. AnyMo combines a Residual FSQ-based motion tokenizer with a masked modeling transformer to enable high-quality motion synthesis under arbitrary modality combinations. Experiments demonstrate AnyMo's ability to generate high-fidelity motion with flexible control over spatial and stylistic attributes, addressing limitations of existing methods constrained by fixed modality configurations.
Unlock motion generation from any combination of text, speech, music, and trajectory inputs with AnyMo, a unified model trained on a massive new multimodal dataset.
Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific architectures, leaving cross-modal interactions and the scaling laws of multimodal-conditioned synthesis largely underexplored. A key bottleneck is the scarcity of large-scale modality-aligned motion data, limiting generalization across diverse control signals. In this work, we introduce OmniHuMo, a large-scale, high-quality dataset comprising over 5,000 hours of motion and 3.2 million sequences with precisely aligned multimodal annotations (e.g., text, speech, music, and trajectory). Leveraging OmniHuMo, we propose AnyMo, a unified multimodal framework combining a Residual FSQ-based motion tokenizer with a scalable masked modeling transformer, enabling high-quality motion synthesis under arbitrary modality combinations. Extensive experiments show that AnyMo achieves high-fidelity synthesis while offering flexible control over both spatial and stylistic attributes.