Search papers, labs, and topics across Lattice.
This paper introduces SeMoCo, a novel semantic-first motion codec designed to enhance motion language modeling by separating semantic and kinematic information within motion tokens. By employing a dual-axis motion generator that autoregressively refines kinematic details while modeling semantic progression, SeMoCo significantly improves the reconstruction accuracy of motion representations. The authors also present the $惟$-MotionVerse dataset, which supports the evaluation of their approach and demonstrates superior performance in text-to-motion generation compared to existing codecs.
SeMoCo achieves unprecedented reconstruction accuracy by decoupling semantic meaning from kinematic details in motion representations.
Discrete motion representations have substantially advanced autoregressive text-to-motion generation. However, most motion tokenizers are optimized for reconstruction and do not explicitly allocate capacity according to semantic role. Action-level meaning and fine-grained kinematic detail must therefore be encoded through the same reconstruction-driven hierarchy. We introduce SeMoCo, a semantic-first motion codec, together with a dual-axis motion generator for language-conditioned motion generation. Each motion token contains one semantic token and a residual sequence of kinematic tokens. The generator models semantic progression across time and autoregressively refines the residual entries. We also construct $惟$-MotionVerse, a large-scale, multi-source human-motion dataset unified under the SOMA representation. Across the reported comparisons, SeMoCo achieves the best reconstruction accuracy among the compared codecs, while strong text-to-motion results demonstrate the effectiveness of its motion tokens for downstream generation.