Mar 19, 2026arXiv:2603.19227

Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer

Chenyang Gu, Chenyang Gu, Mingyuan Zhang, Mingyuan Zhang, Haozhe Xie, Haozhe Xie, Zhongang Cai, Zhongang Cai, Lei Yang, Lei Yang, Ziwei Liu, Ziwei Liu

AI Summary

This paper introduces a three-stage motion generation framework that combines semantic conditioning from discrete token-based generators with the kinematic control of continuous diffusion models. The core innovation is MoTok, a diffusion-based discrete motion tokenizer that uses a diffusion decoder for motion recovery, enabling compact single-layer tokens while preserving motion fidelity. Experiments on HumanML3D demonstrate that this approach significantly improves controllability and fidelity compared to existing methods, especially under strong kinematic constraints, while using fewer tokens.

Key Contribution

Achieve 9x lower trajectory error and 3x better FID in motion generation by using a diffusion-based discrete motion tokenizer that elegantly handles both semantic and kinematic constraints.

Abstract

Prior motion generation largely follows two paradigms: continuous diffusion models that excel at kinematic control, and discrete token-based generators that are effective for semantic conditioning. To combine their strengths, we propose a three-stage framework comprising condition feature extraction (Perception), discrete token generation (Planning), and diffusion-based motion synthesis (Control). Central to this framework is MoTok, a diffusion-based discrete motion tokenizer that decouples semantic abstraction from fine-grained reconstruction by delegating motion recovery to a diffusion decoder, enabling compact single-layer tokens while preserving motion fidelity. For kinematic conditions, coarse constraints guide token generation during planning, while fine-grained constraints are enforced during control through diffusion-based optimization. This design prevents kinematic details from disrupting semantic token planning. On HumanML3D, our method significantly improves controllability and fidelity over MaskControl while using only one-sixth of the tokens, reducing trajectory error from 0.72 cm to 0.08 cm and FID from 0.083 to 0.029. Unlike prior methods that degrade under stronger kinematic constraints, ours improves fidelity, reducing FID from 0.033 to 0.014.

Computer Vision Robotics & Embodied AI World Models & Planning

Citation Metrics

Citations0

Influential citations0

References50

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer

Related Papers