Search papers, labs, and topics across Lattice.
This paper introduces a generative diffusion framework for synthesizing realistic drummer motion from audio, addressing the challenges of high-acceleration dynamics and spatial-temporal precision. By employing a dual-objective loss function, the model achieves centimeter-level stick precision while maintaining natural body dynamics, outperforming existing methods reliant on motion matching or MIDI input. The authors also propose novel evaluation metrics, demonstrating through quantitative analysis and user studies that their system produces motion indistinguishable from real performances, thereby enhancing applications in entertainment and education.
Achieving centimeter-level precision in drummer motion synthesis from audio could revolutionize character animation in music-driven applications.
Music-driven character animation enables and enhances transformative applications in entertainment and interactive education. However, synthesizing realistic drumming motion from audio remains challenging due to the inherent tension between high-acceleration dynamics and the need for extreme spatial-temporal precision. Existing approaches, often reliant on motion matching or MIDI input, struggle with generalizing to diverse real-world audio. Moreover, the field lacks standardized evaluation metrics capable of distinguishing precise drumming from noisy motion. In this paper, we introduce a generative diffusion framework featuring a dual-objective loss function that decouples skeletal integrity from drumstick precision, thus enabling centimeter-level stick precision without sacrificing natural body dynamics. Additionally, leveraging our own dataset and data augmentation strategy, the model generalizes to non-curated, in-the-wild audio. To rigorously evaluate performance, we propose two novel metrics: an impact-to-target distance to quantify spatial precision and an audio-motion correlation score to assess temporal alignment. Our quantitative analysis and user studies demonstrate that our system generates high-quality motion that is often indistinguishable from ground-truth performances.