Search papers, labs, and topics across Lattice.
This paper introduces a new MoCap dataset designed to address the limitations of existing benchmarks for markerless 4D human motion capture, which often fail to capture the complexities of real-world human interactions. The dataset includes synchronized multi-view RGB and depth sequences, ground-truth 3D motion capture from a Vicon system, and SMPL/SMPL-X parameters, covering single and multi-person scenarios with intricate motions and occlusions. Benchmarking state-of-the-art markerless MoCap models on this dataset reveals significant performance degradation, demonstrating the need for more robust approaches and the dataset's value for model development.
Current markerless motion capture models crumble when faced with the messy reality of human interaction: occlusions, rapid movements, and similar clothing styles.
Marker-based motion capture (MoCap) systems have long been the gold standard for accurate 4D human modeling, yet their reliance on specialized hardware and markers limits scalability and real-world deployment. Advancing reliable markerless 4D human motion capture requires datasets that reflect the complexity of real-world human interactions. Yet, existing benchmarks often lack realistic multi-person dynamics, severe occlusions, and challenging interaction patterns, leading to a persistent domain gap. In this work, we present a new dataset and evaluation for complex 4D markerless human motion capture. Our proposed MoCap dataset captures both single and multi-person scenarios with intricate motions, frequent inter-person occlusions, rapid position exchanges between similarly dressed subjects, and varying subject distances. It includes synchronized multi-view RGB and depth sequences, accurate camera calibration, ground-truth 3D motion capture from a Vicon system, and corresponding SMPL/SMPL-X parameters. This setup ensures precise alignment between visual observations and motion ground truth. Benchmarking state-of-the-art markerless MoCap models reveals substantial performance degradation under these realistic conditions, highlighting limitations of current approaches. We further demonstrate that targeted fine-tuning improves generalization, validating the dataset's realism and value for model development. Our evaluation exposes critical gaps in existing models and provides a rigorous foundation for advancing robust markerless 4D human motion capture.