Search papers, labs, and topics across Lattice.
To address the severe scarcity of diverse, in-the-wild biomechanical video data, this work builds SynthGait-19K, a dataset of 19,272 physically grounded synthetic walking videos derived from 6,427 MoCap sequences across 437 subjects with paired SMPL kinematics and clinical gait metrics. Evaluating synthetic-to-real transfer across direct RGB (using a novel GaitXFormer reference model), pose-based, and human mesh recovery (HMR) pipelines shows that synthetic supervision transfers robustly to real-world footage. Crucially, the authors uncover that spatial gait parameters are substantially more vulnerable to sim-to-real domain shift than temporal parameters, and higher surface-level HMR reconstruction accuracy does not reliably yield superior downstream gait parameter estimates.
Better human mesh recovery does not guarantee more accurate gait parameter estimation, exposing a critical disconnect between standard visual 3D pose metrics and downstream biomechanical fidelity.
Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mobility assessment, yet progress is limited by the small scale, restricted viewpoints, and limited visual diversity of existing datasets. We introduce SynthGait-19k, a physically grounded synthetic video dataset containing 19,272 walking videos derived from 6,427 MoCap sequences across 437 subjects, with paired SMPL motion and annotations for six gait parameters. To construct the dataset, we develop Gait2Vid, which unifies heterogeneous MoCap recordings through SMPL and synthesizes diverse RGB walking videos under controllable viewpoints and scene appearances. We assess the generated videos for consistency with their conditioning gait kinematics and validate extracted gait events against force-platform measurements. Using SynthGait-19K, we benchmark direct RGB, pose-based, biomechanical, and human-mesh-recovery approaches and analyze viewpoint, training-data scale, and synthetic-to-real domain shift. We also introduce GaitXFormer as a direct RGB reference model for estimating gait parameters. Synthetic supervision transfers effectively to real videos across both GaitXFormer and a pose-based architecture, demonstrating utility across different representations. We further find that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation.