Search papers, labs, and topics across Lattice.
The paper introduces VSMP-IMU, a novel framework for generating synthetic inertial measurement unit (IMU) data that leverages video-grounded Semantic Motion Programs (SMP) to enhance human activity recognition (HAR) in challenging settings. By separating activity semantics from variations in motion, VSMP-IMU synthesizes realistic IMU signals from input videos, significantly improving performance in low-resource and imbalanced datasets. Experimental results demonstrate that VSMP-IMU achieves a Macro-F1 score of 78.33%, outperforming both real-only training and previous synthetic data generation methods, particularly in tail-class recognition scenarios.
Structured video-grounded semantics can boost synthetic IMU generation, leading to a 19.86% improvement in tail-class recognition over traditional methods.
Wearable human activity recognition (HAR) is often limited by the scarcity of labeled sensor data, especially in low-resource, class-imbalanced, and subject-generalization settings. Synthetic IMU generation can reduce this dependency and enhance HAR machine learning model's performance, but existing approaches face a trade-off without addressing all factors: video-driven methods are visually grounded but sensitive to pose-estimation errors, while text-driven methods are controllable but often weakly grounded in how activities are actually performed. We present VSMP-IMU, a video-grounded framework for controllable synthetic IMU generation based on a structured Semantic Motion Program (SMP), which separates activity-defining semantics from label-preserving variation. Given an input video, VSMP-IMU extracts and augments an SMP, uses it to synthesize motion, converts the motion into virtual IMU signals, and grounds the resulting signals to the target wearable domain. We evaluate VSMP-IMU against state-of-the-art synthetic data generation methods on five public IMU-HAR datasets under leave-one-person-out evaluation. VSMP-IMU achieves an average Macro-F1 of 78.33%, improving over real-only training by 9.77% and over the strongest prior synthetic baseline by 4.04%. In low-resource settings with reduced training data-samples, it improves over real-only training by 18.54% and over the strongest prior synthetic baselines by more than 6% on average. Under long-tail evaluation in imbalanced datasets, it improves tail-class Macro-F1 by 19.86% over Real-only training and by 4.76% over SOTA. These results show that structured video-grounded semantics provide a practical foundation for controllable, wearable-relevant synthetic sensor data generation.