Search papers, labs, and topics across Lattice.
To eliminate the need for wearable sensors in Freezing of Gait (FoG) detection without sacrificing precision, the authors train a vision-only network using cross-modal subspace distillation guided by privileged IMU kinematics and clinical metadata. This approach directly tackles the failure of purely visual systems during turning-in-place tasks, where geometric self-occlusion typically obscures the high-frequency kinematic precursors of FoG. By incorporating velocity and acceleration derivatives alongside confidence-based gating, the resulting inference-time vision-only model matches hardware-sensor predictive fidelity, reaching 85.5% accuracy and 82.4% balanced accuracy.
Camera-only gait tracking can match wearable IMU precision even through severe self-occlusion when supervised via cross-modal subspace distillation from privileged sensor oracles.
Objective assessment of Freezing of Gait (FoG) in Parkinson's disease (PD) relies predominantly on wearable Inertial Measurement Units (IMUs). While IMUs provide optimal kinematic precision, mandatory sensor attachment restricts continuous clinical deployment. Conversely, unobtrusive vision-based alternatives suffer substantial classification errors during turning-in-place tasks, where geometric self-occlusion degrades deterministic skeletal coordinates and obscures the high-frequency precursors required for FoG detection. To resolve these physical observation limits, we propose a supervised cross-modal subspace distillation framework. During optimisation, pre-trained kinematic data from IMU sensors and contextual clinical metadata act as oracles to guide a deployable visual architecture. By incorporating joint velocity and acceleration derivatives, utilising a confidence-based gating mechanism, the visual model mitigates some of the tracking errors during occlusion events. Empirical evaluations confirm this latent alignment transfers the predictive fidelity of hardware sensors directly into the visual representation, yielding $85.5\%$ accuracy, and $82.4\%$ balanced accuracy. All the while maintaining a vision only model at inference.