Search papers, labs, and topics across Lattice.
This paper introduces ReMiX-MAE, a self-supervised multimodal masked pretraining framework designed to learn robust facial representations from RGB-only clinical videos, addressing the limitations of scarce labeled data and the impracticality of deploying thermal and depth signals. By utilizing the newly collected Sympathetic Mediated Pain (SMP) dataset, the authors demonstrate that ReMiX-MAE significantly outperforms traditional RGB-only approaches, especially in challenging multi-class pain assessment scenarios. The method not only enhances performance in clinical settings but also showcases improved transferability across external datasets, underscoring its potential for practical applications in automated pain assessment.
ReMiX-MAE achieves superior performance in pain assessment using only RGB data, revealing the untapped potential of self-supervised learning in data-scarce clinical environments.
Automated pain assessment in real clinics is limited by scarce clinically grounded facial video data with weak labels (often sequence-level self-report) and by the fact that pain cues can be subtle or near-neutral in RGB, while thermal and depth signals are informative yet impractical to deploy routinely. To address these challenges, we propose ReMiX-MAE (Reconstructing Missing Channel Cross-Modal Masked Autoencoder), a self-supervised multimodal masked pretraining framework that learns transferable facial representations from synchronized RGB, thermal, and depth videos and explicitly trains robustness to missing modalities, enabling RGB-only deployment. To fill the gap of clinically grounded facial pain data with video-level self-report and longitudinal treatment trajectories, we collect the Sympathetic Mediated Pain (SMP) dataset with paired pre- and post-recordings across multiple visits. Under RGB-only deployment, we evaluate ReMiX-MAE using both direct feature extraction and pseudo-multimodal features decoded from RGB. ReMiX-MAE consistently outperforms an RGB-only masked autoencoder baseline on SMP, with pseudo-multimodal features providing additional gains in the challenging five-class setting. Across external datasets, ReMiX-MAE further shows more robust and label-efficient transfer than RGB-only baselines, highlighting its advantage in data-limited clinical settings.