Search papers, labs, and topics across Lattice.
This study investigates the efficacy of Self-PreTraining (SPT) for transformer architectures applied to various medical time series tasks, including rehabilitation robotics, stress detection, and Parkinson's disease detection. The authors employ four distinct masking-based objectives to enhance temporal and cross-modal representation learning, systematically varying model depth to assess the interaction between model capacity and pre-training benefits. Results show that SPT improves classification accuracy by 0-6 percentage points across diverse datasets and configurations, particularly benefiting deeper models that leverage enriched representations, highlighting SPT's potential in enhancing performance in data-limited clinical environments.
Self-PreTraining boosts transformer accuracy in medical time series by up to 6 percentage points, even with limited data.
Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical time series. Our objective is to assess the impact of SPT on the performance and scalability of transformer-based models across diverse medical applications, particularly under limited data conditions. We evaluate transformer architectures on three representative medical time-series tasks: rehabilitation robotics (Camargo dataset), stress detection (Non-EEG Stress), and Parkinson's disease detection (Gait Parkinson's Disease). Models are trained either from scratch or through SPT using four masking-based objectives designed to promote temporal and cross-modal representation learning, and we systematically vary model depth to examine how capacity interacts with pre-training benefits. Across datasets and configurations, SPT consistently improves classification accuracy by 0-6 percentage points depending on masking strategy, dataset and architecture, with gains observed not only in multivariate settings but also when models are restricted to simple univariate inputs. The improvements increase for deeper models that can better exploit the enriched temporal representations learned during pre-training. These findings indicate that SPT is a simple and general strategy that enhances transformer performance on medical time-series tasks without requiring task-specific architectural changes, supporting its potential to improve robustness and accuracy in data-limited clinical settings.