Search papers, labs, and topics across Lattice.
The paper introduces STEAM, a hierarchical transfer framework designed to enhance EEG decoding by integrating general-purpose representation learning with paradigm-specific specialization. This model employs a dual-branch spatio-temporal encoder and a shared soft mixture-of-experts (SSMoE) module to facilitate the exchange of complementary representations, significantly improving generalizability and adaptation efficiency. Experimental results across seven datasets demonstrate that STEAM achieves superior performance while maintaining competitive inference costs, underscoring its effectiveness in real-world applications of brain-computer interfaces.
STEAM achieves superior EEG decoding performance with a novel hierarchical pre-training approach that enhances model specialization without starting from scratch.
Brain-computer interfaces (BCIs) have been widely used in motor rehabilitation, disease diagnosis, and other neural engineering scenarios. However, conventional neural signal decoding algorithms often suffer from limited generalizability and high adaptation costs, motivating recent interest in BCI foundation models. Existing approaches still struggle to jointly achieve general transferability, accurate decoding, and efficient downstream adaptation. We present STEAM, a hierarchical transfer framework that reconciles general-purpose representation learning with paradigm-specific specialization in EEG foundation models. The framework is instantiated as a dual-branch spatio-temporal encoder in which a shared soft mixture-of-experts (SSMoE) module aligns the spatial and temporal branches, allowing complementary representations to exchange information through a compact set of soft slots. Across seven downstream datasets and fourteen evaluation settings, STEAM attains the best average rank among the compared methods at a competitive inference cost measured in FLOPs. Building upon the Stage-I general initialization, the hierarchical pre-training strategy further specializes the model to a target paradigm without retraining from scratch, yielding consistent gains in paradigm-specific decoding accuracy.