Search papers, labs, and topics across Lattice.
This study addresses the challenge of reconstructing speech from EEG signals, which suffer from low signal-to-noise ratios and variability across sessions. By employing a contrastive learning approach that leverages repeated EEG responses as positive pairs and integrating variational regularization, the researchers enhance the robustness of the encoder representation space. The results demonstrate a significant reduction in character error rate (CER) while preserving the fidelity of mel-spectrogram reconstruction, indicating effective session-invariance in the learned representations.
Leveraging repeated EEG responses across sessions, this method achieves remarkable session-invariance and reduces character error rates in EEG-to-speech decoding.
Reconstructing heard speech from non-invasive electroencephalography (EEG) is challenging due to a low signal-to-noise ratio (SNR) and inter-session variability. While trial averaging improves the SNR, it is difficult to apply to continuous speech. We instead use repeated EEG responses to the same stimulus across different sessions as positive pairs for contrastive learning, and introduce variational regularization that, combined with this contrastive objective, keeps the encoder representation space broad. Experiments on a Japanese EEG dataset show that combining the session-invariant strategy with variational regularization improves the character error rate (CER) while maintaining mel-spectrogram reconstruction fidelity. Session probing confirms that the encoder representations achieve session-invariance.