Search papers, labs, and topics across Lattice.
This study addresses the limitations of existing semantic decoders in non-invasive brain recordings by introducing a multi-feature fusion framework that integrates both static lexical representations and dynamic contextual representations. By benchmarking linear Naive Concatenation against non-linear Multi-Head Cross-Attention, the authors demonstrate that the latter significantly enhances semantic reconstruction performance. The results indicate that simulating the collaborative modulation between contextual information and core lexical attributes leads to state-of-the-art decoding capabilities, paving the way for improved brain-to-text decoding methods.
Non-linear cross-attention fusion outperforms traditional methods, revealing that effective semantic reconstruction hinges on integrating static and dynamic features of language comprehension.
Continuous semantic reconstruction from non-invasive neural recordings remains limited by the representational mismatch between semantic feature spaces and neural coding patterns, which severely impedes cross-modal alignment between high-noise neural signals and target semantic features. Prior semantic decoders have predominantly relied on static lexical representations or dynamic contextualized representations in isolation. This single-dimension approach inevitably leads to severe information loss, as it fails to account for the human brain's capacity to integrate stable word attributes and dynamic contexts simultaneously.To bridge this gap, this study introduces a multi-feature fusion framework for non-invasive semantic reconstruction, systematically benchmarking two integration approaches: linear Naive Concatenation and non-linear Multi-Head Cross-Attention. Within this framework, our approach complements static lexical representations (W2V) with dynamic contextual representations (GPT) via an interactive gating mechanism to facilitate cooperative processing during language comprehension.Evaluated through extensive semantic reconstruction and text generation experiments, our framework reveals a robust performance hierarchy: Cross-Att > Concat > GPT > W2V. Crucially, the non-linear cross-attention fusion method achieves state-of-the-art performance, demonstrating that neural language decoding benefits from simulating the collaborative modulation between contextual information and core lexical attributes rather than depending on isolated individual features, while also offering a viable non-invasive brain-to-text decoding method.