Search papers, labs, and topics across Lattice.
This paper investigates the feature learning dynamics of multilayer perceptrons (MLPs) in regression tasks with clustered data, revealing that MLPs develop specialized neurons that align with specific predictive features in localized input regions. This challenges the prevailing view of a singular global low-dimensional representation, instead highlighting the emergence of multiple local representations that together cover a high-dimensional space. The key finding demonstrates that this specialization enhances data efficiency in MLPs compared to traditional feature-learning methods reliant on global representations.
MLPs leverage specialized neurons to achieve superior data efficiency by creating localized representations, contradicting the notion of a single global feature space.
Understanding how neural networks learn and organize features is central to understanding their behavior. Much existing theory of feature learning has focused on the emergence of a global low-dimensional predictive geometry. We show that this picture is incomplete. In regression problems with clustered data, we demonstrate that multilayer perceptrons (MLPs) naturally develop monosemantic specialized neurons: individual neurons become strongly aligned with a specific predictive feature relevant to a particular region of the input space. Rather than learning a single global low-dimensional representation, MLPs learn a collection of local low-dimensional representations that can collectively span a high-dimensional space. This specialization provably gives MLPs a data-efficiency advantage over feature-learning methods based on a global low-dimensional representation.