Search papers, labs, and topics across Lattice.
This study investigates the functional equivalence and geometric diversity of neural network approximations, particularly focusing on single-layer networks and multilayer perceptrons under varying noise conditions. By employing the concept of sloppiness and analyzing the eigen spectrum of the Hessian, the authors reveal that numerous functionally indistinguishable networks exist, characterized by low effective rank and significant structural redundancy. The findings culminate in a model selection criterion aimed at optimizing parsimony and inference efficiency, addressing practical identifiability concerns in neural network representations.
A vast number of functionally equivalent neural networks can exhibit striking geometric diversity, challenging assumptions about model uniqueness in approximation tasks.
The Universal Approximation Theorem states that a neural network with a single hidden layer is sufficient to approximate any continuous univariate function on a compact domain to arbitrary error. However, the uniqueness of such neural network representations is not guaranteed, raising questions about practical identifiability. In this work, we address this concern by analyzing functional equivalence and geometric diversity of neural network approximations to a few elementary mathematical functions. The analysis includes an extensive study of single-layer neural networks and multilayer perceptrons under noisy and noise-free conditions. Beyond just network capacity, we study the geometric properties through the lens of sloppiness, characterized by the eigen spectrum of the Hessian of the cost function and the effective rank to quantify the dimensionality of parameter space. The study reveals large equivalence classes of functionally indistinguishable yet geometrically diverse networks that consistently exhibit low effective rank and structural redundancy. Finally, a model select criterion is proposed for identifying optimal models based on parsimony, ease of estimation, and inference efficiency.