Search papers, labs, and topics across Lattice.
This paper establishes the first algorithmic separation between constant-depth and logarithmic-depth neural networks by identifying a specific class of Boolean functions that logarithmic-depth networks can learn efficiently through a hierarchical layerwise coordinate descent approach. The authors demonstrate that these networks can reconstruct hierarchically structured Fourier spectra adaptively, while constant-depth networks face inherent limitations, incurring a constant \( L^2 \) approximation error for certain functions. This distinction is significant as it highlights the computational advantages of deeper architectures in learning complex functions, challenging the traditional view of depth in neural network design.
Logarithmic-depth networks can efficiently learn complex Boolean functions that constant-depth networks cannot, revealing a critical algorithmic separation in neural network capabilities.
Despite the empirical advantages of deep networks over shallow ones, theoretical depth separations largely concern approximation power, while algorithmic results are mostly limited to comparisons between two- and three-layer networks. In this work, we prove the first algorithmic separation between constant-depth and logarithmic-depth networks. Specifically, we identify a class of Boolean functions with hierarchically structured Fourier spectra that logarithmic-depth networks can learn efficiently using layerwise coordinate descent by reconstructing the spectra hierarchically and adaptively. We also exhibit a subclass for which every constant-depth, polynomial-width network with sufficiently regular activations and controlled spectral norms must incur constant $L^2$ approximation error under the uniform distribution over the hypercube.