Search papers, labs, and topics across Lattice.
This study investigates the landscapes of complex-parameterized neural networks from an information-theoretic manifold perspective, utilizing classical optimization guarantees linked to Dolbeault asymptotics. By employing a K盲hler information metric and natural gradient descent, the authors ensure that descent paths remain within the holomorphic tangent bundle, while also addressing the challenges posed by Calabi-Yau manifolds and their ill-conditioned landscapes. Key findings reveal that negative curvature significantly disrupts the loss landscape, highlighting the intricate relationship between geometric properties and neural network optimization guarantees.
Negative curvature can destabilize neural network optimization, revealing critical insights into the geometry of loss landscapes.
We study landscapes for complex-parameterized networks. Our approach is motivated with an information-theoretic manifold perspective of the parameter and via classical optimization guarantees although of complex geometric variety such as through Dolbeault asymptotics. The descent path admits a K\"ahler information metric under a cross-entropy via the Wirtinger Hessian on the log-likelihood potential. We restrict attention to a descent update rule with natural gradient descent via a differentiated loss scaled by the inverse metric, so the descent path remains in the holomorphic tangent bundle. We emphasize Calabi-Yau information manifolds which profane theoretical guarantees via an ill-curvature-conditioned landscape. Under a Calabi-Yau metric, specifically in a non-compact setting with a global potential so defined geometrically rather than invoking the topological requirements of the Calabi conjecture, a wedged nowhere-vanishing holomorphic form is the top exterior product of the K\"ahler form up to constants, yielding a constant determinant condition. Under a fixed determinant, a metric almost low rank up to an eigenvalue tolerance implies a blow-up effect. Moreover, it has been discovered that negative curvature subverts the loss landscape, specifically sectional curvature, so we expand on this and draw interconnections to negative-definite Ricci curvature. Our arguments primarily exist in a geometric analytic modality, although we establish roots in deep learning theory such as through asymptotics at initialization and connections through failure modes of neural network guarantees under vanishing and negative Ricci curvature.