Search papers, labs, and topics across Lattice.
This paper introduces a finite-width geometric framework that elucidates the organization and selective alignment of learned feature geometries in deep neural networks through the use of commutators. It quantifies incompatibilities among weight-generated covariance, gates, and backward sensitivities, revealing that spectral alignment is influenced by transport dynamics and interaction effects rather than being a straightforward outcome of training. Key findings include the identification of four sources contributing to sensitivity-covariance commutator dynamics and the demonstration that cancellation effects persist across various network architectures and tasks, challenging conventional notions of alignment in neural networks.
Spectral alignment in deep neural networks is not a universal outcome of training but a complex interplay of transport dynamics and cancellation effects that varies by layer and scale.
We develop a finite-width geometric framework describing how learned feature geometries are organized, transported, and selectively aligned in deep neural networks. Incompatibility among weight-generated covariance, gates, and backward sensitivities is quantified through three families of commutators: between gates and covariance, between sensitivities and covariance, and between average gradient outer products (AGOPs) and neural feature matrices (NFMs). An exact layerwise identity decomposes the sensitivity-covariance commutator into four sources: downstream transport, adjacent-layer imbalance, pointwise sensitivity fluctuations, and nonlinear gate-covariance interactions. The AGOP-NFM commutator is a singular-value-weighted transport of the internal commutator, explaining why observed feature-side alignment alone does not determine the internal geometry from which it emerges. Buffered localized energies resolve mixing between separated covariance subspaces. We establish spectral-gap, projector-evolution, and stabilization estimates, and formulate conditional Lyapunov principles that yield decay under explicit geometric error-bound or intrinsic-damping assumptions. These criteria do not follow from gradient flow alone and clarify why risk reduction need not imply commutator collapse. Analytic examples and numerical experiments exhibit factorization of spectral and activation geometry, transient growth, and cancellation among nonzero sources. In tested finite-time regimes, cancellation dominated by a negative transport-imbalance interaction persists across depths, widths, and two regression benchmarks. Spectral alignment therefore appears as a layer- and scale-dependent compatibility phenomenon governed by transport, interaction, cancellation, and possible damping, rather than a universal consequence of training.