Search papers, labs, and topics across Lattice.
This paper reinterprets self-attention in transformers through the lens of connection walks on token-position graphs, establishing a unified operator framework for understanding attention mechanisms. The authors prove that single-head attention functions as a connection propagation step with constant transport, while multi-head attention operates as an edge-dependent connection walk with attention-gated mixtures of transports. Empirical results across various transformer scales and structures reveal that effective attention graphs stabilize into geometric operators in deeper layers, supporting the theoretical insights and providing tools for further analysis of transformer models.
Single-head attention is a precise connection propagation step, while multi-head attention reveals complex edge-dependent dynamics that reshape our understanding of transformer geometry.
Self-attention is a ubiquitous primitive in modern sequence models, yet its operator-level geometry is only partially understood. We view a token sequence as a vector field over the token-position graph and identify attention as a connection walk: messages are aggregated by a nonnegative walk matrix while being transported along each edge by a learned linear map. Within this framework, we prove that single-head attention (SHA) is exactly a connection propagation step with constant transport, and that multi-head attention (MHA) is exactly a single edge-dependent connection walk whose effective transport is an attention-gated mixture of headwise transports. We further clarify the conditions under which the corresponding generator reduces to a random-walk connection Laplacian, highlighting the roles of stochasticity, reversibility, and metric-compatible transports. Empirically, we find that trained Transformers across scales (from 124M to 8B) and structures (encoder/decoder) exhibit geometric structure consistent with our theory: effective attention graphs converge to stable geometric operators in deeper layers, learned transports self-organize into approximate scaled isometries, and both phenomena strengthen consistently with scale. Overall, the paper provides a precise connection-walk formalism that links self-attention to classical geometric operators, along with a set of operator-level tools for analyzing transformer models from a geometric perspective.