Search papers, labs, and topics across Lattice.
This paper introduces a continuous geometric framework that models the operations of Transformer architectures as integro-differential equations on a semantic fiber bundle, translating core components into the language of differential geometry and stochastic calculus. The authors validate their framework through extensive experiments across five architectures, revealing quantitative consistency with geometric predictions related to stability and optimization dynamics. Key findings include precise Lipschitz scaling and insights into thermodynamic behavior, highlighting the framework's utility in understanding the limits and dynamics of Large Language Models.
Analyzing Transformers through continuous stochastic differential geometry reveals surprising predictive insights into their stability limits and optimization dynamics.
We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle $\calE = \calM \times \R^d$. Beginning from a single geometric axiom -- that the token sequence forms a discrete $1$-manifold equipped with a canonical measure lattice -- we translate every core component of the modern Transformer (RMSNorm, RoPE, Softmax Attention, FFN, Residual Stream, SGD, Weight Decay) into a cohesive vocabulary of differential geometry, measure theory, and stochastic calculus. The resulting framework yields quantitative predictions spanning entropic optimal transport (Attention as a Schr\"odinger bridge) and non-equilibrium thermodynamics (SGD as It\^{o} diffusion violating detailed balance). We conduct a six-part experimental campaign across five architectures (Qwen3, LLaMA\nobreakdash-3.1, Gemma\nobreakdash-3, GPT-2, Mistral) spanning $124$M to $8$B parameters. The empirical observables are quantitatively consistent with the geometric predictions: the $\epsilon^{-1/2}$ Lipschitz scaling calibration at machine precision ($R^2 = 1.000$), the Lie--Trotter operator-splitting torsion, the symmetric ablation instability confirming the Dual-Law of Topological Stability, the $\calO(1/\sqrt{k})$ thermodynamic suppression of Poincar\'e recurrence on the RoPE torus, the thermodynamic context-limit phase transition, and the Non-Equilibrium Steady State parameter vortex -- verified across two optimizers (AdamW and Pure SGD) to exclude momentum artifacts. The results demonstrate that analyzing Transformers through the lens of continuous stochastic differential geometry provides a predictive descriptive vocabulary for the stability limits, context bounds, and optimization dynamics of Large Language Models.