Search papers, labs, and topics across Lattice.
This paper formulates a mathematical framework for understanding slow thinking and active perception, introducing a theory called "active lifting" that focuses on sampling latent sequences to minimize uncertainty. By positioning slow thinking models within a representation and sampler hierarchy, the authors derive a comprehensive design space that facilitates the development of these models and their training processes. Key findings suggest that this framework not only enhances the inference capabilities of slow thinking models but also offers insights into human-like perception and potential solutions to policy collapse in AI systems.
Active lifting reveals a novel pathway to enhance slow thinking models, bridging cognitive theory and practical AI applications.
As part of a series on first-principles modeling of cognitive functions, this paper attempts to provide a mathematical formulation of thinking and perception. It formally derives slow thinking or more generally, active perception, and encompasses the design, training and inference of slow thinking large language models. Our starting point is the lifting and projection of probability distributions on the observable and latent paces,with the objective of representing complex data distributions by simple function families such as neural networks. A theory called “active lifting“ is proposed, based on the sampling of latent sequences and an intrinsic drive to reduce uncertainty with maximum rate. It derives a large design space, containing the slow thinking models in a subspace that we call the static theory. These models are positioned on the representation hierarchy and sampler hierarchy induced by the static theory, and can be upgraded by climbing the two hierarchies. Active lifting further derives an inference process with an internal time axis, and a training objective that resembles minimum-length coding as well as the invention of languages. Thus, it characterizes the agency of perception, including the emergence of the slow thinking formats. Technical by-products of this theory include a three-stage pathway for improving slow thinking models, a unified approach to constructing encoders and generative models for all data modalities, a priori formation of human-like visual representations, and a possible solution to policy collapse.