A First-Principles Theory of Slow Thinking and Active Perception
Summary
A "First-Principles Theory of Slow Thinking and Active Perception," by Hongkang Yang, Zhi-Qin John Xu, Feiyu Xiong, and Weinan E, published on 2026/05/11 in Journal of Machine Learning, offers a mathematical formulation for cognitive functions. This theory formally derives slow thinking and active perception, covering the design, training, and inference of slow thinking large language models. Its foundation, "active lifting," uses lifting and projection of probability distributions on observable and latent spaces to represent complex data via neural networks. It operates by sampling latent sequences and intrinsically reducing uncertainty at a maximum rate. The theory defines a design space for slow thinking models, positioned on representation and sampler hierarchies, and outlines a three-stage improvement pathway. It also derives an inference process with an internal time axis and a training objective resembling minimum-length coding. Technical by-products include a unified approach for encoders and generative models across data modalities, a priori human-like visual representations, and a potential solution to policy collapse.
Key takeaway
For AI Scientists and Machine Learning Engineers designing or improving large language models, this first-principles theory of slow thinking and active perception offers a novel mathematical foundation. You should explore the "active lifting" framework to develop LLMs with emergent slow thinking capabilities. This approach provides a structured three-stage pathway for model upgrades and a unified method for constructing encoders and generative models across all data modalities, potentially offering solutions to challenges like policy collapse.
Key insights
"Active lifting" theory mathematically derives slow thinking and active perception, informing LLM design, training, and inference from first principles.
Principles
- Complex data distributions can be represented by simple function families.
- Reduce uncertainty with maximum rate drives active perception.
- Models can be upgraded by climbing representation and sampler hierarchies.
Method
Proposes "active lifting" based on sampling latent sequences and an intrinsic drive to reduce uncertainty. It derives an inference process with an internal time axis and a minimum-length coding training objective.
In practice
- Three-stage pathway for improving slow thinking models.
- Unified approach for encoders and generative models across modalities.
- A priori formation of human-like visual representations.
Topics
- Slow Thinking
- Active Perception
- Large Language Models
- Active Lifting Theory
- Generative Models
- Cognitive Modeling
Best for: Research Scientist, AI Scientist, Machine Learning Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by cs.AI updates on arXiv.org.