Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The idea that LLMs just predict the next statistical character is wrong. To accurately predict the next word in a physics paper, for instance, an AI must create an internal model of the physical world. This implies that scaling current architectures can lead to deeper understanding, not just pattern matching.

Related Insights

The seemingly simple task of next-token prediction, when perfected, requires a model to understand concepts as deeply as the source. To accurately predict what Einstein would say in a new situation, a system must be as intelligent as Einstein, proving prediction is fundamental to intelligence.

A core debate in AI is whether LLMs, which are text prediction engines, can achieve true intelligence. Critics argue they cannot because they lack a model of the real world. This prevents them from making meaningful, context-aware predictions about future events—a limitation that more data alone may not solve.

The complexity in LLMs isn't intelligence emerging in silicon; it reflects our own. These models are deep because they encode the vast, causally powerful structure of human language and culture. We are looking at a high-resolution imprint of our own collective mind.

The next major leap in AI may come from "world models," which aim to give LLMs an experiential, physical understanding of concepts like space and physics. This mirrors the difference between knowing facts from a book and having real-world experience.

Training a language model to predict the next amino acid in a sequence forces it to learn the protein's 3D structure. To make accurate predictions, the model must understand an amino acid's physical microenvironment, effectively deriving 3D spatial relationships from 1D sequence data alone. This demonstrates emergent capabilities of LLMs in biology.

Startups and major labs are focusing on "world models," which simulate physical reality, cause, and effect. This is seen as the necessary step beyond text-based LLMs to create agents that can truly understand and interact with the physical world, a key step towards AGI.

The argument that LLMs are just "stochastic parrots" is outdated. Current frontier models are trained via Reinforcement Learning, where the signal is not "did you predict the right token?" but "did you get the right answer?" This is based on complex, often qualitative criteria, pushing models beyond simple statistical correlation.

A non-obvious aspect of LLM training is that to accurately predict text describing reality (e.g., a lab result), the AI must learn to model the underlying physics or logic. This inherently trains it to be more capable than the human who merely observed the result, forcing emergent intelligence.

Large Language Models are limited because they lack an understanding of the physical world. The next evolution is 'World Models'—AI trained on real-world sensory data to understand physics, space, and context. This is the foundational technology required to unlock physical AI like advanced robotics.

We can now prove that LLMs are not just correlating tokens but are developing sophisticated internal world models. Techniques like sparse autoencoders untangle the network's dense activations, revealing distinct, manipulable concepts like "Golden Gate Bridge." This conclusively demonstrates a deeper, conceptual understanding within the models.

LLMs Build World Models to Predict Text, Refuting Simpler Explanations | RiffOn