Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Countering Yann LeCun, Runway's cofounder argues that predicting video frames at scale *is* learning world dynamics. As models scale, their ability to simulate physics predictably improves on benchmarks. This aligns with the "bitter lesson" of AI: general methods that leverage computation and data ultimately outperform specialized, human-designed ones.

Related Insights

Sora 2's most significant advancement is not its visual quality, but its ability to understand and simulate physics. The model accurately portrays how water splashes or vehicles kick up snow, demonstrating a grasp of cause and effect crucial for true world-building.

Startups and major labs are focusing on "world models," which simulate physical reality, cause, and effect. This is seen as the necessary step beyond text-based LLMs to create agents that can truly understand and interact with the physical world, a key step towards AGI.

GI discovered their world model, trained on game footage, could generate a realistic camera shake during an in-game explosion—a physical effect not part of the game's engine. This suggests the models are learning an implicit understanding of real-world physics and can generate plausible phenomena that go beyond their source material.

The paradigm shift with AI is not an abandonment of physical laws. Instead of using supercomputers to approximate solutions to physics equations, AI learns the patterns governed by those laws directly from historical data. The ultimate goal is to forecast direct impacts, not just variables.

To perform complex edits like 'knock over this water glass,' a model must understand physics, causality, and object relationships. This requirement inadvertently builds a form of visual intelligence that serves as a precursor to more sophisticated world models for applications like robotics.

Large Language Models are limited because they lack an understanding of the physical world. The next evolution is 'World Models'—AI trained on real-world sensory data to understand physics, space, and context. This is the foundational technology required to unlock physical AI like advanced robotics.

Cuban believes today's LLMs, trained on text and images, are a limited step. The next leap will be "worldview" models trained on the fundamental physics of the real world, using data from video and sensors to understand cause and effect, not just language patterns.

The concept of a 'world model' is evolving from action-conditioned video predictors to single, multimodal models like Google's Omni. Omni demonstrates a deep, scalable understanding of the world, shown through nuanced video editing, representing a more practical approach than traditional, computationally expensive architectures.

Prof. Cho outlines two competing visions for world models. One camp believes in high-fidelity, step-by-step prediction (e.g., video generation). The other, which he and Yann LeCun favor, argues for abstract, high-level latent models that can plan without simulating every detail, akin to human thinking.

Demis Hassabis sees video generation as more than a content tool; it's a step toward building AI with "world models." By learning to generate realistic scenes, these models develop an intuitive understanding of physics and causality, a foundational capability for AGI to perform long-term planning in the real world.