Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Instead of creating separate models for understanding the past (e.g., explaining a video) and simulating the future, NVIDIA is merging these capabilities. Their Cosmos model shares a core world representation, creating a unified "omni model" that handles diverse inputs (text, video, action) and outputs.

Related Insights

The next major leap in AI may come from "world models," which aim to give LLMs an experiential, physical understanding of concepts like space and physics. This mirrors the difference between knowing facts from a book and having real-world experience.

Human understanding is the ability to connect new information to a global, unified model of the universe. Until recently, AI models were isolated (e.g., a chess model). The major advance with large multimodal models is their ability to create a single, cohesive reality model, enabling true, generalizable understanding.

Startups and major labs are focusing on "world models," which simulate physical reality, cause, and effect. This is seen as the necessary step beyond text-based LLMs to create agents that can truly understand and interact with the physical world, a key step towards AGI.

Large language models are insufficient for tasks requiring real-world interaction and spatial understanding, like robotics or disaster response. World models provide this missing piece by generating interactive, reason-able 3D environments. They represent a foundational shift from language-based AI to a more holistic, spatially intelligent AI.

Large Language Models are limited because they lack an understanding of the physical world. The next evolution is 'World Models'—AI trained on real-world sensory data to understand physics, space, and context. This is the foundational technology required to unlock physical AI like advanced robotics.

The concept of a 'world model' is evolving from action-conditioned video predictors to single, multimodal models like Google's Omni. Omni demonstrates a deep, scalable understanding of the world, shown through nuanced video editing, representing a more practical approach than traditional, computationally expensive architectures.

For the first time, Atlas combines the traditionally separate fields of creative pixel generation (like text-to-video) and precise 3D reconstruction into one architecture. This dual capability allows it to both imagine and accurately map physical spaces.

Traditional simulators are rule-based, programmed with physics equations. World models pioneer "neural simulation," which learns the physics of the world implicitly from massive datasets of visual observations. It's a pattern-recognition approach to predicting outcomes, rather than one based on pre-defined rules.

World Labs posits that "world models"—AI focused on visual and physical understanding—represent a new general-purpose platform, similar to LLMs for text. These models can generate, simulate, and reconstruct physical worlds, with applications spanning from robotics and construction to entertainment and VR.

Black Forest Labs is developing multimodal models that understand and generate images, video, and audio while also predicting actions. This convergence means the same fundamental technology used as a creative tool for filmmaking can also be deployed as the 'brain' for a physical robot, unifying the digital and physical worlds under a single AI paradigm.