We scan new podcasts and send you the top 5 insights daily.
Humanoid running has progressed rapidly because locomotion on flat, rigid surfaces is straightforward to model in simulation and transfer to the real world. In contrast, manipulation requires modeling complex contact dynamics, friction, and deformable objects like cloth. Where physics is harder to model, current simulations break down, leaving manipulation reliant on slower, harder real-world data collection.
A surprise technical leap—from 'dreamlike' simulations to models with robust object permanence—dramatically accelerated expert timelines for solving dexterous robotics. This breakthrough allows for vast, cheap generation of high-quality synthetic training data.
Leading roboticist Ken Goldberg clarifies that while legged robots show immense progress in navigation, fine motor skills for tasks like tying shoelaces are far beyond current capabilities. This is due to challenges in sensing and handling deformable, unpredictable objects in the real world.
A major hurdle in robotics is the laborious collection of real-world training data. Atlas accelerates this by creating high-fidelity simulations from sparse real-world images ("real-to-sim"), enabling rapid training and randomization of robotic policies without extensive data capture.
The choice between simulation and real-world data depends on a task's core difficulty. For locomotion, complex reactive behavior is harder to capture than simple ground physics, favoring simulation. For manipulation, complex object physics are harder to simulate than simple grasping behaviors, favoring real-world data.
Robotic intelligence has two components. "Reasoning," which involves creating a plan, is quickly being solved by AI. The other, harder part is "movement"—the robot's physical dexterity to execute that plan reliably in a complex environment without tripping or failing.
Ken Goldberg quantifies the challenge: the text data used to train LLMs would take a human 100,000 years to read. Equivalent data for robot manipulation (vision-to-control signals) doesn't exist online and must be generated from scratch, explaining the slower progress in physical AI.
Self-driving cars, a 20-year journey so far, are relatively simple robots: metal boxes on 2D surfaces designed *not* to touch things. General-purpose robots operate in complex 3D environments with the primary goal of *touching* and manipulating objects. This highlights the immense, often underestimated, physical and algorithmic challenges facing robotics.
Unlike LLMs trained on vast digital text, humanoid robots need immense amounts of real-world physical data to learn simple tasks. It's estimated that 100 million hours—over 11,000 years' worth—of interaction data is needed to create truly smart, useful humanoids, highlighting the scale of the challenge.
Despite industry hype, humanoid robots are not imminent. They lack the massive datasets of real-world, unpredictable interactions needed to operate safely and usefully in a home environment, which is far more complex than a structured factory floor.
For physical AI, the primary constraint is not the cost of data but its fundamental non-existence. Unlike software AI, you can't advance without deploying robots "in the wild" to capture edge cases—a classic chicken-and-egg problem that simulation alone cannot solve and capital cannot easily buy.