Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Unlike LLMs trained on vast digital text, humanoid robots need immense amounts of real-world physical data to learn simple tasks. It's estimated that 100 million hours—over 11,000 years' worth—of interaction data is needed to create truly smart, useful humanoids, highlighting the scale of the challenge.

Related Insights

Unlike cars, which gather data passively, humanoid robots need active training. To solve this, Musk's strategy is to build a physical 'academy' of 10,000-30,000 Optimus robots performing self-play on various tasks, using this real-world data to close the 'sim-to-real' gap from millions of simulated robots.

Progress in robotics for household tasks is limited by a scarcity of real-world training data, not mechanical engineering. Companies are now deploying capital-intensive "in-field" teams to collect multi-modal data from inside homes, capturing the complexity of mundane human activities to train more capable robots.

Generalist CEO Pete Florence provides a tier list for robotics training data. He ranks "lived experience of the physical world" as S-tier, emphasizing the irreplaceable value of high-quality, real-world data. In contrast, he rates synthetic data from world models as F-tier, suggesting it is far less effective.

Ken Goldberg quantifies the challenge: the text data used to train LLMs would take a human 100,000 years to read. Equivalent data for robot manipulation (vision-to-control signals) doesn't exist online and must be generated from scratch, explaining the slower progress in physical AI.

Neurobotics posits that true physical AI requires more than just vision-language models; it needs a "nervous system" and reflexes. They advocate for training robots in physical "gyms" to collect embodied data, arguing that complex physical tasks cannot be learned solely by watching videos.

According to Agility Robotics' co-founder, perception is now a largely solved problem. The new frontier is generating training data for robot control—the specific torque commands and sensor inputs for actions. Unlike text or images for LLMs, this data does not exist on the internet and must be painstakingly created.

The humanoid robot industry is stalled by a data paradox: robots need vast amounts of real-world data from factory tasks to become useful, but they cannot be deployed in factories until they are already useful. This catch-22 forces companies to rely on simulated data, slowing the transition from entertainment props to industrial tools.

Despite industry hype, humanoid robots are not imminent. They lack the massive datasets of real-world, unpredictable interactions needed to operate safely and usefully in a home environment, which is far more complex than a structured factory floor.

The "bitter lesson" (scale and simple models win) works for language because training data (text) aligns with the output (text). Robotics faces a critical misalignment: it's trained on passive web videos but needs to output physical actions in a 3D world. This data gap is a fundamental hurdle that pure scaling cannot solve.

For physical AI, the primary constraint is not the cost of data but its fundamental non-existence. Unlike software AI, you can't advance without deploying robots "in the wild" to capture edge cases—a classic chicken-and-egg problem that simulation alone cannot solve and capital cannot easily buy.