Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

For physical AI, the primary constraint is not the cost of data but its fundamental non-existence. Unlike software AI, you can't advance without deploying robots "in the wild" to capture edge cases—a classic chicken-and-egg problem that simulation alone cannot solve and capital cannot easily buy.

Related Insights

The primary challenge in robotics AI is the lack of real-world training data. To solve this, models are bootstrapped using a combination of learning from human lifestyle videos and extensive simulation environments. This creates a foundational model capable of initial deployment, which then generates a real-world data flywheel.

According to Figure's CEO, the company's biggest challenge is no longer hardware reliability but acquiring enormous amounts of diverse, high-quality data. This data is essential for pre-training their Helix AI model to generalize and handle countless real-world scenarios in homes and commercial settings.

The rapid progress of many LLMs was possible because they could leverage the same massive public dataset: the internet. In robotics, no such public corpus of robot interaction data exists. This “data void” means progress is tied to a company's ability to generate its own proprietary data.

Progress in robotics for household tasks is limited by a scarcity of real-world training data, not mechanical engineering. Companies are now deploying capital-intensive "in-field" teams to collect multi-modal data from inside homes, capturing the complexity of mundane human activities to train more capable robots.

The future of valuable AI lies not in models trained on the abundant public internet, but in those built on scarce, proprietary data. For fields like robotics and biology, this data doesn't exist to be scraped; it must be actively created, making the data generation process itself the key competitive moat.

Ken Goldberg quantifies the challenge: the text data used to train LLMs would take a human 100,000 years to read. Equivalent data for robot manipulation (vision-to-control signals) doesn't exist online and must be generated from scratch, explaining the slower progress in physical AI.

The humanoid robot industry is stalled by a data paradox: robots need vast amounts of real-world data from factory tasks to become useful, but they cannot be deployed in factories until they are already useful. This catch-22 forces companies to rely on simulated data, slowing the transition from entertainment props to industrial tools.

Brett Adcock states that Figure AI's "Helix 2" neural net provides the right technical stack for general robotics. The biggest remaining obstacle is not hardware but the immense data required to train the robot for a wide distribution of tasks. The company plans to spend nine figures on data acquisition in 2026 to solve this.

Despite industry hype, humanoid robots are not imminent. They lack the massive datasets of real-world, unpredictable interactions needed to operate safely and usefully in a home environment, which is far more complex than a structured factory floor.

For robotics companies, market dominance hinges on a data flywheel effect. This requires rapidly deploying robots into real-world environments, even at a financial loss, because each unit acts as a data source. A small lead in data collection today translates into a massive competitive advantage tomorrow.