We scan new podcasts and send you the top 5 insights daily.
Simply collecting more data from a deployed robot isn't enough to create a powerful learning flywheel. If the tasks are repetitive (e.g., a million car welds), the model won't generalize. Data must be diverse, acting more like an 'education program' than a fungible commodity to drive real capability growth.
Counterintuitively, training AI models with data from disparate physical domains, like mining, improves the performance of systems in completely different areas, such as self-driving cars. This cross-domain learning suggests that a broad understanding of the physical world is key to robust, real-world AI.
The primary challenge in robotics AI is the lack of real-world training data. To solve this, models are bootstrapped using a combination of learning from human lifestyle videos and extensive simulation environments. This creates a foundational model capable of initial deployment, which then generates a real-world data flywheel.
To build generalist robots, the most effective approach is pre-training foundation models on internet-scale video datasets, not just simulation or tele-operated data. This vast, diverse data provides a deep, implicit understanding of physics and object interaction that is impossible to replicate in controlled environments, enabling true generalization.
Figure is observing that data from one robot performing a task (e.g., moving packages in a warehouse) improves the performance of other robots on completely different tasks (e.g., folding laundry at home). This powerful transfer learning, enabled by deep learning, is a key driver for scaling general-purpose capabilities.
For physical AI systems like robots, data quality hinges on diversity, not just quantity. A robot trained to make a bed in one specific lighting condition may fail completely if the lighting changes or the bed is moved. This brittleness highlights a key challenge: training data must capture a wide variety of contexts and edge cases to enable real-world generalization.
A flashy robot demo typically uses a highly controlled, pristine environment tailored to one task. True progress lies in a robot performing a mundane task reliably in any novel situation—a feat of generalization that is much harder to showcase visually and less exciting to a layperson.
Adopting the true 'foundation model ethos' from LLMs is difficult for roboticists. It means a warehouse automation company should collect data from kitchen robots. This breadth, while seemingly unrelated, builds a generalist model that better handles the weird edge cases in the target domain than a narrowly trained specialist model.
The adoption of powerful AI architectures like transformers in robotics was bottlenecked by data quality, not algorithmic invention. Only after data collection methods improved to capture more dexterous, high-fidelity human actions did these advanced models become effective, reversing the typical 'algorithm-first' narrative of AI progress.
For physical AI, the primary constraint is not the cost of data but its fundamental non-existence. Unlike software AI, you can't advance without deploying robots "in the wild" to capture edge cases—a classic chicken-and-egg problem that simulation alone cannot solve and capital cannot easily buy.
For robotics companies, market dominance hinges on a data flywheel effect. This requires rapidly deploying robots into real-world environments, even at a financial loss, because each unit acts as a data source. A small lead in data collection today translates into a massive competitive advantage tomorrow.