Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Adopting the true 'foundation model ethos' from LLMs is difficult for roboticists. It means a warehouse automation company should collect data from kitchen robots. This breadth, while seemingly unrelated, builds a generalist model that better handles the weird edge cases in the target domain than a narrowly trained specialist model.

Related Insights

Counterintuitively, training AI models with data from disparate physical domains, like mining, improves the performance of systems in completely different areas, such as self-driving cars. This cross-domain learning suggests that a broad understanding of the physical world is key to robust, real-world AI.

To build generalist robots, the most effective approach is pre-training foundation models on internet-scale video datasets, not just simulation or tele-operated data. This vast, diverse data provides a deep, implicit understanding of physics and object interaction that is impossible to replicate in controlled environments, enabling true generalization.

According to Figure's CEO, the company's biggest challenge is no longer hardware reliability but acquiring enormous amounts of diverse, high-quality data. This data is essential for pre-training their Helix AI model to generalize and handle countless real-world scenarios in homes and commercial settings.

The rapid progress of many LLMs was possible because they could leverage the same massive public dataset: the internet. In robotics, no such public corpus of robot interaction data exists. This “data void” means progress is tied to a company's ability to generate its own proprietary data.

The Physical Intelligence thesis is that a foundation model learning from diverse data can achieve a "physical understanding" of the world, making it easier to adapt to new tasks than building single-purpose robots from scratch. Generality leverages broader data, which is ultimately a more scalable approach.

Figure is observing that data from one robot performing a task (e.g., moving packages in a warehouse) improves the performance of other robots on completely different tasks (e.g., folding laundry at home). This powerful transfer learning, enabled by deep learning, is a key driver for scaling general-purpose capabilities.

For physical AI systems like robots, data quality hinges on diversity, not just quantity. A robot trained to make a bed in one specific lighting condition may fail completely if the lighting changes or the bed is moved. This brittleness highlights a key challenge: training data must capture a wide variety of contexts and edge cases to enable real-world generalization.

Simply collecting more data from a deployed robot isn't enough to create a powerful learning flywheel. If the tasks are repetitive (e.g., a million car welds), the model won't generalize. Data must be diverse, acting more like an 'education program' than a fungible commodity to drive real capability growth.

Counterintuitively, the best way to train a robot foundation model isn't to start with vast human video datasets. Research indicates that starting with real, embodied robot data provides a physical 'grounding' that allows the model to more effectively absorb and contextualize other data sources, like human videos, later on.

For unpredictable situations where a robot has no prior training data (e.g., a "gas leak" sign), multimodal LLMs can provide the necessary world knowledge to reason and act appropriately. This solves the long-standing robotics problem of how to handle the long tail of real-world scenarios.