We scan new podcasts and send you the top 5 insights daily.
Building narrow AI creates high performance on one or two tasks, but the nth task costs as much to train as the first. In contrast, building generalist baselines endows models with broad common sense and physical reasoning. Once this broad base is established, tuning a robot to high reliability and mastery on any arbitrary task becomes significantly faster and cheaper.
To build generalist robots, the most effective approach is pre-training foundation models on internet-scale video datasets, not just simulation or tele-operated data. This vast, diverse data provides a deep, implicit understanding of physics and object interaction that is impossible to replicate in controlled environments, enabling true generalization.
The path to a general-purpose AI model is not to tackle the entire problem at once. A more effective strategy is to start with a highly constrained domain, like generating only Minecraft videos. Once the model works reliably in that narrow distribution, incrementally expand the training data and complexity, using each step as a foundation for the next.
The Physical Intelligence thesis is that a foundation model learning from diverse data can achieve a "physical understanding" of the world, making it easier to adapt to new tasks than building single-purpose robots from scratch. Generality leverages broader data, which is ultimately a more scalable approach.
Figure is observing that data from one robot performing a task (e.g., moving packages in a warehouse) improves the performance of other robots on completely different tasks (e.g., folding laundry at home). This powerful transfer learning, enabled by deep learning, is a key driver for scaling general-purpose capabilities.
A flashy robot demo typically uses a highly controlled, pristine environment tailored to one task. True progress lies in a robot performing a mundane task reliably in any novel situation—a feat of generalization that is much harder to showcase visually and less exciting to a layperson.
Adopting the true 'foundation model ethos' from LLMs is difficult for roboticists. It means a warehouse automation company should collect data from kitchen robots. This breadth, while seemingly unrelated, builds a generalist model that better handles the weird edge cases in the target domain than a narrowly trained specialist model.
Runway’s robotics thesis is that pre-training on massive, easily available third-person video data (e.g., people performing tasks) is more scalable and effective than relying on expensive, limited teleoperation or first-person data. This general world knowledge can then be fine-tuned for specific robotic tasks.
Sunday Robotics found that as they scaled up pre-training data and compute for their laundry-folding robot, it developed the ability to learn a new task from a single demonstration. This suggests that complex abilities like one-shot learning don't need to be explicitly programmed but can emerge from scaled-up general training.
By solving the core "intelligence" problem with a foundation model, the barrier to entry for creating novel robotic applications and form factors will dramatically decrease. This will enable a "Cambrian explosion" of hardware creativity, as builders will no longer need to solve AI from scratch for each new idea.
Unlike older robots requiring precise maps and trajectory calculations, new robots use internet-scale common sense and learn motion by mimicking humans or simulations. This combination has “wiped the slate clean” for what is possible in the field.