Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Robots can generalize skills to new hardware by using an intermediate 'thinking' step. The model generates an image of the next desired milestone (e.g., a half-folded shirt). This visual goal is embodiment-agnostic and easier to create than new motor commands, allowing the robot to then solve for the actions to match the image.

Related Insights

The Physical Intelligence thesis is that a foundation model learning from diverse data can achieve a "physical understanding" of the world, making it easier to adapt to new tasks than building single-purpose robots from scratch. Generality leverages broader data, which is ultimately a more scalable approach.

The cutting edge of physical AI involves more than just programming a robot's response to a stimulus ("policy"). It also requires a "world capability"—a virtual twin that simulates and predicts outcomes, allowing the physical robot to choose intelligent actions based on those predictions.

Figure is observing that data from one robot performing a task (e.g., moving packages in a warehouse) improves the performance of other robots on completely different tasks (e.g., folding laundry at home). This powerful transfer learning, enabled by deep learning, is a key driver for scaling general-purpose capabilities.

A new model architecture allows robots to vary their internal 'thinking' iterations at test time. This lets practitioners trade response speed for decision accuracy on a case-by-case basis, boosting performance on complex tasks without needing to retrain the model.

Instead of simulating photorealistic worlds, robotics firm Flexion trains its models on simplified, abstract representations. For example, it uses perception models like Segment Anything to 'paint' a door red and its handle green. By training on this simplified abstraction, the robot learns the core task (opening doors) in a way that generalizes across all real-world doors, bypassing the need for perfect simulation.

Sunday Robotics found that as they scaled up pre-training data and compute for their laundry-folding robot, it developed the ability to learn a new task from a single demonstration. This suggests that complex abilities like one-shot learning don't need to be explicitly programmed but can emerge from scaled-up general training.

Neurological studies show the human brain maps a tool's tip as if it were our hand. This implies that a powerful physical intelligence should not be tied to a specific body (e.g., a humanoid) but should be a general "brain" capable of controlling any embodiment, from a bulldozer to a multi-fingered hand.

Robots have become so capable at low-level physical tasks that the primary bottleneck has shifted to "mid-level reasoning"—interpreting a scene and choosing the correct next action. This means improvement can come from high-level language-based coaching, not just more physical demonstration data, which is a major breakthrough.

Intuition Robotics' core bet is that the transfer from simulated to physical worlds is unlocked by a shared action interface. Since many real-world robots like drones and arms are already operated with game controllers, an agent trained in diverse gaming environments only needs to adapt to a new visual world, not an entirely new action space.

Manipulating deformable objects like towels was long considered one of the final, hardest challenges in robotics due to their infinite variations. The fact that Figure's neural networks can now successfully fold laundry indicates that the core technological hurdles for truly general-purpose robots have been overcome.