We scan new podcasts and send you the top 5 insights daily.
Pure visual feedback is insufficient for delicate real-world manipulation tasks such as stacking brittle items or assembling components. To avoid crushing materials, robots require real-time end-effector force sensing and physical back-drivability/compliance. Having explicit torque and pressure signals allows neural policies to make safe adjustments that vision alone cannot deduce.
Instead of loading robots with costly sensors for touch or force, powerful learning models can infer physical properties from simple cameras. A wrist camera can act as a "touch sensor in disguise" by observing local deformations, dramatically lowering hardware costs and complexity for scalable robotics.
Unlike digital AI safety which emphasizes guardrails against bad actors and deceptive behavior, the core safety challenge in physical robotics is basic operational competence. A humanoid robot falling over or misjudging sensor inputs poses serious physical danger to humans entirely through accidental clumsiness rather than malevolent intent.
Leading roboticist Ken Goldberg clarifies that while legged robots show immense progress in navigation, fine motor skills for tasks like tying shoelaces are far beyond current capabilities. This is due to challenges in sensing and handling deformable, unpredictable objects in the real world.
Humanoid running has progressed rapidly because locomotion on flat, rigid surfaces is straightforward to model in simulation and transfer to the real world. In contrast, manipulation requires modeling complex contact dynamics, friction, and deformable objects like cloth. Where physics is harder to model, current simulations break down, leaving manipulation reliant on slower, harder real-world data collection.
Neurobotics posits that true physical AI requires more than just vision-language models; it needs a "nervous system" and reflexes. They advocate for training robots in physical "gyms" to collect embodied data, arguing that complex physical tasks cannot be learned solely by watching videos.
According to Agility Robotics' co-founder, perception is now a largely solved problem. The new frontier is generating training data for robot control—the specific torque commands and sensor inputs for actions. Unlike text or images for LLMs, this data does not exist on the internet and must be painstakingly created.
No single data collection method will solve robotics. Teleoperation provides high-precision ground truth for specific robots but scales poorly and degrades when hardware updates. Universal Manipulation Interface (UMI) data offers better scale and precision via sensors but requires hardware maintenance. Egocentric human video provides massive scale but lacks precise end-effector force signals and suffers from human-to-robot physical embodiment mismatch.
Surgeons perform intricate tasks without tactile feedback, relying on visual cues of tissue deformation. This suggests robotics could achieve complex manipulation by advancing visual interpretation of physical interactions, bypassing the immense difficulty of creating and integrating artificial touch sensors.
Classical robots required expensive, rigid, and precise hardware because they were blind. Modern AI perception acts as 'eyes', allowing robots to correct for inaccuracies in real-time. This enables the use of cheaper, compliant, and inherently safer mechanical components, fundamentally changing hardware design philosophy.
Manipulating deformable objects like towels was long considered one of the final, hardest challenges in robotics due to their infinite variations. The fact that Figure's neural networks can now successfully fold laundry indicates that the core technological hurdles for truly general-purpose robots have been overcome.