Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Language models execute identically regardless of device, but robotics software is constrained by the physical robot body. Keerthana Gopalakrishnan places current robotics in the 'GPT-2 era' because generic robotic brains cannot reliably control arbitrary, unseen robot embodiments out of the box, nor can they reliably generalize from few-shot demonstrations across hard, diverse tasks.

Related Insights

While language models understand the world through text, Demis Hassabis argues they lack an intuitive grasp of physics and spatial dynamics. He sees 'world models'—simulations that understand cause and effect in the physical world—as the critical technology needed to advance AI from digital tasks to effective robotics.

While AI computation improves exponentially, physical robot hardware evolves very slowly. A robot's hand is vastly inferior to a human's, which has millions of sensors and self-healing capabilities. This physical limitation is the primary barrier to creating AIs that can operate effectively in the real world.

Robotic intelligence has two components. "Reasoning," which involves creating a plan, is quickly being solved by AI. The other, harder part is "movement"—the robot's physical dexterity to execute that plan reliably in a complex environment without tripping or failing.

The robotics field has a scalable recipe for AI-driven manipulation (like GPT), but hasn't yet scaled it into a polished, mass-market consumer product (like ChatGPT). The current phase focuses on scaling data and refining systems, not just fundamental algorithm discovery, to bridge this gap.

Ken Goldberg quantifies the challenge: the text data used to train LLMs would take a human 100,000 years to read. Equivalent data for robot manipulation (vision-to-control signals) doesn't exist online and must be generated from scratch, explaining the slower progress in physical AI.

Neurobotics posits that true physical AI requires more than just vision-language models; it needs a "nervous system" and reflexes. They advocate for training robots in physical "gyms" to collect embodied data, arguing that complex physical tasks cannot be learned solely by watching videos.

Robots can generalize skills to new hardware by using an intermediate 'thinking' step. The model generates an image of the next desired milestone (e.g., a half-folded shirt). This visual goal is embodiment-agnostic and easier to create than new motor commands, allowing the robot to then solve for the actions to match the image.

While China's humanoid hardware demonstrates impressive locomotion in programmed tasks, the major obstacle to widespread deployment is the "robot brain." Current AI lacks the ability to autonomously navigate unpredictable, real-world environments, making massive data collection the current R&D focus.

Unlike LLMs trained on vast digital text, humanoid robots need immense amounts of real-world physical data to learn simple tasks. It's estimated that 100 million hours—over 11,000 years' worth—of interaction data is needed to create truly smart, useful humanoids, highlighting the scale of the challenge.

The "bitter lesson" (scale and simple models win) works for language because training data (text) aligns with the output (text). Robotics faces a critical misalignment: it's trained on passive web videos but needs to output physical actions in a 3D world. This data gap is a fundamental hurdle that pure scaling cannot solve.