Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Sergey Levine explains that humanoid robotics is currently focused on identifying fundamental, scalable technologies, similar to the pre-transformer era of LSTMs. The industry is not yet in a predictable, industrial-scale growth phase like current LLMs, but is instead assembling the necessary puzzle pieces for future scaling.

Related Insights

Insiders in top robotics labs are witnessing fundamental breakthroughs. These “signs of life,” while rudimentary now, are clear precursors to a rapid transition from research to widely adopted products, much like AI before ChatGPT’s public release.

Vision Language Action models (VLAs) have not yet produced a 'ChatGPT moment' for robotics. Consequently, investor enthusiasm and capital are increasingly flowing towards the alternative 'World Model' approach, which learns physics from video, even though it has yet to demonstrate superior tangible results.

Large Language Models are limited because they lack an understanding of the physical world. The next evolution is 'World Models'—AI trained on real-world sensory data to understand physics, space, and context. This is the foundational technology required to unlock physical AI like advanced robotics.

The adoption of humanoid robots will mirror that of autonomous vehicles: focus on achievable, single-task applications first. Instead of a complex, general-purpose home robot, the market will first embrace robots trained for specific, repeatable industrial tasks like warehouse logistics or shelf stocking.

The adoption of powerful AI architectures like transformers in robotics was bottlenecked by data quality, not algorithmic invention. Only after data collection methods improved to capture more dexterous, high-fidelity human actions did these advanced models become effective, reversing the typical 'algorithm-first' narrative of AI progress.

The robotics field has a scalable recipe for AI-driven manipulation (like GPT), but hasn't yet scaled it into a polished, mass-market consumer product (like ChatGPT). The current phase focuses on scaling data and refining systems, not just fundamental algorithm discovery, to bridge this gap.

Ken Goldberg quantifies the challenge: the text data used to train LLMs would take a human 100,000 years to read. Equivalent data for robot manipulation (vision-to-control signals) doesn't exist online and must be generated from scratch, explaining the slower progress in physical AI.

According to Agility Robotics' co-founder, perception is now a largely solved problem. The new frontier is generating training data for robot control—the specific torque commands and sensor inputs for actions. Unlike text or images for LLMs, this data does not exist on the internet and must be painstakingly created.

While China's humanoid hardware demonstrates impressive locomotion in programmed tasks, the major obstacle to widespread deployment is the "robot brain." Current AI lacks the ability to autonomously navigate unpredictable, real-world environments, making massive data collection the current R&D focus.

Unlike older robots requiring precise maps and trajectory calculations, new robots use internet-scale common sense and learn motion by mimicking humans or simulations. This combination has “wiped the slate clean” for what is possible in the field.