Sergey Levine explains that humanoid robotics is currently focused on identifying fundamental, scalable technologies, similar to the pre-transformer era of LSTMs. The industry is not yet in a predictable, industrial-scale growth phase like current LLMs, but is instead assembling the necessary puzzle pieces for future scaling.
Sergey Levine finds that robots making 'sensible' mistakes, like putting utensils in an oven when a drawer is stuck, is a positive sign. These errors demonstrate a degree of common-sense reasoning, similar to a child's logic, which is a significant advancement over random or nonsensical failures.
Simply collecting more data from a deployed robot isn't enough to create a powerful learning flywheel. If the tasks are repetitive (e.g., a million car welds), the model won't generalize. Data must be diverse, acting more like an 'education program' than a fungible commodity to drive real capability growth.
Attempts to simplify robotics by creating highly structured environments are doomed to fail, much like early autonomous driving efforts to instrument highways. The 1% of real-world exceptions will always break the system. True progress comes from tackling messy, unstructured environments head-on, as Waymo did in San Francisco.
Counterintuitively, the best way to train a robot foundation model isn't to start with vast human video datasets. Research indicates that starting with real, embodied robot data provides a physical 'grounding' that allows the model to more effectively absorb and contextualize other data sources, like human videos, later on.
The main risk to humanoid robotics adoption isn't a lack of impressive capabilities, but the failure to achieve near-perfect reliability. Unlike an LLM where a human can simply re-prompt, a robot's value is in its autonomy. Bridging the gap from 95% to 100% reliability for autonomous tasks is the critical, unsolved challenge.
Adopting the true 'foundation model ethos' from LLMs is difficult for roboticists. It means a warehouse automation company should collect data from kitchen robots. This breadth, while seemingly unrelated, builds a generalist model that better handles the weird edge cases in the target domain than a narrowly trained specialist model.
The rise of Chinese robotics highlights a key U.S. weakness: a fragmented ecosystem. Instead of just focusing on ML models, the U.S. needs to reinvest in the entire stack—including domestic hardware manufacturing, supply chains, and R&D—to create the healthy, holistic environment necessary for leadership.
Robots can generalize skills to new hardware by using an intermediate 'thinking' step. The model generates an image of the next desired milestone (e.g., a half-folded shirt). This visual goal is embodiment-agnostic and easier to create than new motor commands, allowing the robot to then solve for the actions to match the image.
The hardest problem in robotics is generalization—performing a task reliably with new objects in new environments. However, a single demonstration of this looks far less impressive than a highly rehearsed, acrobatic feat. The true technical achievement is only visible across many trials, making it hard to appreciate from a typical demo video.
