Humanoid running has progressed rapidly because locomotion on flat, rigid surfaces is straightforward to model in simulation and transfer to the real world. In contrast, manipulation requires modeling complex contact dynamics, friction, and deformable objects like cloth. Where physics is harder to model, current simulations break down, leaving manipulation reliant on slower, harder real-world data collection.
Building narrow AI creates high performance on one or two tasks, but the nth task costs as much to train as the first. In contrast, building generalist baselines endows models with broad common sense and physical reasoning. Once this broad base is established, tuning a robot to high reliability and mastery on any arbitrary task becomes significantly faster and cheaper.
Language models execute identically regardless of device, but robotics software is constrained by the physical robot body. Keerthana Gopalakrishnan places current robotics in the 'GPT-2 era' because generic robotic brains cannot reliably control arbitrary, unseen robot embodiments out of the box, nor can they reliably generalize from few-shot demonstrations across hard, diverse tasks.
Separating robotic autonomy into a high-level reasoning model (like Gemini Robotics ER) and a low-level vision-language-action (VLA) policy causes practical bottlenecks. Beyond the communication and processing latency of reasoning models, sequential multi-step tasks suffer from compounded errors where the success rates of both models multiply across every handoff, creating compounding failure points during execution.
A standard 128k context window holds roughly three minutes of dense multimodal visual history. To maintain longer operational horizons, roboticists must adopt context engineering: converting older visual memories into compressed textual narratives and summarizations of past actions while retaining raw visual tokens only for immediate, high-precision actions.
Within roughly 15 months between Gemini Robotics releases, multi-fingered hardware and control policies matured enough to overtake parallel grippers. Multi-fingered hands can replicate every task that parallel grippers performed at the frontier while unlocking complex dexterous actions like tying trash bags, rendering grippers legacy technology for frontier dexterity research.
When a gripper robot makes a mistake, humans tolerate it because it looks mechanical. But because humanoids look human, users instinctively expect human-level common sense and competence. Consequently, humanoid manipulation errors or clumsy movements are judged far more harshly, raising the baseline reliability bar that humanoids must satisfy before public deployment.
Unlike digital AI safety which emphasizes guardrails against bad actors and deceptive behavior, the core safety challenge in physical robotics is basic operational competence. A humanoid robot falling over or misjudging sensor inputs poses serious physical danger to humans entirely through accidental clumsiness rather than malevolent intent.
No single data collection method will solve robotics. Teleoperation provides high-precision ground truth for specific robots but scales poorly and degrades when hardware updates. Universal Manipulation Interface (UMI) data offers better scale and precision via sensors but requires hardware maintenance. Egocentric human video provides massive scale but lacks precise end-effector force signals and suffers from human-to-robot physical embodiment mismatch.
Pure visual feedback is insufficient for delicate real-world manipulation tasks such as stacking brittle items or assembling components. To avoid crushing materials, robots require real-time end-effector force sensing and physical back-drivability/compliance. Having explicit torque and pressure signals allows neural policies to make safe adjustments that vision alone cannot deduce.
