We scan new podcasts and send you the top 5 insights daily.
No single data collection method will solve robotics. Teleoperation provides high-precision ground truth for specific robots but scales poorly and degrades when hardware updates. Universal Manipulation Interface (UMI) data offers better scale and precision via sensors but requires hardware maintenance. Egocentric human video provides massive scale but lacks precise end-effector force signals and suffers from human-to-robot physical embodiment mismatch.
Unlike LLMs that train on the existing internet, robotics lacks a pre-training dataset for the physical world. This forces companies like Encore to build a full-stack solution combining a software platform for data management with human-led operations for data collection, annotation, and even real-time remote robot piloting for exception handling.
To build generalist robots, the most effective approach is pre-training foundation models on internet-scale video datasets, not just simulation or tele-operated data. This vast, diverse data provides a deep, implicit understanding of physics and object interaction that is impossible to replicate in controlled environments, enabling true generalization.
To overcome the data bottleneck in robotics, Sunday developed gloves that capture human hand movements. This allows them to train their robot's manipulation skills without needing a physical robot for teleoperation. By separating data gathering (gloves) from execution (robot), they can scale their training dataset far more efficiently than competitors who rely on robot-in-the-loop data collection methods.
Robotics company OneX designs its robot hands to be biomechanically identical to human hands not for aesthetics, but for data transfer. This allows them to train models on vast amounts of existing human video, which then 'just works' on the robot, bypassing the need for extensive simulation or teleoperation data.
Physical Intelligence demonstrated an emergent capability where its robotics model, after reaching a certain performance threshold, significantly improved by training on egocentric human video. This solves a major bottleneck by leveraging vast, existing video datasets instead of expensive, limited teleoperated data.
Runway’s robotics thesis is that pre-training on massive, easily available third-person video data (e.g., people performing tasks) is more scalable and effective than relying on expensive, limited teleoperation or first-person data. This general world knowledge can then be fine-tuned for specific robotic tasks.
Simply collecting more data from a deployed robot isn't enough to create a powerful learning flywheel. If the tasks are repetitive (e.g., a million car welds), the model won't generalize. Data must be diverse, acting more like an 'education program' than a fungible commodity to drive real capability growth.
According to Agility Robotics' co-founder, perception is now a largely solved problem. The new frontier is generating training data for robot control—the specific torque commands and sensor inputs for actions. Unlike text or images for LLMs, this data does not exist on the internet and must be painstakingly created.
Counterintuitively, the best way to train a robot foundation model isn't to start with vast human video datasets. Research indicates that starting with real, embodied robot data provides a physical 'grounding' that allows the model to more effectively absorb and contextualize other data sources, like human videos, later on.
The "bitter lesson" (scale and simple models win) works for language because training data (text) aligns with the output (text). Robotics faces a critical misalignment: it's trained on passive web videos but needs to output physical actions in a 3D world. This data gap is a fundamental hurdle that pure scaling cannot solve.