We scan new podcasts and send you the top 5 insights daily.
Unlike vision models that leverage vast internet corpora of image-text pairs, foundation models for sensor data (radar, vibration) face a major hurdle: a near-total absence of publicly available, paired sensor-language data. This forces researchers to develop entirely new techniques for sensor-language alignment, distinct from standard VLM methods.
The rapid progress of many LLMs was possible because they could leverage the same massive public dataset: the internet. In robotics, no such public corpus of robot interaction data exists. This “data void” means progress is tied to a company's ability to generate its own proprietary data.
Large Language Models are limited because they lack an understanding of the physical world. The next evolution is 'World Models'—AI trained on real-world sensory data to understand physics, space, and context. This is the foundational technology required to unlock physical AI like advanced robotics.
The push toward physical AI and spatial intelligence is primarily a strategy to overcome data scarcity for training general models. By creating simulated 3D environments, researchers can generate the novel, complex data that is currently unavailable but crucial for advancing AI into the real world.
Ken Goldberg quantifies the challenge: the text data used to train LLMs would take a human 100,000 years to read. Equivalent data for robot manipulation (vision-to-control signals) doesn't exist online and must be generated from scratch, explaining the slower progress in physical AI.
Neurobotics posits that true physical AI requires more than just vision-language models; it needs a "nervous system" and reflexes. They advocate for training robots in physical "gyms" to collect embodied data, arguing that complex physical tasks cannot be learned solely by watching videos.
AI can generate art because it was trained on the internet's vast trove of images. It struggles with physical tasks like washing dishes because there is virtually no first-person video data for such actions. Solving this data-gathering problem is key to advancing robotics.
According to Agility Robotics' co-founder, perception is now a largely solved problem. The new frontier is generating training data for robot control—the specific torque commands and sensor inputs for actions. Unlike text or images for LLMs, this data does not exist on the internet and must be painstakingly created.
Current multimodal models shoehorn visual data into a 1D text-based sequence. True spatial intelligence is different. It requires a native 3D/4D representation to understand a world governed by physics, not just human-generated language. This is a foundational architectural shift, not an extension of LLMs.
Unlike LLMs trained on vast digital text, humanoid robots need immense amounts of real-world physical data to learn simple tasks. It's estimated that 100 million hours—over 11,000 years' worth—of interaction data is needed to create truly smart, useful humanoids, highlighting the scale of the challenge.
The "bitter lesson" (scale and simple models win) works for language because training data (text) aligns with the output (text). Robotics faces a critical misalignment: it's trained on passive web videos but needs to output physical actions in a 3D world. This data gap is a fundamental hurdle that pure scaling cannot solve.