Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Unlike the drone or self-driving car industries, which benefit from vast open-source datasets and research, the maritime sector is a 'data desert.' This is due to a lack of commercial adoption and the immediate classification of advanced subsea data, forcing pioneers to build from scratch.

Related Insights

Unlike LLMs that train on the existing internet, robotics lacks a pre-training dataset for the physical world. This forces companies like Encore to build a full-stack solution combining a software platform for data management with human-led operations for data collection, annotation, and even real-time remote robot piloting for exception handling.

The rapid progress of many LLMs was possible because they could leverage the same massive public dataset: the internet. In robotics, no such public corpus of robot interaction data exists. This “data void” means progress is tied to a company's ability to generate its own proprietary data.

While large language models (LLMs) converge by training on the same public internet data, autonomous driving models will remain distinct. Each company must build its own proprietary dataset from its unique sensor stack and vehicle fleet. This lack of a shared data foundation means different automakers' AI driving behaviors and capabilities will likely diverge over time.

While aggregating compute is a known challenge for open source AI, the more critical, less-discussed problem is aggregating data. Closed-source labs spend billions creating complex reinforcement learning (RL) environments and proprietary datasets. Without a concerted, non-commercial effort to create and open-source these data assets, open source models risk falling behind permanently.

Unlike consumer AI trained on public internet data, industrial AI requires vast, proprietary datasets from the physical world (e.g., sensor readings from a submarine hull). Gecko Robotics is building this data corpus via its robots, creating an advantage that's difficult to replicate.

Generic tech companies can't easily dominate industrial AI. Training models requires proprietary operational data that isn't public, creating "data friction." Furthermore, solving problems in a refinery versus a hospital requires deep, sector-specific domain knowledge, preventing a one-size-fits-all approach.

The future of valuable AI lies not in models trained on the abundant public internet, but in those built on scarce, proprietary data. For fields like robotics and biology, this data doesn't exist to be scraped; it must be actively created, making the data generation process itself the key competitive moat.

According to Agility Robotics' co-founder, perception is now a largely solved problem. The new frontier is generating training data for robot control—the specific torque commands and sensor inputs for actions. Unlike text or images for LLMs, this data does not exist on the internet and must be painstakingly created.

Unlike radar, which operates in the consistent medium of air, sonar performance is heavily distorted by local oceanic conditions. This means AI models trained on sonar data from one location cannot be reliably transferred to another, necessitating low-cost, scalable hardware for real-time data collection.

For physical AI, the primary constraint is not the cost of data but its fundamental non-existence. Unlike software AI, you can't advance without deploying robots "in the wild" to capture edge cases—a classic chicken-and-egg problem that simulation alone cannot solve and capital cannot easily buy.