Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

While large language models are trained on scraped internet data, a significant advantage lies in the physical world. Startups that digitize unique, offline knowledge—like the expertise of a retiring machinist—can build defensible moats that larger platforms can't easily replicate.

Related Insights

DoorDash is creating a unique data moat by digitizing physical-world information unavailable on the internet, like hyper-local parking data or real-time store inventory. This proprietary dataset, which LLMs cannot currently access, becomes a key strategic asset for building specialized AI models.

Unlike digital AI trained on public internet data, physical AI models require vast, private datasets collected from real-world operations like mines. The ability to collect this proprietary data, often in restricted locations with government approval, creates a powerful and defensible competitive advantage.

Generic tech companies can't easily dominate industrial AI. Training models requires proprietary operational data that isn't public, creating "data friction." Furthermore, solving problems in a refinery versus a hospital requires deep, sector-specific domain knowledge, preventing a one-size-fits-all approach.

OpenAI believes it has sufficient coding data. The next data advantage lies in capturing "knowledge work" tasks—data not on the public internet. This may require novel approaches like acquiring failed startups for their internal data from tools like Slack.

Since LLMs are commodities, sustainable competitive advantage in AI comes from leveraging proprietary data and unique business processes that competitors cannot replicate. Companies must focus on building AI that understands their specific "secret sauce."

Since Large Language Models are trained on public internet data, their answers become commoditized. Cultivate a private network of narrow-topic experts you can text for unique insights. This creates an informational advantage that AI cannot currently replicate.

As AI application layers become easier to clone, the sustainable competitive advantage is moving down the tech stack. Companies with unique, last-mile user interaction data can build proprietary models that are cheaper and better, creating a data flywheel and a moat that is difficult for competitors to replicate.

As AI models become commoditized, the ultimate defensibility comes from exclusive access to a unique dataset. A startup with a slightly inferior model but a comprehensive, proprietary dataset (e.g., all legal records) will beat a superior, general-purpose model for specialized tasks, creating a powerful long-term advantage.

Companies create defensibility by generating unique, non-public data through their operations (e.g., legal case outcomes). This proprietary data improves their own models, creating a feedback loop and a compounding advantage that large, generalist labs like OpenAI cannot replicate.

While general models are powerful, true competitive advantage will come from hyper-specialized AI. This requires training models on vast amounts of proprietary data stored within a company or on a factory floor, creating a moat that general models cannot replicate.