Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The race for superior AI models has moved beyond the public internet. Companies like Micro One are offering huge sums (e.g., $800,000) for access to private corporate data from platforms like Slack, Notion, and Gmail, creating a new and lucrative market for proprietary training data.

Related Insights

The industry has already exhausted the public web data used to train foundational AI models, a point underscored by the phrase "we've already run out of data." The next leap in AI capability and business value will come from harnessing the vast, proprietary data currently locked behind corporate firewalls.

A new market has emerged where defunct startups sell their entire operational histories—including codebases, internal communications, and go-to-market data—to AI labs and data brokers. This creates a new form of salvage value, turning years of failed effort into a valuable corpus for training next-generation models.

LLMs have hit a wall by scraping nearly all available public data. The next phase of AI development and competitive differentiation will come from training models on high-quality, proprietary data generated by human experts. This creates a booming "data as a service" industry for companies like Micro One that recruit and manage these experts.

Public internet data has been largely exhausted for training AI models. The real competitive advantage and source for next-generation, specialized AI will be the vast, untapped reservoirs of proprietary data locked inside corporations, like R&D data from pharmaceutical or semiconductor companies.

In an era of commoditized LLMs, the real competitive advantage lies in unique, proprietary datasets. These datasets, when combined with AI models, create a defensible moat that software alone cannot replicate. This is why major tech companies are aggressively acquiring data-rich companies.

With public data exhausted, AI companies are seeking proprietary datasets. After being rejected by established firms wary of sharing their 'crown jewels,' these labs are now acquiring the codebases of failed startups for tens of thousands of dollars as a novel source of high-quality training data.

Google outbid competitors for Spirit Airlines' internal data—not for customer info, but for mundane emails, Slacks, and meeting transcripts. This signals a new phase in AI training where unstructured corporate communications are highly valued for teaching AI agents how organizations actually function and communicate.

As algorithms become more widespread, the key differentiator for leading AI labs is their exclusive access to vast, private data sets. XAI has Twitter, Google has YouTube, and OpenAI has user conversations, creating unique training advantages that are nearly impossible for others to replicate.

The initial AI boom was fueled by scraping the public internet. Cuban predicts the next phase will be dominated by exclusive data deals. Content owners, like medical journals, will protect their IP and auction it to the highest-bidding AI companies, creating valuable data silos.

The era of building frontier AI models on easily scraped internet data is ending. The next competitive advantage lies in securing unique, proprietary, real-world datasets that reflect complex physical interactions, such as endoscopy videos or 3D object data. Synthetic data is proving insufficient, making access to this "reality" data the key differentiator.