Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Google outbid competitors for Spirit Airlines' internal data—not for customer info, but for mundane emails, Slacks, and meeting transcripts. This signals a new phase in AI training where unstructured corporate communications are highly valued for teaching AI agents how organizations actually function and communicate.

Related Insights

The industry has already exhausted the public web data used to train foundational AI models, a point underscored by the phrase "we've already run out of data." The next leap in AI capability and business value will come from harnessing the vast, proprietary data currently locked behind corporate firewalls.

A new market has emerged where defunct startups sell their entire operational histories—including codebases, internal communications, and go-to-market data—to AI labs and data brokers. This creates a new form of salvage value, turning years of failed effort into a valuable corpus for training next-generation models.

The vast majority of enterprise information, previously trapped in formats like PDFs and documents, was largely unusable. AI, through techniques like RAG and automated structure extraction, is unlocking this data for the first time, making it queryable and enabling new large-scale analysis.

Meta's controversial keystroke logging is a data collection effort to capture the full context of white-collar work. The goal is to train AI on the reasoning, trade-offs, and discussions that lead to a final product—a much richer signal for agentic AI than the final code or document alone.

Public internet data has been largely exhausted for training AI models. The real competitive advantage and source for next-generation, specialized AI will be the vast, untapped reservoirs of proprietary data locked inside corporations, like R&D data from pharmaceutical or semiconductor companies.

OpenAI believes it has sufficient coding data. The next data advantage lies in capturing "knowledge work" tasks—data not on the public internet. This may require novel approaches like acquiring failed startups for their internal data from tools like Slack.

The most valuable data for training enterprise AI is not a company's internal documents, but a recording of the actual work processes people use to create them. The ideal training scenario is for an AI to act like an intern, learning directly from human colleagues, which is far more informative than static knowledge bases.

As algorithms become more widespread, the key differentiator for leading AI labs is their exclusive access to vast, private data sets. XAI has Twitter, Google has YouTube, and OpenAI has user conversations, creating unique training advantages that are nearly impossible for others to replicate.

For an AI agent to be effective, "context" isn't just data access. It's understanding an organization's fluid, internal shorthand—definitions, acronyms, and unwritten rules like "top spenders in EMEA." This evolving knowledge is often buried in emails and meeting transcripts, not formal documents.

The initial AI boom was fueled by scraping the public internet. Cuban predicts the next phase will be dominated by exclusive data deals. Content owners, like medical journals, will protect their IP and auction it to the highest-bidding AI companies, creating valuable data silos.