Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Tech giants like Google are now bidding on the data of bankrupt companies, such as Spirit Airlines, not for their physical assets but for their unique, real-world datasets. This data is highly valuable for training AI models, creating a new and unexpected market for corporate information that was previously considered dormant.

Related Insights

A new market has emerged where defunct startups sell their entire operational histories—including codebases, internal communications, and go-to-market data—to AI labs and data brokers. This creates a new form of salvage value, turning years of failed effort into a valuable corpus for training next-generation models.

Unlike debt-laden startups, tech giants are funding AI buildouts with cash and can weather a downturn. They fully expect smaller, leveraged competitors to go bankrupt, creating a strategic opportunity to purchase their data center assets for pennies on the dollar, thereby reducing their own future capital expenditures.

Public internet data has been largely exhausted for training AI models. The real competitive advantage and source for next-generation, specialized AI will be the vast, untapped reservoirs of proprietary data locked inside corporations, like R&D data from pharmaceutical or semiconductor companies.

Cuban identifies a massive, overlooked opportunity: acquiring the intellectual property (patents, data, designs) from millions of defunct businesses. This "dead IP" could be aggregated and sold at a high premium to foundational model companies desperate for unique training data.

With public data exhausted, AI companies are seeking proprietary datasets. After being rejected by established firms wary of sharing their 'crown jewels,' these labs are now acquiring the codebases of failed startups for tens of thousands of dollars as a novel source of high-quality training data.

Google outbid competitors for Spirit Airlines' internal data—not for customer info, but for mundane emails, Slacks, and meeting transcripts. This signals a new phase in AI training where unstructured corporate communications are highly valued for teaching AI agents how organizations actually function and communicate.

As AI commoditizes software creation, the primary source of sustainable value shifts from the software itself to the unique, high-quality data that AI agents use for decision-making. Businesses must re-center their strategy around data as the core asset.

In the current M&A landscape, data-centric startups are more valuable than application-layer companies. Acquirers, particularly large tech firms, need proprietary data sets to train, run, and customize their AI models. This demand makes companies with unique data assets highly attractive takeover targets, with some seeing a tenfold increase in inquiries.

The rumored acquisition of Pinterest by OpenAI is driven by its 200 billion user-tagged images, a 'goldmine' for AI training. This demonstrates that large, well-structured datasets are becoming critical strategic assets and key drivers for M&A activity in the AI sector.

The era of building frontier AI models on easily scraped internet data is ending. The next competitive advantage lies in securing unique, proprietary, real-world datasets that reflect complex physical interactions, such as endoscopy videos or 3D object data. Synthetic data is proving insufficient, making access to this "reality" data the key differentiator.