We scan new podcasts and send you the top 5 insights daily.
Huge amounts of valuable information are legally orphaned, such as patents from the Soviet Union. This 'stateless data,' along with other undigitized historical artifacts, represents a massive untapped resource for training AI and uncovering knowledge, posing unique challenges for digitization, ownership, and access.
The industry has already exhausted the public web data used to train foundational AI models, a point underscored by the phrase "we've already run out of data." The next leap in AI capability and business value will come from harnessing the vast, proprietary data currently locked behind corporate firewalls.
A new market has emerged where defunct startups sell their entire operational histories—including codebases, internal communications, and go-to-market data—to AI labs and data brokers. This creates a new form of salvage value, turning years of failed effort into a valuable corpus for training next-generation models.
Many businesses overlook their most valuable existing data assets. Years of unstructured data, like support tickets detailing integration issues and customer problems, are invaluable for training specialized AI models. This 'boring' data can become a key source of competitive advantage when activated with an LLM.
Public internet data has been largely exhausted for training AI models. The real competitive advantage and source for next-generation, specialized AI will be the vast, untapped reservoirs of proprietary data locked inside corporations, like R&D data from pharmaceutical or semiconductor companies.
Advanced reactor concepts proven decades ago were not pursued because the foundational research was stored in non-digitized, paper archives with limited distribution. This information asymmetry created a major barrier for new companies seeking to commercialize the technology.
While aggregating compute is a known challenge for open source AI, the more critical, less-discussed problem is aggregating data. Closed-source labs spend billions creating complex reinforcement learning (RL) environments and proprietary datasets. Without a concerted, non-commercial effort to create and open-source these data assets, open source models risk falling behind permanently.
Contrary to the belief that AI requires perfect, clean data, the biggest opportunity lies in building technology that can find signals in messy, diverse data sets across different modalities and organisms. The tech should solve the data problem, not wait for it to be solved.
Cuban identifies a massive, overlooked opportunity: acquiring the intellectual property (patents, data, designs) from millions of defunct businesses. This "dead IP" could be aggregated and sold at a high premium to foundational model companies desperate for unique training data.
A massive opportunity for AI lies in unearthing and recording experts' tacit, unwritten knowledge—the "knack" for doing things that is lost when they die. This "dark data," once fed into models, will unlock immense, currently inaccessible value.
Tech giants like Google are now bidding on the data of bankrupt companies, such as Spirit Airlines, not for their physical assets but for their unique, real-world datasets. This data is highly valuable for training AI models, creating a new and unexpected market for corporate information that was previously considered dormant.