Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

A contrarian view on "destructive scanning" of old books for AI training. Instead of a loss, it's a form of preservation. This process transfers knowledge from a format no longer widely consumed (physical books) into a dynamic, interactive system (AI), making the information more accessible to the world.

Related Insights

A new market has emerged where defunct startups sell their entire operational histories—including codebases, internal communications, and go-to-market data—to AI labs and data brokers. This creates a new form of salvage value, turning years of failed effort into a valuable corpus for training next-generation models.

The static PDF is an inefficient medium for knowledge transfer. The future may be interactive AI models that hold the research, allowing users to dynamically query, expand, and explore concepts, making science more accessible and breaking the compress/decompress cycle of papers.

Avoid processing raw data into summaries and then deleting the source. AI technology improves so rapidly that you'll want to re-process the original, raw data with future, more capable models to generate superior outputs and system upgrades, preventing irreversible information loss.

Anthropic's 'Project Panama' involved buying and physically destroying one million books to scan their pages. This expensive, analog process created a unique, high-quality dataset that is difficult for competitors relying on easily accessible digital libraries to replicate, forming a powerful and defensible competitive advantage in the AI race.

A massive opportunity for AI lies in unearthing and recording experts' tacit, unwritten knowledge—the "knack" for doing things that is lost when they die. This "dark data," once fed into models, will unlock immense, currently inaccessible value.

Historically, the value of content IP like scripts and music declined sharply 30-60 days after release. AI tools can now "reimagine" these dormant libraries quickly and cost-effectively, creating new derivative works. This presents a massive, previously untapped opportunity to unlock new revenue streams from back catalogs.

Contrary to the goal of perfect data retention, 'machine unlearning' is becoming a critical capability. The ability for an AI to forget is essential for privacy (removing user data), correcting biases from flawed training data, and adapting to new information, mirroring a core, beneficial aspect of human cognition.

The success of AI is creating a long-term data scarcity problem. By obviating the need for human-curated knowledge platforms like Stack Overflow, AI is eliminating the very sources of high-quality, structured data required for training future models. This creates a self-defeating cycle where AI's utility today undermines its improvement tomorrow.

A significant business opportunity exists in creating a service that properly parses, manages, and provides API access to the vast library of out-of-copyright books. This would allow AI agents and tools to easily ingest and utilize classic literature and knowledge, creating a new content distribution layer.

Anthropic maintains a competitive edge by physically acquiring and digitizing thousands of old books, creating a massive, proprietary dataset of high-quality text. This multi-year effort to build a unique data library is difficult to replicate and may contribute to the distinct quality of its Claude models.