Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Anthropic operates on a double standard. It argues for the 'fair use' right to ingest and learn from all the world's copyrighted material for free. However, it deems it IP theft for anyone to train on its own models' outputs—which courts have ruled are not even copyrightable in the first place.

Related Insights

Major AI labs are protesting that Chinese companies are "stealing" their models via distillation. However, these same labs built their foundational models by training on vast amounts of copyrighted material without permission, a practice the host calls "IP theft," undermining their public standing on the issue.

Proprietary labs argue against 'distillation' (using their model outputs for training) while they have built their own models on vast amounts of copyrighted data. This opposition is an anti-competitive tactic, as model outputs are not copyrightable and distillation helps smaller, open players to compete.

The practice of "smart distillation"—using a frontier model to guide and train a smaller model—operates in a legal and ethical gray area. It is more sophisticated than simple copying ("dumb distillation") and resembles how enterprises fine-tune models, complicating narratives about IP theft in AI development.

There is a profound hypocrisy in the AI industry's stance on intellectual property. Companies that built their foundational models by scraping the entire internet are now seeking regulatory protection to prevent others from distilling or learning from their models—mirroring how the music industry fought Napster after profiting from an open ecosystem.

The controversial practice of AI 'distillation' is not IP theft but a modern form of competitive benchmarking. It's akin to how early Google submitted queries to Yahoo to compare and improve its own search results. The focus is on learning from a competitor's public output, not stealing their underlying software or code.

US AI labs' efforts to prevent foreign rivals from distilling their models face accusations of hypocrisy. Critics point out that these labs train their own models on vast amounts of public data without permission. This "pot calling the kettle black" dynamic complicates legal and ethical arguments against industrial-scale distillation.

If a company like Meta uses Anthropic's AI to rewrite its codebase, it creates a legally ambiguous dataset. While enterprise contracts typically prevent labs from training on customer data, the reverse is also likely restricted, raising questions about whether the customer can train its own future models on this AI-augmented corpus.

AI companies protest when competitors "distill" their models, calling it a violation. This stance is deeply ironic, as it mirrors the complaints of artists and creators whose work was scraped without permission to build the original models. The industry fails to acknowledge this double standard.

US copyright law's "fair use" doctrine, which allows AI models to be trained on vast datasets of copyrighted material, is a key competitive advantage. This legal framework, an artifact of American law, enables more rapid and powerful LLM development compared to countries with more restrictive copyright regimes.

Anthropic's argument that Chinese AI model distillation is 'IP theft' is a potentially fatal legal mistake. This assertion can be used against them in lawsuits from content creators like the New York Times, as Anthropic's own models are built by 'distilling' public content, effectively confessing their product is based on stolen IP.