Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Unsealed documents from the New York Times lawsuit reveal internal communications where key employees acknowledged training models on copyrighted data was ethically and legally dubious. One Microsoft director even questioned if it could be called "fair use," undermining their public legal defense.

Related Insights

Major AI labs are protesting that Chinese companies are "stealing" their models via distillation. However, these same labs built their foundational models by training on vast amounts of copyrighted material without permission, a practice the host calls "IP theft," undermining their public standing on the issue.

The AI industry was largely built on scraping data without permission. Apple's lawsuit frames OpenAI's alleged theft of hardware secrets as part of this same culture. This narrative makes OpenAI culturally vulnerable in a legal battle, as it appears to be a pattern of behavior rather than an isolated incident.

Proprietary labs argue against 'distillation' (using their model outputs for training) while they have built their own models on vast amounts of copyrighted data. This opposition is an anti-competitive tactic, as model outputs are not copyrightable and distillation helps smaller, open players to compete.

There is a profound hypocrisy in the AI industry's stance on intellectual property. Companies that built their foundational models by scraping the entire internet are now seeking regulatory protection to prevent others from distilling or learning from their models—mirroring how the music industry fought Napster after profiting from an open ecosystem.

Satya Nadella argues that when enterprises use third-party AI, they give away valuable proprietary knowledge through their prompts and data. This "Reverse Information Paradox" means companies pay twice: once with money, and again by training the vendor's model with their core intellectual property.

US AI labs' efforts to prevent foreign rivals from distilling their models face accusations of hypocrisy. Critics point out that these labs train their own models on vast amounts of public data without permission. This "pot calling the kettle black" dynamic complicates legal and ethical arguments against industrial-scale distillation.

AI companies protest when competitors "distill" their models, calling it a violation. This stance is deeply ironic, as it mirrors the complaints of artists and creators whose work was scraped without permission to build the original models. The industry fails to acknowledge this double standard.

In the landmark NYT v. OpenAI copyright case, the DOJ filed a statement supporting the idea that training models on copyrighted text is fair use. This position prioritizes U.S. innovation and competition with foreign rivals over creators' rights.

Companies like OpenAI knowingly use copyrighted material, calculating that the market cap gained from rapid growth will far exceed the eventual legal settlements. This strategy prioritizes building a dominant market position by breaking the law, viewing fines as a cost of doing business.

Anthropic's argument that Chinese AI model distillation is 'IP theft' is a potentially fatal legal mistake. This assertion can be used against them in lawsuits from content creators like the New York Times, as Anthropic's own models are built by 'distilling' public content, effectively confessing their product is based on stolen IP.