Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Schiff explains that frontier AI companies chose to vacuum up protected creative intellectual property without licensing agreements because competitive pressure drove them to prioritize speed over legality. Rather than establishing legal rights beforehand, developers opted to assume massive litigation risk later, leaving Congress to resolve retroactively what training data must be retained and disclosed to copyright holders.

Related Insights

Major AI labs are protesting that Chinese companies are "stealing" their models via distillation. However, these same labs built their foundational models by training on vast amounts of copyrighted material without permission, a practice the host calls "IP theft," undermining their public standing on the issue.

Unlike Google and Meta who own vast video libraries, OpenAI lacked training data for Sora. Their solution was a legally aggressive "opt-out" policy for copyrighted material, effectively shifting the burden to IP holders and turning IP licensing, not just data access, into the next competitive frontier.

There is a profound hypocrisy in the AI industry's stance on intellectual property. Companies that built their foundational models by scraping the entire internet are now seeking regulatory protection to prevent others from distilling or learning from their models—mirroring how the music industry fought Napster after profiting from an open ecosystem.

The administration's policy document expresses its belief that training AI on copyrighted material is not a violation. However, rather than proposing legislation, it advocates for allowing the judiciary to resolve the contentious "fair use" issue, effectively punting the decision to the courts and avoiding a difficult political battle.

Unsealed documents from the New York Times lawsuit reveal internal communications where key employees acknowledged training models on copyrighted data was ethically and legally dubious. One Microsoft director even questioned if it could be called "fair use," undermining their public legal defense.

US copyright law's "fair use" doctrine, which allows AI models to be trained on vast datasets of copyrighted material, is a key competitive advantage. This legal framework, an artifact of American law, enables more rapid and powerful LLM development compared to countries with more restrictive copyright regimes.

The market reality is that consumers and businesses prioritize the best-performing AI models, regardless of whether their training data was ethically sourced. This dynamic incentivizes labs to use all available data, including copyrighted works, and treat potential fines as a cost of doing business.

While US AI companies navigate complex licensing deals with IP holders, Chinese firms like ByteDance appear to be using copyrighted material, such as specific actors' voices, without restriction. This lack of legal friction allows them to generate highly specific and realistic content that Western labs are hesitant to produce.

Companies like OpenAI knowingly use copyrighted material, calculating that the market cap gained from rapid growth will far exceed the eventual legal settlements. This strategy prioritizes building a dominant market position by breaking the law, viewing fines as a cost of doing business.

While an AI model itself may not be an infringement, its output could be. If you use AI-generated content for your business, you could face lawsuits from creators whose copyrighted material was used for training. The legal argument is that your output is a "derivative work" of their original, protected content.

AI Developers Knowingly Risked Copyright Infringement to Outpace Competitors in Model Training | RiffOn