We scan new podcasts and send you the top 5 insights daily.
It's a tenuous and difficult argument to suggest that training an AI on the public internet is acceptable, but training an AI on the output of another AI is not. Aaron Levie notes the logical inconsistency, suggesting the ethical line is blurry at best, especially when the original model provider is paid for the API usage during distillation.
As more of the public internet and code repositories are generated by LLMs, any new model trained on this public data is, in effect, being 'distilled' from other models. This complicates accusations of direct distillation and blurs the line for what constitutes original training data.
Major AI labs are protesting that Chinese companies are "stealing" their models via distillation. However, these same labs built their foundational models by training on vast amounts of copyrighted material without permission, a practice the host calls "IP theft," undermining their public standing on the issue.
Anthropic operates on a double standard. It argues for the 'fair use' right to ingest and learn from all the world's copyrighted material for free. However, it deems it IP theft for anyone to train on its own models' outputs—which courts have ruled are not even copyrightable in the first place.
Proprietary labs argue against 'distillation' (using their model outputs for training) while they have built their own models on vast amounts of copyrighted data. This opposition is an anti-competitive tactic, as model outputs are not copyrightable and distillation helps smaller, open players to compete.
The practice of "smart distillation"—using a frontier model to guide and train a smaller model—operates in a legal and ethical gray area. It is more sophisticated than simple copying ("dumb distillation") and resembles how enterprises fine-tune models, complicating narratives about IP theft in AI development.
There is a profound hypocrisy in the AI industry's stance on intellectual property. Companies that built their foundational models by scraping the entire internet are now seeking regulatory protection to prevent others from distilling or learning from their models—mirroring how the music industry fought Napster after profiting from an open ecosystem.
As more of the internet and code repositories are generated by leading AI models, any new model trained on this public data inadvertently "distills" the knowledge and quirks of those proprietary systems. This blurs the line between original training and outright copying.
The debate over distilling from other AI models is becoming moot. The internet is now so saturated with AI-generated content ('AI slop') that any new model trained on web data is already, by default, being trained on the outputs of its predecessors. Pure 'human data' is a dwindling resource.
US AI labs' efforts to prevent foreign rivals from distilling their models face accusations of hypocrisy. Critics point out that these labs train their own models on vast amounts of public data without permission. This "pot calling the kettle black" dynamic complicates legal and ethical arguments against industrial-scale distillation.
AI companies protest when competitors "distill" their models, calling it a violation. This stance is deeply ironic, as it mirrors the complaints of artists and creators whose work was scraped without permission to build the original models. The industry fails to acknowledge this double standard.