We scan new podcasts and send you the top 5 insights daily.
Accusations that Chinese labs cheat by copying US models are misleading. The practice, known as distillation, is common across the industry (including by Elon Musk's xAI) and academia. Now, with Chinese labs dominating open source, American startups are increasingly building on top of Chinese models.
When a company distills knowledge from a competitor's AI, it's not just scraping pre-training data. It's a highly efficient process of extracting the model's intelligence, reasoning patterns, and skills. This is more akin to an apprentice directly interacting with and learning from a world-class expert than simply reading the same textbooks the expert used.
Despite impressive models from companies like DeepSeek, China's AI ecosystem is heavily reliant on "distilling"—essentially copying and refining—open-source models from the US. This dependency on an external innovation engine is a major weakness in their national strategy to achieve genuine AI leadership and self-sufficiency.
China is gaining an efficiency edge in AI by using "distillation"—training smaller, cheaper models from larger ones. This "train the trainer" approach is much faster and challenges the capital-intensive US strategy, highlighting how inefficient and "bloated" current Western foundational models are.
Challenging the narrative of pure technological competition, Jensen Huang points out that American AI labs and startups significantly benefited from Chinese open-source contributions like the DeepSeek model. This highlights the global, interconnected nature of AI research, where progress in one nation directly aids others.
China is rapidly closing the AI gap not through pure innovation but through "distillation"—systematically querying Western frontier models via their APIs to harvest their reasoning processes. This allows them to train their own models to a near-frontier level at a fraction of the cost, bypassing years of foundational research.
Hugging Face's CEO dismisses the controversy around distillation, framing it as a widespread technique that offers only a marginal boost. It doesn't determine a model's fundamental quality—'if you suck, you suck with or without distillation'—and he questions the merit of 'unfair competition' claims from dominant, trillion-dollar companies.
Chinese labs use 'smart distillation,' a sophisticated technique where a frontier model acts as a 'teacher' to guide a smaller model's judgment and data labeling. This is viewed as a legitimate and efficient catch-up method, distinct from simply copy-pasting answers.
Leading Chinese AI models like Kimi appear to be primarily trained on the outputs of US models (a process called distillation) rather than being built from scratch. This suggests China's progress is constrained by its ability to scrape and fine-tune American APIs, indicating the U.S. still holds a significant architectural and innovation advantage in foundational AI.
Unable to build frontier models from scratch, some Chinese companies gain a competitive edge by using "scale distillation." This involves training smaller, open models on the outputs of larger, proprietary US models, effectively piggybacking on American R&D to create capable, low-cost alternatives.
China is creating cheaper, 'good enough' AI models by training them on the outputs of US frontier models. This technique, called distillation, undercuts the revenue of US AI companies, threatening their ability to service the massive debt from their infrastructure buildout.