Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

To determine if one AI model has been trained on the output of another (a process called distillation), analysts use semantic similarity charts. These charts compare the diction and phrasing of different models. High correlation between a new model and existing ones, like Chinese model Kimi K3 and Anthropic's models, suggests potential distillation, though it isn't definitive proof.

Related Insights

Accusations that Chinese labs cheat by copying US models are misleading. The practice, known as distillation, is common across the industry (including by Elon Musk's xAI) and academia. Now, with Chinese labs dominating open source, American startups are increasingly building on top of Chinese models.

As more of the public internet and code repositories are generated by LLMs, any new model trained on this public data is, in effect, being 'distilled' from other models. This complicates accusations of direct distillation and blurs the line for what constitutes original training data.

Leading AI labs, despite intense competition, are collaborating through the Frontier Model Forum to detect and prevent Chinese firms from creating imitation models. This rare alliance is driven by the shared existential threat that 'adversarial distillation' poses to their business models and to U.S. national security.

As more of the internet and code repositories are generated by leading AI models, any new model trained on this public data inadvertently "distills" the knowledge and quirks of those proprietary systems. This blurs the line between original training and outright copying.

Despite intense domestic rivalry, top US AI labs like OpenAI, Anthropic, and Google are collaborating to detect "adversarial distillation"—where Chinese firms copy their models. This rare cooperation shows the shared commercial and national security threat from foreign competitors outweighs their direct competition.

A new form of analysis compares the semantic similarities (e.g., diction, phrasing) of outputs from different AI models. This technique is being used to create 'fingerprints' that can suggest if one model was illicitly 'distilled' or trained on the outputs of another, a key concern in the AI arms race.

Chinese labs use 'smart distillation,' a sophisticated technique where a frontier model acts as a 'teacher' to guide a smaller model's judgment and data labeling. This is viewed as a legitimate and efficient catch-up method, distinct from simply copy-pasting answers.

New open-weight models like Inkling are not entirely 'pure'; they use 'distillation light' from other open models (e.g., Kimi). Since those models may be distilled from closed-source giants like OpenAI, it creates a multi-layered dependency chain where traits and biases are passed down, blurring the lines between truly independent and derivative models.

Leading Chinese AI models like Kimi appear to be primarily trained on the outputs of US models (a process called distillation) rather than being built from scratch. This suggests China's progress is constrained by its ability to scrape and fine-tune American APIs, indicating the U.S. still holds a significant architectural and innovation advantage in foundational AI.

Unable to build frontier models from scratch, some Chinese companies gain a competitive edge by using "scale distillation." This involves training smaller, open models on the outputs of larger, proprietary US models, effectively piggybacking on American R&D to create capable, low-cost alternatives.