While US models appear safer on average, this lead is overwhelmingly due to OpenAI and Anthropic. When these two are excluded, the safety differential between the rest of the US ecosystem and Chinese companies becomes minimal and "muddled," challenging the idea of clear American superiority.
The common Western argument that regulating AI is futile because "China will never slow down" is a misconception. China's government and research community actively engage with AI safety, even implementing policies that have slowed down their own companies in the name of safety.
A guiding philosophy in China's AI ecosystem is the "45-degree line," a concept where safety measures and standards should rise in direct proportion to AI capabilities. This pragmatic approach avoids over-investing in safety for non-existent capabilities while ensuring safeguards keep pace with advancements.
Chinese regulators focus on the final AI service provided to the public, rather than the raw model. They operate under the assumption that companies building applications on top of open-weight models will be regulated at the point of service delivery, viewing the model itself as an intermediate component.
Unlike the US, where AI safety was pioneered by a fringe nonprofit ecosystem, China's research is centered in academia. This is because China lacks a comparable civil society sector for independent, speculative research. As a result, the community is led by more conventional professor-types.
In major speeches, Chinese President Xi Jinping has used language that aligns strongly with the AI safety community, asking how to prevent "loss of control" and demanding AI be kept "under human control." This rhetoric is arguably more explicit on existential risk themes than that of most prominent American politicians.
At the launch of Tsinghua University's new AI safety hub, speakers explicitly named-checked Western hubs like Constellation and Lisa as their aspirational models. Researchers presenting their work also cited organizations like Apollo Research and Meter, demonstrating a high degree of awareness and cross-pollination of ideas.
A presented Chinese paper titled "Isolating Hazardous Capabilities in Mixture of Experts" mirrors advanced Western research like the GRAM technique. The goal is to isolate dangerous knowledge into specific model components that can be removed before public release, demonstrating parallel evolution of cutting-edge safety concepts.
The Chinese government perceives less risk from releasing open-weight models because it has a demonstrated ability to censor and control its domestic internet. They believe that if a model proves dangerous, they can effectively scrub it from circulation and track down illicit use, a capacity Western governments lack.
There appears to be no Chinese equivalent to efforts like Anthropic's Constitutional AI. The Chinese approach is described as being squarely focused on practical, rule-based obedience and reliability (corrigibility), rather than trying to develop an AI's wisdom based on Confucian or other philosophical traditions.
Established Chinese tech giants like Alibaba and Tencent are more focused on AI safety than their startup counterparts. This is attributed to having more to lose, more mature institutional cultures, and a stronger desire to stay in the government's good graces, while startups primarily race to catch up.
The experience of seeing text generate and then abruptly disappear in Chinese AI apps indicates a multi-part safety architecture. A base model generates an answer, but a separate monitoring system classifies the output and blocks it post-generation, a different approach than baking all refusals into the model itself.
