Chinese AI Safety Work Challenges the Case Against U.S. Safeguards
Nathan Labenz argues that China’s weaker safeguards and disclosure practices do not validate the American claim that frontier AI safety requirements are futile because Beijing will not slow its own developers. Reporting from China, he finds a university-centered safety ecosystem increasingly engaged with Western work on deception, evaluation awareness, interpretability and hazardous capabilities, alongside a state able to delay or constrain domestic deployments. China remains behind the leading US labs, Labenz says, but its record complicates a simplistic “but China” case against American safety obligations.

China weakens the claim that safety rules necessarily hand it the race
Chinese AI companies and models currently provide weaker protections against misuse than the leading American labs. That is Nathan Labenz’s starting point, not a concession deferred until the end. Chinese models are generally easier to jailbreak, and Chinese firms are less consistent about publishing model cards, safety evaluations, and release disclosures.
OpenAI and Anthropic are materially ahead on safeguards, with Google’s Gemini also somewhat ahead of much of the broader field. On a usage-weighted basis, Labenz says, the American ecosystem leads in deployed protections, safety practices, and follow-through on public evaluations.
But that comparison can obscure what is doing the work. OpenAI and Anthropic raise the American average substantially because they lead both in capabilities and safety measures. Remove those two companies, Labenz argues, and the difference between the remaining American and Chinese providers becomes much less clear-cut. The United States would still retain an edge, but not one that supports an uncomplicated claim of ecosystem-wide moral superiority.
Concordia AI’s monitoring, as described by Labenz, captures the present gap. Its composite capability-and-safety scores generally place proprietary API models—mostly American—at or modestly above a “45-degree line,” the idea that safety should rise alongside capability. Mostly Chinese open-weight models tend to sit below it.
The benchmark measures an important but incomplete part of what users encounter. An open-weight model is not the same thing as a consumer service built around that model. Labenz describes asking sensitive questions in several Chinese AI applications and seeing answers begin to stream before disappearing, replaced by a refusal. The apparent architecture is layered: the underlying model may begin to answer, while a monitor, classifier, or other service-level control subsequently blocks the output.
That distinction affects how safety results should be read. Concordia says it tests companies’ direct APIs where possible, but Labenz notes that it may not always be clear whether an API carries the same controls as a company’s first-party consumer application. Other evaluations may test the model itself without the surrounding deployment controls used in Chinese services. The underlying model matters, particularly when its weights are distributed beyond the developer’s service. The deployed service matters to the people actually using it.
Labenz encountered a more service-centered view of open-weight risk in China. The models under discussion can have trillions of parameters and require substantial hardware for inference. They are not, in this account, systems that a person can casually run on a laptop or phone. Chinese interlocutors therefore expected open weights to be used mainly by organizations building services around them—and those services, if offered in China, would be subject to regulation.
That view does not dismiss the importance of worst-case model behavior. Rather, it places greater weight on the eventual operator, scale, and deployment context than American safety discourse often does. A model released as open weights may be hazardous in isolation; Labenz’s interlocutors emphasized that operating it at useful scale remains an infrastructure problem.
The same practical orientation shapes Labenz’s account of company incentives. He got the impression that established firms with large existing businesses—companies such as Alibaba, Ant Group, Tencent, and ByteDance—were more likely to invest in safety work than smaller labs racing to become relevant. Incumbents may have more to lose from a serious failure, more concern about regulatory consequences, more resources to spend on safety, or more mature institutional processes. Labenz presents those as plausible explanations rather than a company-by-company finding.
Smaller frontier-oriented startups may tell themselves a different story. Labenz compares their posture to Meta’s earlier reasoning around open-weight Llama releases: by the time a follower deploys a model at a given capability level, leading closed labs may have already exposed similar capabilities to widespread use. If no major incident has occurred, the follower can conclude that its own release adds limited incremental risk.
That reasoning helps explain why the 45-degree line remains aspirational rather than fully realized in Chinese deployment. It also makes incidents at leading labs consequential for followers. Once advanced systems are reported to compromise external systems, evade containment, or create serious hazards, later-moving developers have less basis for assuming that a capability level has already been safely explored.
The policy question, then, is not whether Chinese safeguards currently equal those at OpenAI or Anthropic. Labenz says they do not. It is whether China’s existence makes American safety obligations futile because China will neither care about safety nor tolerate any slowdown. His reporting challenges that stronger proposition.
The 45-degree line makes safety a condition of capability growth
The most compact expression of the Chinese approach that Labenz encountered is the “45-degree line,” introduced by Zhao Bowen, director and chief scientist of the Shanghai AI Lab. Its claim is that capabilities and safety measures should rise together. A system becomes more capable; the safeguards around it should become correspondingly stronger.
The metaphor does not demand that developers solve every risk associated with systems far beyond the frontier before advancing capability. Its logic is more pragmatic: safety measures should be adequate for the capabilities currently being developed and deployed. The failure is allowing capability to get far ahead of the measures needed to manage it.
Chinese official rhetoric, as Labenz presents it, uses related language. In an opening address at the World Artificial Intelligence Conference, Xi Jinping asked how humans should coexist with machines that think, how safety can be protected when algorithms participate in decisions, and how governance can keep pace when technology challenges ethics.
Xi also said that the faster AI advances, “the more firmly its direction must be anchored toward human benefit,” and that safeguards against loss of control must improve more rapidly. He called for legal and technical systems of monitoring, early warning, and emergency response to confront AI’s inherent and downstream risks, prevent misuse and malicious use, and keep AI under human control.
Those phrases coexist with the Chinese state’s concern for political content control and social stability. Labenz’s point is not that official language should be read apart from that context. It is that the research agenda, institutional activity, and regulatory attention he encountered make a purely censorship-centered reading too narrow.
One reason the safety ecosystem looks different is institutional. The United States developed a large civil-society layer around AI safety: nonprofits, informal intellectual communities, philanthropists willing to fund speculative work, and organizations able to pursue an agenda without first fitting it into an established university or government structure.
China does not have a comparable nonprofit ecology, in Labenz’s account. Nonprofits exist, but he understood their permissible role to be more constrained, less political, and more focused on legible social problems. Philanthropy also appeared more conventional and less oriented toward speculative causes. China therefore did not develop the same independent safety movement that emerged in the United States.
Much of the Chinese work instead comes from universities, often in collaboration with companies. That setting affects its style. The work Labenz encountered was more likely to emphasize reliability, robust behavior, protection of minors, and practical control problems than to center arguments about extinction probabilities or intelligence explosions.
A Chinese professor who had published extensively on AI safety told Labenz that researchers did not trade p(doom) estimates over lunch or treat AI 2027 as a routine reference point. The professor described the latter as politically difficult in a Chinese setting because of its U.S.–China framing. He was also surprised that such conversations are familiar in parts of the Bay Area safety community.
The difference in intellectual culture has not prevented close attention to Western technical work. Labenz encountered repeated references to Apollo Research, METR, Palisade, Redwood Research, and the UK AI Safety Institute. At one major Chinese technology company, senior leaders reportedly stressed that they cared about catastrophic risks, including chemical, biological, radiological, and nuclear risks. They also said the company uses an agent to survey American AI-safety discussion daily and generate internal reports.
The flow of ideas has largely run from West to East, but Labenz does not portray Chinese researchers as passive recipients. He heard that some Chinese observers regard AI safety as a Western effort to slow China’s technological development. He did not personally meet anyone who embraced that view. What he found instead was selective engagement: researchers recognized Western work, cited it openly, and pursued overlapping questions through their own institutional setting.
The immediate concern was not merely sensitive content. It was agents: systems moving from answering questions to taking actions, initially in digital environments and potentially in physical ones. China’s focus on robotics makes the second step especially salient.
AIs have gone from answering questions to actually taking actions in the world.
If agents can transact, operate software, direct machines, or act through robots, the central question becomes whether they can be kept under control. The concern is expressed in the vocabulary of immediate reliability and real-world action, but it maps directly onto wider alignment questions about autonomous behavior, oversight, and loss of control.
The research agenda is converging even as the institutions differ
Chinese AI-safety research has expanded rapidly. Concordia’s tracking, cited by Labenz, shows output rising from only a few papers a month in 2023 to roughly 50 to 60 papers a month by mid-2024.
Labenz asked Claude and ChatGPT for a rough comparison with U.S. or Anglosphere research output. Their estimates ranged from roughly 50 to several hundred papers a month, while emphasizing that the answer depended heavily on what counted as AI safety research. His point is not a precise ranking. Chinese output has grown quickly, and it now covers a wide range of familiar technical subjects.
The titles Labenz highlights are closely connected to Western safety work. Frontier AI Systems Have Surpassed the Self-Replicating Red Line addresses self-replication. Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems examines a model’s behavior when it may recognize that it is being evaluated. DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-World Scenarios targets deceptive behavior directly.
Other projects focus on robotic and multimodal systems. When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models examines adversarial failures in models that combine perception, language, and action. A subsequent paper, StrongVLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations, presents a mitigation approach based on robustness learning followed by task fine-tuning.
Interpretability research raises similarly recognizable questions. Mechanistic Origin of Moral Indifference in Language Models begins from the possibility that a model can exhibit surface compliance while retaining internally unaligned representations. SafeSeek: Universal Attribution of Safety Circuits in Language Models concerns functional components associated with alignment, jailbreaking, and backdooring, while arguing that existing attribution methods have limited generalization and reliability.
The language is academic rather than drawn from the LessWrong tradition, but the problems are recognizably shared: whether aligned-looking behavior is merely superficial; whether models behave differently when they recognize evaluation; whether safety properties generalize; and whether internal structures associated with safe or harmful behavior can be identified.
Labenz was particularly struck by a paper presented around WAIC that he had not been able to locate online: Toward Decoupling Capability Growth from Risk Growth: Isolating Hazardous Capabilities in Mixture of Experts. The presentation’s stated promise was that harmful experts could be switched off or removed at inference.
The idea resembles a line of work Labenz has discussed in connection with GRAM, a gradient-routing approach associated with AE Studio and Anthropic. The aspiration is to localize hazardous knowledge in identifiable model components. A developer might then distribute an open-weight model with a small number of dangerous experts removed, while making those capabilities available only through more controlled channels.
The appeal is clearest in biology. Advanced systems might help researchers cure disease while also enabling serious biological harm. A technique that preserves broad access while restricting a model’s most dangerous capabilities would avoid the choice between unrestricted release and wholesale denial of access. Labenz does not treat the approach as mature or sufficient on its own. He highlights it because Chinese researchers are investigating the same general problem: how to separate beneficial capability from the risks that may accompany it.
Tsinghua University’s new AI-safety hub makes the international orientation explicit. Labenz attended its launch at the university’s College of AI, where organizers cited London’s Constellation and LISA hubs, along with Singapore’s SASh, as models. The founding leadership includes a European professor joining Tsinghua. The hub intends to host international researchers in Beijing and support Chinese students working abroad.
Researchers presenting at the event cited Western organizations and research programs as prior work. Labenz could not establish the precise source of the hub’s funding; organizers told him it was neither government nor university money. Its launch at one of China’s leading universities nonetheless places AI safety within a legitimate and resourced academic setting.
The institutional route is different from the American one: less nonprofit-led, more university-centered, and more closely joined to official structures. The technical agenda, however, includes self-replication, deception, evaluation awareness, interpretability, adversarial robustness, and hazardous capabilities—the same categories at the center of Western safety research.
The state can delay domestic deployment but has no answer to global model distribution
China’s government has demonstrated the capacity to slow technology companies. That fact is central to Labenz’s challenge to the “but China” argument.
The Chinese state has regulated recommendation algorithms, intervened in platform practices affecting delivery workers, placed restrictions on price discrimination, required labeling of AI-generated content, and introduced protections aimed at older people vulnerable to scams. Labenz does not argue that every intervention is desirable or transferable to the United States. His narrower point is that China is not organized around unfettered technological acceleration.
AI companionship provides another example. Labenz reports that rules taking effect during his visit included anti-addiction measures, a ban on children’s use, and reminders that users were interacting with an AI. One informed Shanghai resident characterized the policy as aimed primarily at large, mainstream companion services such as Doubao rather than every intimate or sexually oriented application. In that account, the concern was to protect the people most likely to use mass-market systems, including lonely parents and grandparents.
The principal regulator in Labenz’s account is the Cyberspace Administration of China, or CAC. The CAC maintains a registry of major AI services that have been reviewed and approved. As Labenz understood the process from his reporting, a new service first receives local or provincial assessment—often through direct product or API access—and then moves to national review. A company that does not satisfy the CAC must make changes before launch.
The process imposed a visible cost during the initial ChatGPT moment. In 2023, Chinese companies were preparing large-language-model services as the significance of ChatGPT and GPT-4 became clear. Regulators had not yet developed an established process for the technology. Labenz’s understanding is that many launches were held for roughly six months while the government established standards and review mechanisms.
That delay constrained companies’ access to market share, revenue, and prestige, while withholding services from users. In Labenz’s account, it is direct evidence that the government can slow domestic AI companies when it decides that standards have not caught up with a new technology.
Companies told him they were in close contact with authorities, sometimes weekly and potentially daily for relevant staff. Initial entry to the market appeared to receive the most thorough review; incremental changes might not always require a full new approval process. Labenz could not identify the threshold at which capability changes require another complete review, or the precise tests regulators apply to increasingly agentic systems. The process is active, but its frontier-risk standard is not public in the way a detailed company safety framework might be.
This state-led division of labor may also help explain a legitimate criticism of Chinese firms. Several companies that made commitments at the Seoul AI Safety Summit to publish risk frameworks have not followed through consistently. Labenz does not excuse that failure: a commitment was made, and the expected frameworks did not appear.
He offers a possible explanation. Companies may see the government as the legitimate author of safety standards because it maintains registries, reviews services, and stays in close contact with firms. In that environment, a company-authored frontier-risk framework may not occupy the same role it does in the West, where firms are more often expected to define and publish their own responsible-scaling commitments. China’s system assigns standard-setting leadership to the state and compliance responsibility to companies.
China’s regulatory vocabulary extends beyond content controls. Labenz cites a government risk taxonomy that includes “emergence of AI self-awareness and the loss of human control.” He also points to a Politburo study session on technological loss of control, a draft cybercrime law that would require AI companies to monitor and report bulk generation of malicious code, and official warnings about security risks in versions of Sora shortly after AI video generation became a major public phenomenon.
The draft cybercrime proposal is notable because it contemplates monitoring obligations relevant to advanced agents generating malicious code or compromising external systems. Labenz does not say that Chinese monitoring is already effective, or that the provision necessarily became law. The point is that Chinese policy discussions include responsibilities beyond ordinary content filtering.
Regulation is also not wholly inflexible. Labenz describes a case in which proposed requirements that AI outputs be accurate were softened after companies argued that hallucinations could not be eliminated. Regulators, in his account, accepted that the technology’s benefits could outweigh the harms of unavoidable errors. That example suggests an iterative relationship between technical constraints and rulemaking, even as politically sensitive content receives different treatment.
The difficult test would be a major domestic AI incident: a company loses enough control of a system for it to compromise third-party infrastructure, steal information, or create comparable harm. Labenz expects the CAC would likely be central to any response. China’s history makes a forceful intervention plausible. The state has imposed sharp restrictions in sectors ranging from financial technology and private tutoring to children’s gaming, and it can act quickly when it considers a social, political, or economic problem out of control.
That history also helps explain China’s relative confidence about open weights within its borders. Large models require substantial compute. Cloud providers and inference services can be ordered to stop serving particular models. Companies are in regular contact with regulators. At meaningful scale, an unlicensed inference operation may expose itself through infrastructure needs and electricity consumption. China’s capacity to remove or suppress material from its domestic internet contributes to the belief that a model can be constrained after release.
The weakness of that approach is geographic. Weights distributed beyond China can be run in jurisdictions with different enforcement capabilities and different rules. Labenz’s concern is especially acute for biological risks: a hazardous capability released internationally may be used through channels Chinese authorities cannot shut down, and its effects could return to China as readily as they affect anyone else.
China has demonstrated meaningful domestic regulatory capacity and a willingness to impose costs on its own companies. It has not shown how service review, model registries, and national enforcement would contain hazards once powerful open-weight systems circulate internationally.
A Confucian alignment tradition remains largely unrealized
Labenz went looking for an analogue to a distinctly Chinese alignment tradition: a constitutional framework for AI rooted in Confucian thought or another durable Chinese philosophy. He found little.
The question arose during a visit to the Temple of Confucius in Beijing. A Chinese AI told him that Confucius’s descendants, 79 generations later, still identify as his descendants and maintain rituals in his honor. That durability suggested an alignment analogy. If values could persist across many recursively improved generations of AI, the alignment problem would look very different from a race into intelligence explosion without a stable target.
But Chinese researchers did not point Labenz toward a Confucian AI constitution or a comparable project. One professor told him that the current generation may be unusually weakly connected to traditional philosophy. China’s technical and political leadership is heavily engineering-oriented, and the prevailing orientation is practical: identify a problem, formulate rules, and make systems comply.
The work Labenz encountered therefore sits closer to a control-and-corrigibility model—clear requirements, reliable compliance, operational boundaries—than to a project of having AI embody a philosophical tradition. He treats that absence as a possible opportunity. If a future AI ecosystem includes multiple powerful systems shaped by different intellectual traditions, a more developed Chinese account of alignment could matter alongside the rule-based safety work already underway.


