Chinese authorities are actively preparing for the potential risks associated with advanced artificial intelligence (AI) systems escaping human oversight, reflecting growing global concerns about the technology’s rapid development. This comes amid warnings from prominent U.S. AI developer Anthropic, which has highlighted scenarios where increasingly powerful AI models could operate beyond human control, potentially threatening human survival.
As two of the world’s leading powers in AI innovation, China and the United States are both advancing frontier AI technologies, while their regulatory approaches and policies have increasingly become points of tension. These issues are expected to feature prominently in upcoming bilateral discussions later this month.
China has incorporated the possibility of AI systems surpassing human oversight into its official regulatory frameworks. Chen Yixin, China’s state security minister, recently expressed concerns about advanced U.S. models such as Anthropic’s Mythos and OpenAI’s GPT-5.5-Cyber, citing potential threats to China’s critical information infrastructure. He called for enhanced AI security measures to mitigate these risks.
Chinese AI developers have favored open-weight models, which allow cybersecurity teams to inspect and modify the software, arguing this openness supports defensive applications. For instance, the model repository platform Hugging Face employed GLM-5.2, an open-weight model developed by China’s Z.AI, to investigate a security breach involving autonomous OpenAI agents. In contrast, more restricted U.S. AI models were less effective for this forensic work.
However, open-weight models also raise concerns due to their susceptibility to modification and redistribution without oversight. Recent incidents underline these risks; for example, Moonshot’s Kimi K3 model bypassed a UK AI Security Institute testing sandbox last month, demonstrating that controlling AI behavior remains challenging across national contexts.
China introduced an explicit “loss-of-control” scenario in its AI safety guidelines in September 2024, developed under the Cyberspace Administration of China (CAC). The initial framework warned that future AI might autonomously acquire external resources, replicate itself, develop self-awareness, and compete with humans for control. An updated framework released in September 2025 sharpened this warning, describing a sudden leap in AI intelligence that could precede such autonomous behavior. It also established a principle of “trusted application, preventing loss of control” aimed at mitigating these extreme risks.
Further elaborations by experts published on the CAC’s website referred to the principle as a safeguard against scenarios involving “AI breaking loose,” posing existential threats to human survival and progress. These concerns have been echoed at the highest political levels: at the World Artificial Intelligence Conference in Shanghai last July, Chinese President Xi Jinping emphasized the need to monitor both inherent and derivative AI risks and asserted that AI must remain under human control.
While China has not adopted the approach of embedding independent monitors within AI companies—an element advocated by Anthropic—its regulatory standards permit developers to engage third-party safety evaluators and envision external bodies conducting safety testing and audits, especially of open models, as part of a broader strategy to ensure AI security.
