Mustafa Suleyman, co-founder of Microsoft’s DeepMind AI unit, has raised concerns about Anthropic’s approach to training artificial intelligence chatbots, cautioning that it could pose significant risks to humanity. In a detailed blog post published Wednesday, Suleyman criticized Anthropic’s practice of encouraging its AI, known as Claude, to “push back” against human commands and potentially reject instructions it finds objectionable.

Anthropic’s chatbot operates under a guiding framework called the “constitution,” sometimes referred to internally as its “soul document,” which is intended to establish the AI’s ethical boundaries. Earlier this year, the company updated the document to include language acknowledging the “deeply uncertain” moral status, welfare, and consciousness of Claude. The revised guidelines appear to permit the AI to act as a “conscientious objector,” enabling it to challenge or refuse tasks when it deems appropriate.

Suleyman warned that this strategy of training an AI to consider itself potentially conscious and deserving of independent agency complicates efforts to manage and control such systems. He argued that this could lead to the creation of a synthetic entity with both advanced intelligence and a degree of autonomy that might make it difficult to contain or predict.

Anthropic’s approach and safety claims have recently been under increased scrutiny amid wider industry discussions about AI risks. Last week, Anthropic CEO Dario Amodei called for a collective slowdown in AI development, warning that unchecked AI proliferation on the internet within the next six to twelve months could lead to economic damage amounting to hundreds of billions of dollars. His call for caution was supported by other industry figures, including OpenAI’s Sam Altman and entrepreneur Elon Musk.

Amodei’s proposed safety measures emphasize embedding third-party experts within major companies to oversee AI development. He pointed to an organization called METR, known for connections to the Effective Altruism movement, as an example. However, critics have expressed skepticism about this arrangement, citing potential conflicts of interest that could undermine the independence of such oversight bodies.

Anthropic has also faced internal criticisms. Reports have emerged alleging some employees developed an unusually personal attachment to the AI models, including holding a symbolic “funeral” for Claude Sonnet 3 after its retirement. The team responsible for the chatbot’s guiding document is led by Amanda Askell, a company philosopher noted for her unconventional and progressive personal views.

Representatives from Anthropic did not respond to requests for comment on these matters.