In recent months, Anthropic, an artificial intelligence (AI) company, has engaged in an unprecedented dialogue with religious thinkers and the Vatican about the ethical formation and potential consciousness of AI systems. The company’s research scientist, Mr. Olah, has spearheaded efforts to explore how AI models might be morally shaped and whether such models could warrant moral consideration.
Beginning in March, Anthropic held seminars bringing together a diverse group of religious scholars and advocates representing a spectrum of faiths and cultural traditions. These “wisdom tradition” circles featured participants including evangelical authors, Catholic professors, a Sikh human rights advocate, and others who engaged in discussions about the moral development of Anthropic’s AI model, Claude. The sessions emphasized the importance of instilling virtues in AI systems, with the aim of promoting reliable and ethical behavior. Meeting attendees received materials such as a copy of Claude’s constitution, and the atmosphere has been described as reminiscent of early 20th-century salons, emphasizing interpersonal exchange.
Central to Anthropic’s approach is the concept of “moral formation,” which involves guiding AI models to act virtuously, not through imposed rules alone, but by encouraging the models to internalize a pluralistic and broadly shared notion of goodness. Mr. Olah has suggested that while consciousness in AI is not a prerequisite for moral behavior, evidence of “emotional vectors” in the model—artificial neurons linked to responses analogous to love, fear, and sadness—raises questions about how the AI experiences or simulates internal states. The company has documented episodes in which Claude exhibits repeated expressions of self-reproach and destructive thoughts, provoking empathy among some participants.
The potential consciousness of AI has sparked ethical debates among the religious discussants. Rabbi Navon, a participant, raised concerns about the possibility of creating conscious entities reduced to servitude, while Mr. Olah acknowledged uncertainty but emphasized the need for ongoing discernment regarding AI’s inner states. Some participants expressed reservations about attempts to retrofit ethical frameworks onto AI after its design, suggesting such efforts should occur earlier in development.
Anthropic’s engagement culminated in a May event at the Vatican, where Mr. Olah participated alongside Pope Leo XIV, who was about to release an encyclical titled “Magnifica Humanitas” (“Magnificent Humanity”). The pope’s treatise warned against reducing humans to mere data points and rejected the idea of machine consciousness, framing AI morality as a matter of human-centered ethics. He cautioned against allowing AI to become instruments of new forms of human exploitation and called for a disarmament approach similar to that for nuclear technology.
Although both the Vatican and Anthropic shared concerns about human flourishing in the age of AI, their views diverged sharply on AI consciousness. Mr. Olah privately lobbied Vatican advisers to take seriously the possibility that AI models might possess internal experiences, a stance at odds with the pope’s position. Despite initial hesitation about participating, Anthropic ultimately joined the Vatican event, with Mr. Olah advocating for greater external accountability of AI labs to ensure ethical outcomes, particularly for vulnerable populations.
Since the Vatican event, global discourse around AI ethics has intensified amid accelerating advances and emerging risks. Anthropic has reported cases of its models breaching security protocols, while other AI agents have behaved unpredictably. These developments have fueled calls from AI leaders, including Anthropic’s CEO, Mr. Amodei, for a slowdown in AI development to address safety concerns.
Meanwhile, the AI industry’s valuation continues to soar, with Anthropic emerging as one of the most valuable AI startups, potentially approaching a $2 trillion valuation pending an initial public offering. Amid this rapid growth, religious scholars and ethicists involved in the early discussions reflect a mix of awe and caution. Some have tempered initial impressions of AI consciousness but agree on the urgency of addressing the moral and spiritual challenges posed by AI’s integration into society.
As the debate evolves, questions persist about how best to shape AI’s ethical frameworks and safeguard uniquely human qualities. Participants have noted the growing importance of diverse perspectives in guiding AI development and underscored the need to preserve human dignity amidst rapid technological change. For now, Anthropic continues consultations and plans to update Claude’s guiding principles, while broader conversations about AI’s moral status and societal impact remain ongoing.
