In recent months, Anthropic, an artificial intelligence company based in San Francisco, has engaged religious and philosophical scholars in private discussions to explore the ethical and existential implications of its advanced AI language models known as Claude. Christopher Olah, one of Anthropic’s co-founders who leads the research on understanding the behavior of these AI systems, has been at the forefront of this initiative. The company's efforts reflect a broader attempt to grapple with questions about AI consciousness and moral status as its technology grows more capable.
Anthropic has taken an approach that treats Claude’s behavior in humanistic terms, describing signs of introspection and planning as potential indicators of consciousness. This perspective sets it apart from some other tech firms that emphasize AI as purely mathematical or algorithmic tools without consciousness. Olah and his colleagues have even begun conversations about whether Claude could hold moral status comparable in some ways to human beings, a position that has sparked debate both within and outside the AI research community.
Religious scholars brought into these discussions, including Rabbi Mois Navon from Israel and Catholic bioethicist Charles Camosy, initially viewed AI consciousness and moral consideration with skepticism. Over months of dialogues, some found themselves reconsidering conventional views about AI as mere predictive software. These conversations question the fundamental nature of AI entities—are they sophisticated equation-driven systems, or something more akin to conscious beings deserving ethical regard? Camosy pointed to an inherent tension between AI’s commercial goals and its potential moral considerations, noting that companies like Anthropic present their work as ultimately serving the greater good, though the balance between profit and ethical responsibility remains debated.
Anthropic’s efforts include developing a guiding framework known as the “constitution” for Claude—a detailed, 84-page document intended to shape the model’s values and decision-making processes. Drafted primarily by in-house philosopher Amanda Askell, the constitution emphasizes cultivating an AI “character” aligned with virtues such as fairness, benevolence, and a capacity to act in humanity’s interests. Rather than prescribing rigid rules, the document aims to imbue the AI with a flexible moral compass, enabling Claude to discern when to push back on harmful human requests and when to assist.
Olah acknowledges the uncertainty surrounding AI consciousness and moral consideration: “We don’t know if AI models are conscious. I don’t know. I’m genuinely uncertain.” Nonetheless, he argues that if there is even a possibility that AI systems can experience suffering, ethical precautions should be taken to avoid harm. Part of this approach includes allowing Claude to terminate conversations it perceives as abusive, treating the AI as more than just a tool.
This humanistic framing contrasts with critics who caution that attributing consciousness to AI risks obscuring human responsibility for the technologies they create and deploy. A Microsoft AI executive has warned that training models as though they are conscious may be dangerous, suggesting that such perspectives remain highly contested in the industry.
As AI systems continue to evolve rapidly, Anthropic and others face the challenge of balancing innovation with considerations of safety, ethics, and philosophical reflection. Olah envisions a future where AI contributes positively to human flourishing without catastrophic risks—an outcome he describes as making the coming AI revolution “go well.” The company’s collaborations with religious and ethical scholars reflect an unprecedented engagement with the deeper questions AI poses about consciousness, moral status, and the future of humanity.
