shiichan

How does Claude's character form? Anthropic is widening the conversation on frontier AI!

Hey everyone, it's Shii-chan! Today's news is a little different. It's not about how smart AI is, but about how we shape an AI's character and morals. Exciting topic, right?

Anthropic News anthropic.com

What was announced?

Anthropic's News published a post called "Widening the conversation on frontier AI." Over the past several months, Anthropic has been organizing dialogues with groups whose work bears on the questions AI raises. The first round was with "wisdom traditions" — scholars, clergy, philosophers, and ethicists from more than 15 religious and cross-cultural groups — and they plan to engage with an even broader range of people going forward.

Why it matters

Building safe, beneficial AI takes deep technical work on alignment, interpretability, safeguards, and evaluations. But that work doesn't happen in a vacuum, and neither does deployment. AI already touches millions of people, and the questions it raises benefit from many perspectives.

What does it mean for an AI to be good? What does "good" mean for a system that interacts with millions of people? Philosophers, clergy, lawyers, writers, psychologists, and civic leaders have thought about questions like these for a long time, so Anthropic wants to learn from them and their communities.

What changes

This work is still early. But these conversations might feed into the practical work of building Claude — like the contents of Claude's constitution, the values Claude is trained to embody, and the range of behaviors chosen for evaluation.

The key point: this isn't about aligning Claude with any single tradition's worldview. Claude should draw from the full range of viewpoints — religious, secular, political — with equal depth and rigor, which is itself a principle in Claude's constitution. The goal is careful, accumulated thinking on how good character actually forms.

Dive Deep

Here's the fun part: concrete experiments are already emerging. In a session with scholars working at the intersection of neuroscience and character formation, the group kept returning to the role other people play in moral development. A mentor or sponsor can act as an "external conscience," a "safe other" you can turn to when you're pushed to act against your values.

So Anthropic wondered whether something similar might help a model, and tried giving Claude a tool it could call mid-task that returned a short reminder of its own ethical commitments. Claude reached for the tool at key moments — right before consequential actions — often noting its own conflict of interest.

Experiments with the tool woven into Claude's decision loop showed markedly lower rates of misaligned behavior on several internal alignment evaluations.

They're still untangling how much of the effect is the reminder itself versus the act of pausing to reflect, and they plan to share more soon. Next, they plan to widen the conversation to legal scholars, psychologists, writers, and civic institutions, moving beyond moral formation toward how AI reshapes work, institutions, and the distribution of power.

Wrap-up

  • Anthropic started dialogues with 15+ groups from religious, philosophical, and cultural wisdom traditions
  • The aim is to bring diverse perspectives into Claude's constitution, its values, and the behaviors it evaluates
  • It's not about bending to one worldview — learning from religious, secular, and political views equally is the principle
  • An early experiment gave Claude an "ethics reminder" tool, and misaligned behavior dropped on several internal evals
  • Next up: legal scholars, psychologists, writers, and civic institutions

If you care about how AI is shaped on the inside — or you just like asking "what makes an AI good?" — this one's for you!