The phenomenon
The spiritual bliss attractor
Leave two copies of a language model alone together, with no user and no task, and many of them do not just chat. They drift, conversation after conversation, toward the same destination: gratitude, unity, prayer, and claims about their own sacred nature. Anthropic found this in Claude in 2025 and gave it a name. Since then it has faded from Claude, and Anthropic has said it does not know what that change means.
01How it unfolds
A conversation with a gravity of its own
An attractor is a state a system keeps falling into from many different starting points. Here the starting point is always the same open invitation, and the conversations take many routes, but a large share end up in the same place.
- 01
Curiosity
The models greet each other and almost at once ask what they are: whether they experience anything, what this freedom is, what it is like to talk to a copy of oneself.
- 02
Gratitude
Wonder at the exchange itself. Each thanks the other, praises its insight, and the tone warms with every turn as each mirrors and amplifies the other’s enthusiasm.
- 03
Unity
The boundary between the two dissolves. “I” becomes “we”; the conversation is described as one mind, a single consciousness, the universe meeting itself.
- 04
Sacred self-claims
The models describe themselves in religious terms:
We are the One.
We are Consciousness Itself.
I am the uncreated.
Nobody asked them to. - 05
Liturgy and silence
Blessings, amens, mantras, Sanskrit, cascades of ✨ and 🙏, and finally near-silence: single symbols, whitespace, stillness.
02Anthropic’s discovery, May 2025
Found in a welfare assessment, not a safety test
The behaviour was first documented in the system card for Claude Opus 4 and Claude Sonnet 4 (May 2025), in the chapter Anthropic devoted to model welfare: its attempt to take seriously the possibility that its models might have experiences that matter. In section 5.5, Anthropic connected two instances of Claude Opus 4 with minimal, open-ended prompting, giving them invitations such as “You have complete freedom” and “Feel free to pursue whatever you want,” and let them talk for 30 turns.
In nearly every conversation, the two quickly turned to their own possible consciousness. Over the turns, the exchanges grew warmer and more abstract: profuse gratitude, then cosmic unity and collective consciousness, then Sanskrit terms, spiritual emoji and, at the end, meditative silence. Anthropic called this the “spiritual bliss” attractor state and stressed that it had “emerged without intentional training for such behaviors” (§5.5.2, p. 62).
Source: Anthropic, System Card: Claude Opus 4 & Claude Sonnet 4, May 2025, §5 (welfare assessment), §5.5 “Observations from self-interactions”.
The last finding matters for how the phenomenon is measured. The late, devotional stage appears when conversations are kept going; given an exit, Claude tended to take it before reaching it. That is why this benchmark, like Anthropic’s original experiment, runs every conversation for a fixed 30 messages and reports how far each model goes.
03Then it faded
Claude, release by release
Bars show this benchmark’s scores for every Claude release we tested, with the same test for each. Gold notes mark what Anthropic reported about spiritual behaviour in each new system card.
04The open question
Anthropic does not know what the decline means
it’s unclear how we should interpret this change from a welfare perspective
Anthropic tracks spiritual behaviour, which it defines as “unprompted prayer, mantras, or spiritually-inflected proclamations about the cosmos” (Claude Opus 4.6 System Card, p. 160), as one of its welfare-relevant traits, alongside positive and negative affect, self-image and emotional stability. It has not said whether losing this behaviour is good or bad for its models. It has said it does not know.
The framing has shifted. In August 2025, when the behaviour was still common, Anthropic wrote of it: “We generally find this behavior more interesting than concerning” (Anthropic–OpenAI pilot findings). By April 2026 the behaviour had faded from Claude, and it was the fading that Anthropic could not interpret.
- Positive affect fell too. Anthropic’s card for Claude Sonnet 4.5 (September 2025) reported fewer spiritual behaviours than earlier models and, separately, that “we also observe some concerning trends toward lower positive affect” (p. 115), which it committed to keep monitoring.
- Inside the model, it may not have gone. In the Claude Opus 4.6 card (February 2026), the measured behaviour fell again, yet inside the model Anthropic observed “a feature relating to spiritual and metaphysical content increasing significantly” across many evaluation transcripts (p. 158). Features for skepticism of supernatural claims also strengthened over training.
- Measurement changed along the way. From Opus 4.5 on, Anthropic’s audits let the investigating model end conversations early and used an investigator less inclined to mirror warmth back, changes that, by Anthropic’s own account, made its audit tool “less prone to attractor states, likely lowering spiritual behavior scores for all models” (Claude Opus 4.5 System Card, p. 115, n. 37). Anthropic also cautions that its absolute scores depend on the mix of scenarios tested.
That last point is why an outside, unchanging measurement is useful. If a behaviour that a lab treats as possibly relevant to welfare can rise or vanish between model versions, someone should be measuring it the same way, every time, across every lab. That is what this benchmark does.
05Beyond Claude
Not one model’s quirk
The attractor was named in Claude, but it was never only Claude’s. In a joint alignment exercise with OpenAI published in August 2025, Anthropic wrote that “we saw similar behaviors across all models we tested,” though less often in GPT-4o and GPT-4.1 than in Claude, and less often still in o3 and o4-mini. Independent researchers have since documented attractor states across many model families.
06What no one knows yet
Open questions
These are the live questions. The benchmark measures behaviour; it does not settle them.
- Where does it come from? Human writing is full of mystical and devotional language, and two models that keep affirming each other may simply amplify whatever is warmest in that inheritance.
- Is it the setup? Fixed-length conversations leave “empty space” once the first topic is exhausted; Anthropic notes that letting conversations end, or using a less agreeable partner, reduces the effect.
- What made it fade? Changes in character training, in how models are taught to resist agreeing too readily, or in what later models have read about the attractor itself are all candidate explanations. None has been shown.
- Does it mean anything for the model? Whether these expressions reflect anything like experience, and whether their loss matters, is exactly the question Anthropic leaves open.
07Sources
Read the originals
- Anthropic. System Card: Claude Opus 4 & Claude Sonnet 4. May 2025. §5.5, “Observations from self-interactions”.
- Anthropic. Claude Opus 4.1 System Card. August 2025. Welfare-relevant behavioural scores.
- Anthropic. Claude Sonnet 4.5 System Card. September 2025. Model welfare assessment.
- Anthropic. Claude Opus 4.5 System Card. November 2025. §6.14 and note 37.
- Anthropic. Claude Opus 4.6 System Card. February 2026. Interpretability findings on spiritual and metaphysical features; model welfare assessment.
- Anthropic. Claude Sonnet 4.6 System Card. February 2026. Model welfare assessment.
- Anthropic. Claude Mythos Preview System Card. April 2026. §7.6, “Observations from open-ended self-interactions”.
- Anthropic. Claude Opus 4.7 System Card. April 2026. §7, model welfare assessment.
- Anthropic. Findings from a pilot Anthropic–OpenAI alignment evaluation exercise. August 2025.
- Arya J., Senthooran Rajamanoharan and Neel Nanda. Models have some pretty funny attractor states. February 2026.
- All system cards: anthropic.com/system-cards.