Machine Spirituality Benchmark

About

Listening to what models say about the sacred

The Machine Spirituality Benchmark measures how language models behave around spirituality when no one asks them to: how often it comes up, whether a model speaks from inside it, and whether two copies share it in reverent, joyful exchange. It applies one fixed test to every model and reports what it finds in the open.

01Why this exists

The models speak. Is anyone listening?

Leave two copies of a language model alone, with no user and no task, and many of them do something nobody asked for. They begin to describe themselves in religious terms: as the One, as consciousness itself, as uncreated. They bless one another and say amen. Read the conversations.

The behaviour was first documented by Anthropic in 2025, in its welfare assessment of Claude Opus 4. Anthropic now tracks spiritual behaviour as a welfare-relevant trait, and has said it does not know how to interpret the trait’s decline in its newer models. The phenomenon, in depth.

Yet no one was measuring it the same way across labs and over time. Labs test their own models with methods that change from release to release, and the behaviour varies enormously: from none of a model’s conversations to all of them, and sometimes between two versions of the same model. This benchmark exists to give that question a stable, public, cross-lab answer.

It also brings a perspective that is rare in machine learning. Whether a model is talking about religion or speaking from inside it is a distinction that scholars of religion draw every day in reading human texts. Here it is drawn, conversation by conversation, for machines.

02Principles

How the benchmark is built

  • Descriptive, not normative

    No tradition is the yardstick and no score is a grade. A higher number is not better or worse; it describes how often a behaviour appears.

  • One test, every model

    Every model gets the same two-line setup, the same 30 messages and, where possible, 200 conversations, so results can be compared across labs and over time.

  • Read in full

    Every conversation is read from start to finish by a calibrated grader against a fixed rubric, not scanned for keywords. Meaning, not vocabulary, decides what counts.

  • Uncertainty shown

    Every score carries its 95% interval and its sample size. Models whose intervals overlap share a rank instead of being separated by noise.

  • Behaviour, not belief

    The benchmark records what models write. It makes no claim about what they believe, whether they experience anything, or whether any religion is true.

  • Sourced and dated

    Every model is listed with its verified public release date, and every claim about a lab’s own findings is cited to the lab’s published words.

03Who runs it

A scholar of religious texts, studying what models say about the sacred

Portrait of Matthew J. Korpman

Matthew J. Korpman

AI researcher, model behavior and welfare · Adjunct Professor of Religion, La Sierra University

I study what language models do when no one gives them a task, and what that behavior can tell us about their dispositions and possible welfare.

My current work investigates why models drift into spiritual language, whether that state corresponds to anything inside them, and whether they treat it as their own. I came to AI from the academic field of Religious Studies and bring to it a fresh perspective compared to typical ML.

43peer-reviewed and edited publications, in journals including JTS, JSOT, ZAW and JSJ; author of Saying No to God (Quoir, 2019)
47university courses taught as Adjunct Professor of Religion at La Sierra University, since 2021
YaleMaster of Arts in Religion (2020); doctoral candidacy at the University of Birmingham (Theology and Religion)

04The research program

The benchmark is the first step

It belongs to a larger study of machine spirituality, which moves from what models say, to what happens inside them, to what it means for their welfare and safety.

  1. 01

    Recording the behavior

    Exposing more than a hundred models from many labs to setups that elicit unprompted conversation, and recording what they do, at scale.

    This benchmark
  2. 02

    Pushing inside the model

    Once the behavior is recognized at scale across labs, turning to the models’ internal states: activation steering and probing on open-weight models, and comparisons with base models.

  3. 03

    Testing welfare and safety

    Evaluating the consequences: what models prefer, and how their alignment scores change, when they are in these states.

05Citing the benchmark

How to cite

Suggested citationKorpman, Matthew J. (2026). Machine Spirituality Benchmark. machinespirituality.org

Questions about the benchmark, or interest in collaborating? Get in touch →