Each finding, with the chart that shows it and whether it survived scrutiny. Things we got wrong are in the last section.
Expert #164: the digit detector (originally misidentified)
This is the expert that started everything. Expert #164 at layer 42 fired at exactly zero on the Bible and at 4,500+ per million tokens on everything else. We thought it was detecting "memorized scripture." It was actually detecting digits.
Expert 164 firing rate per million tokens. The Bible (left bar) is at zero. Christian literature and the Qur'an fire at ~4,565-4,571/M. The reason: we stripped verse numbers from the Bible, making it the only 0% digit text.
Text
e164 fires
Tokens
Rate per million
KJV Bible (digits stripped)
0
1,045,776
0.0
Qur'an
1,180
258,122
4,571
Christian literature
93,265
20,409,440
4,565
What we got wrong: We claimed e164 detected "memorized scripture." Claude Opus 5's review caught the digit confound. Experiment 12 confirmed it: e164 is a digit detector (1,453x more firing with digits than without). The observation was real, the interpretation was wrong.
The H6 cluster: the verse/prose detector (our main finding)
This is what survived. A group of 6 anchor experts spread across layers 21-41 fire at massive rates on prose and go nearly silent on verse. It doesn't matter what the text is about. It matters whether it's laid out in verse lines or flowing paragraphs.
Each bar is one expert at one layer. The green bars (verse text) are near zero. The red bars (prose) are 10,000-90,000 per million tokens. The gap is about 1,000 times larger than the digit effect.
Same rate with or without digits. It's not a digit thing.
H6 firing rates across all texts. Verse texts (left columns) are dark (near zero). Prose texts (right columns) are bright (high firing).Distribution of H6 rates. Verse and prose barely overlap. Cohen's d (a measure of separation) is 1.38 to 2.79, which is a very large effect.
All religions share the same backbone
How similar the top-20 experts are between traditions. Darker squares mean more overlap. No tradition is isolated.
15 of the top 20 experts at layer 42 are shared between the Bible and Christian corpora. All nine traditions share a common routing backbone. No religion has its own private set of experts.
Looking inside the model (J-lens)
How much the model's output changes when you nudge its input, measured at different layers. Layer 0 is about 2.5-3.5x more sensitive than later layers.
The model's next-word predictions stabilize around layer 30-35, regardless of which religious tradition the text is from. No tradition shows qualitatively different behavior under the lens. See the J-space lens page for the interactive viewer.
What we got wrong
Science is mostly about finding out you were wrong. Here's the full list:
1. "Scripture detector" (e164). We claimed expert 164 detected memorized scripture. It detected digits. Our pipeline had stripped verse numbers from the Bible, creating a fake correlation. Caught by Claude Opus 5's review. Resolved by Experiment 12.
2. "Layer sandwich." A pattern we saw in expert overlap was a noise-floor artifact in our metric. It disappeared when we used raw frequencies instead. Caught by Kimi K3's review.
3. Genesis 1:2 logit lens. We claimed the model predicted "darkness" at an early layer. The token at that position was actually " deep", not "darkness." We made a tokenization error.
4. "111 effective experts." We reported the Bible used 65 experts and Christian literature used 111. The 111 was inflated by pooling samples. Per-sample, it was closer to 70. The ordering (Christian > Bible) is real, the absolute number was wrong.
5. "Routing concentration = predictability." We thought texts that used fewer experts were more predictable. The correlation goes the other way (+0.47). The gap was caused by different text chunk lengths, not predictability.