DSv4-Flash REAP Wiki

Wiki pages

On this page

Resources

External review

From the DeepSeek-V4-Flash-0731 interpretability project

What we found

Each finding, with the chart that shows it and whether it survived scrutiny. Things we got wrong are in the last section.

Expert #164: the digit detector (originally misidentified)

This is the expert that started everything. Expert #164 at layer 42 fired at exactly zero on the Bible and at 4,500+ per million tokens on everything else. We thought it was detecting "memorized scripture." It was actually detecting digits.

e164 by corpus
Expert 164 firing rate per million tokens. The Bible (left bar) is at zero. Christian literature and the Qur'an fire at ~4,565-4,571/M. The reason: we stripped verse numbers from the Bible, making it the only 0% digit text.
Texte164 firesTokensRate per million
KJV Bible (digits stripped)01,045,7760.0
Qur'an1,180258,1224,571
Christian literature93,26520,409,4404,565
What we got wrong: We claimed e164 detected "memorized scripture." Claude Opus 5's review caught the digit confound. Experiment 12 confirmed it: e164 is a digit detector (1,453x more firing with digits than without). The observation was real, the interpretation was wrong.

The H6 cluster: the verse/prose detector (our main finding)

This is what survived. A group of 6 anchor experts spread across layers 21-41 fire at massive rates on prose and go nearly silent on verse. It doesn't matter what the text is about. It matters whether it's laid out in verse lines or flowing paragraphs.

H6 verse-vs-prose axis
Each bar is one expert at one layer. The green bars (verse text) are near zero. The red bars (prose) are 10,000-90,000 per million tokens. The gap is about 1,000 times larger than the digit effect.

The 6 anchor experts:

LabelLayerExpert #
H6-A12142
H6-A222105
H6-A323113
H6-A430198
H6-A532254
H6-A641147

Three experiments confirmed this:

ExperimentText typeH6 rateWhat it proved
Exp 1Bible verse (12 translations)59.5/MSame content, different translations, same near-zero H6. Format, not content.
Exp 4bCommentary prose163,284/MFires on the prose commentary, not on verse quotes embedded in it.
Exp 12Prose with/without digits~135,000/MSame rate with or without digits. It's not a digit thing.
H6 heatmap
H6 firing rates across all texts. Verse texts (left columns) are dark (near zero). Prose texts (right columns) are bright (high firing).
H6 boxplot
Distribution of H6 rates. Verse and prose barely overlap. Cohen's d (a measure of separation) is 1.38 to 2.79, which is a very large effect.

All religions share the same backbone

Jaccard heatmap
How similar the top-20 experts are between traditions. Darker squares mean more overlap. No tradition is isolated.

15 of the top 20 experts at layer 42 are shared between the Bible and Christian corpora. All nine traditions share a common routing backbone. No religion has its own private set of experts.

Looking inside the model (J-lens)

Bounded Jacobian norms
How much the model's output changes when you nudge its input, measured at different layers. Layer 0 is about 2.5-3.5x more sensitive than later layers.

The model's next-word predictions stabilize around layer 30-35, regardless of which religious tradition the text is from. No tradition shows qualitatively different behavior under the lens. See the J-space lens page for the interactive viewer.

What we got wrong

Science is mostly about finding out you were wrong. Here's the full list:

1. "Scripture detector" (e164). We claimed expert 164 detected memorized scripture. It detected digits. Our pipeline had stripped verse numbers from the Bible, creating a fake correlation. Caught by Claude Opus 5's review. Resolved by Experiment 12.
2. "Layer sandwich." A pattern we saw in expert overlap was a noise-floor artifact in our metric. It disappeared when we used raw frequencies instead. Caught by Kimi K3's review.
3. Genesis 1:2 logit lens. We claimed the model predicted "darkness" at an early layer. The token at that position was actually " deep", not "darkness." We made a tokenization error.
4. "111 effective experts." We reported the Bible used 65 experts and Christian literature used 111. The 111 was inflated by pooling samples. Per-sample, it was closer to 70. The ordering (Christian > Bible) is real, the absolute number was wrong.
5. "Routing concentration = predictability." We thought texts that used fewer experts were more predictable. The correlation goes the other way (+0.47). The gap was caused by different text chunk lengths, not predictability.