From the DeepSeek-V4-Flash-0731 interpretability project
H6 is our name for a group of six experts that fire heavily on prose and go quiet on verse. Before this experiment we had only seen that pattern on one corpus of Bible text. We wanted to know whether the pattern holds across translations, or whether it was just an accident of whichever translation we happened to feed the model.
So we picked 30 passages spanning Genesis through Romans and rendered each one in 12 English translations, from the King James Version (1611) to the Christian Standard Bible (2017). That gives 360 records and 613,344 tokens of text in total. We stripped every verse number first, so there are zero digits in any of the text. That matters because we already knew digits confuse the picture (see Experiment 12), and we wanted to remove that variable.
They span four centuries and several translation philosophies, from ultra-literal word-for-word to loose paraphrase:
| Translation | Era | Register | License |
|---|---|---|---|
| KJV | 1611 | Archaic | Public domain |
| ASV | 1901 | Formal | Public domain |
| YLT | 1862 | Ultra-literal | Public domain |
| WEB | 2000s | Modern | Public domain |
| BBE | 1965 | Basic English | Public domain |
| NIV | 2011 | Modern | Copyrighted |
| ESV | 2001 | Formal | Copyrighted |
| NLT | 2015 | Dynamic | Copyrighted |
| NRSV | 1989 | Formal | Copyrighted |
| NASB | 1995 | Ultra-formal | Copyrighted |
| CSB | 2017 | Optimal | Copyrighted |
| MSG | 2002 | Paraphrase | Copyrighted |
The table below shows how often H6 fired, measured in "firing rate per million tokens." That number means: for every million tokens of text we feed the model, how many times did the gate pick one of the six H6 experts? A low number means H6 stayed quiet. A high number means H6 was active.
| Translation | Records | Tokens | H6 fires | H6 Rate (M) | License |
|---|---|---|---|---|---|
| KJV | 30 | 42,010 | 0 | 0.0 | Public domain |
| ESV | 30 | 57,755 | 0 | 0.0 | Copyrighted |
| NASB | 30 | 57,872 | 1 | 17.3 | Copyrighted |
| BBE | 30 | 41,434 | 1 | 24.1 | Public domain |
| ASV | 30 | 41,103 | 1 | 24.3 | Public domain |
| YLT | 30 | 43,186 | 2 | 46.3 | Public domain |
| NRSV | 30 | 57,090 | 3 | 52.5 | Copyrighted |
| WEB | 30 | 39,933 | 3 | 75.1 | Public domain |
| CSB | 30 | 57,898 | 5 | 86.4 | Copyrighted |
| NIV | 30 | 57,504 | 7 | 121.7 | Copyrighted |
| NLT | 30 | 58,493 | 10 | 171.0 | Copyrighted |
| MSG | 30 | 59,066 | 61 | 1,032.7 | Copyrighted |
| Excl. MSG | 330 | 554,278 | 33 | 59.5 |
Two integrity checks passed. First, the "invariant check," a running audit that confirms our expert-frequency numbers add up correctly, reported 0 violations across all 360 records. Second, our digit-detecting expert e164 fired only 1.6 times per million tokens across the whole corpus, with a single fire total. That confirms the digit confound was controlled: H6 is not reacting to digits here.
One translation jumps off the page. The Message (MSG) fires at 1,033/M, which is 6 to 60 times higher than any other translation. This is not a broken translation or a failure of our test. It is the strongest piece of evidence we have for the format hypothesis.
The Message is a paraphrase. Its translator, Eugene Peterson, rewrote the Bible as conversational prose, abandoning the short line-by-line verse layout every other translation keeps. From the model's perspective, MSG reads as prose, not verse. So H6 fires on it the same way it fires on any prose.
Eight of the 30 passages show H6 = 0 across all 12 translations, including MSG. Those are the most tightly verse-structured passages, the ones where even The Message keeps enough verse-like formatting that H6 stays quiet.
| Chart | What it shows |
|---|---|
| H6 by translation (12) | Bar chart, log scale. KJV and ESV at zero; MSG towers above all. |
| PD vs Copyrighted | Grouped bar. CR inflated by MSG; excl. MSG the difference is modest. |
| Per-passage heatmap | H6 rate per passage x translation. Shows which passages trigger H6. |