In ELT, corpora are nothing new- we’ve long used them to explore how language is actually used. Whether it’s the British National Corpus for identifying common collocations or learner corpora like the Cambridge Learner Corpus for spotting recurring errors, corpus work has usually focused on describing what is typical in a variety of contexts. This new study takes that idea further- using corpus comparison not just to describe language, but to uncover who produced it.
The Bible wasn’t written in a single sitting by one author- it’s the product of centuries of writing, editing & re-editing. That’s not controversial among scholars. But exactly who wrote which parts? That’s where things get tricky.
A new study in PLOS One (Faigenbaum-Golovin et al., 2025) uses a statistical method known as Higher Criticism to identify authorship based on subtle differences in word frequencies. Instead of relying solely on stylistic judgement, the authors ran a kind of computational “linguistic fingerprint” test on ancient Hebrew texts.
The study
Step 1: They identified three “ground-truth” corpora of biblical Hebrew, agreed on by most experts:
- D (oldest layers of Deuteronomy, c. 7th century BCE)
- DtrH (Deuteronomistic History—Joshua to Kings, 7th– 6th century BCE)
- P (Priestly texts, late exilic or post-exilic, c. 6th– 5th century BCE)
Step 2: They lemmatised the Hebrew (grouping all forms of a word together) & built a frequency dictionary of 1,447 lemmas.
Step 3: They compared texts by calculating how much their word-use patterns deviated from each corpus using a binomial test aggregated via the Higher Criticism method.
Step 4: They tested the method’s robustness—accuracy was around 84–86%, even with short chapters (~10 verses).
The findings
- D & DtrH were much closer to each other than to P, reflecting their shared cultural & theological background.
- Priestly texts stood out clearly- often identifiable by specific vocabulary (e.g., “gold”, “ark”, “cubit”) tied to worship & ritual.
- Misclassifications mostly occurred between D & DtrH, reinforcing the scholarly view that their scribes likely belonged to the same “school”.
- The method successfully classified disputed texts like the Ark narrative: 1 Samuel 4–6 was unassignable to any corpus, but 2 Samuel 6 fit well with DtrH.
The authors of the study didn’t name individual biblical authors, rather they treated each corpus as the product of a distinct scribal “milieu” or tradition—so in practical terms, there were three author groups.
The significance? This study offers a rare combination of transparent, statistically rigorous authorship attribution & interpretable results for ancient, multilayered texts- something that’s long challenged linguistics.
It sidesteps a major limitation of stylometry -the statistical analysis of literary style- by moving beyond broad measures like average sentence length or common word counts, instead focusing on distinctive word-frequency patterns that can work even with short, complex texts. For linguists, this opens up new possibilities for analysing authorship in everything from medieval chronicles to modern disputed documents.
Teacher Takeaways?
- Corpus comparison works in ELT too: I guess you could adapt this idea to compare learner essays with model texts, helping students see exactly which words or structures are typical of a given genre.
- Explainable AI is key: The study identifies the specific words that tip the scales toward one author or another. In teaching, this is the kind of transparency students need—rather than just telling them their writing “sounds informal”, show them which function words, collocations, or grammatical choices signal informality.
- Vocabulary as a fingerprint: Even in English learning, small, repeated linguistic habits can give away a learner’s L1 background, proficiency, or exposure (e.g., preference for certain discourse markers or overuse of basic verbs like “do” & “make”). Analysing these patterns over time can help track progress & target teaching at persistent gaps.
Could frequency analysis reveal something about our students’ progress that traditional marking might miss?



Leave a Reply