tl;dr-ELT

too long; didn’t read- ELT

In ELT, corpora are nothing new- we’ve long used them to explore how language is actually used. Whether it’s the British National Corpus for identifying common collocations or learner corpora like the Cambridge Learner Corpus for spotting recurring errors, corpus work has usually focused on describing what is typical in a variety of contexts. This new study takes that idea further- using corpus comparison not just to describe language, but to uncover who produced it.

The Bible wasn’t written in a single sitting by one author- it’s the product of centuries of writing, editing & re-editing. That’s not controversial among scholars. But exactly who wrote which parts? That’s where things get tricky.

A new study in PLOS One (Faigenbaum-Golovin et al., 2025) uses a statistical method known as Higher Criticism to identify authorship based on subtle differences in word frequencies. Instead of relying solely on stylistic judgement, the authors ran a kind of computational “linguistic fingerprint” test on ancient Hebrew texts.

The study

Step 1: They identified three “ground-truth” corpora of biblical Hebrew, agreed on by most experts:

  • D (oldest layers of Deuteronomy, c. 7th century BCE)
  • DtrH (Deuteronomistic History—Joshua to Kings, 7th– 6th century BCE)
  • P (Priestly texts, late exilic or post-exilic, c. 6th– 5th century BCE)

Step 2: They lemmatised the Hebrew (grouping all forms of a word together) & built a frequency dictionary of 1,447 lemmas.

Step 3: They compared texts by calculating how much their word-use patterns deviated from each corpus using a binomial test aggregated via the Higher Criticism method.

Step 4: They tested the method’s robustness—accuracy was around 84–86%, even with short chapters (~10 verses).

The findings

  • D & DtrH were much closer to each other than to P, reflecting their shared cultural & theological background.
  • Priestly texts stood out clearly- often identifiable by specific vocabulary (e.g., “gold”, “ark”, “cubit”) tied to worship & ritual.
  • Misclassifications mostly occurred between D & DtrH, reinforcing the scholarly view that their scribes likely belonged to the same “school”.
  • The method successfully classified disputed texts like the Ark narrative: 1 Samuel 4–6 was unassignable to any corpus, but 2 Samuel 6 fit well with DtrH.

The authors of the study didn’t name individual biblical authors, rather they treated each corpus as the product of a distinct scribal “milieu” or tradition—so in practical terms, there were  three author groups.

The significance? This study offers a rare combination of transparent, statistically rigorous authorship attribution & interpretable results for ancient, multilayered texts- something that’s long challenged linguistics.

It sidesteps a major limitation of stylometry -the statistical analysis of literary style- by moving beyond broad measures like average sentence length or common word counts, instead focusing on distinctive word-frequency patterns that can work even with short, complex texts. For linguists, this opens up new possibilities for analysing authorship in everything from medieval chronicles to modern disputed documents.

Teacher Takeaways?

  • Corpus comparison works in ELT too: I guess you could adapt this idea to compare learner essays with model texts, helping students see exactly which words or structures are typical of a given genre.
  • Explainable AI is key: The study identifies the specific words that tip the scales toward one author or another. In teaching, this is the kind of transparency students need—rather than just telling them their writing “sounds informal”, show them which function words, collocations, or grammatical choices signal informality.
  • Vocabulary as a fingerprint: Even in English learning, small, repeated linguistic habits can give away a learner’s L1 background, proficiency, or exposure (e.g., preference for certain discourse markers or overuse of basic verbs like “do” & “make”). Analysing these patterns over time can help track progress & target teaching at persistent gaps.

Could frequency analysis reveal something about our students’ progress that traditional marking might miss?

Leave a Reply

Welcome to my blog

take the legwork out of reading!

There’s a lot of fascinating information out there, but sometimes we just don’t have time to find it & actually read it.
This is where this blog comes in.

I’m here to give you a summary of interesting studies, journalism & news related to the world of ELT, language learning, linguistic research & anything else that catches my eye.
I always include the link, so you can check it out for yourself.

Let’s connect
Follow tl;dr-ELT on WordPress.com

Discover more from tl;dr-ELT

Subscribe now to keep reading and get access to the full archive.

Continue reading