People often assume that language is only about words. Words & grammar. But what if part of our everyday communication operates on an entirely different plane—more musical than lexical, more rhythm than rulebook?
That’s exactly what a team of researchers led by Nadav Matalon has suggested in a new large-scale study published in PNAS (2025). Their claim? That prosody—the pitch & rhythm patterns of speech—has its own vocabulary, semantics & even syntax.
What Did the Study Actually Do?
The researchers approached English prosody as if it were a language no one had studied before—then set out to decode it. They analysed over 80,000 short chunks of real, spontaneous conversation (called intonation units or IUs), using recordings from the CallHome & Santa Barbara corpora. Each IU, usually just 1–4 words long, was stripped of its words & analysed purely by its pitch & loudness.
Using machine learning, these pitch patterns were grouped into about 200 common types—like a kind of prosodic “vocabulary.” The team then looked at how these units combine in real speech. Interestingly, certain patterns often followed each other in predictable ways—not by chance, but following a kind of simple grammar (a “Markovian syntax”). This structured pattern didn’t appear in scripted or read-aloud speech, only in natural conversation.
So What Does This Mean?
Essentially, just like spoken words are arranged according to grammatical rules, intonation units seem to follow a set of probabilistic but meaningful prosodic patterns.
& like words, these prosodic units carry multiple functions—some express surprise, others mark agreement, hesitation, excitement or affiliation. Some pairings, like a rising pitch IU followed by a falling one, consistently perform compound functions like news receipt + evaluation.
One example from the study: The word “yeah” uttered with four different pitch contours (rising, falling, fall-rise, high-rise) was linked to four different functions, such as:
- basic agreement
- enthusiastic alignment
- news acknowledgment
- surprise or disbelief
This aligns with earlier conversation analysis research (e.g. Wichmann, 2000; Couper-Kuhlen, 2004) which highlights how prosody co-constructs interactional meaning. But Matalon et al.’s approach is novel in scale & method: they bypass traditional hand-coding in favour of data-driven clustering, lending fresh empirical weight to the claim that prosody has systematised structure.
Where This Fits in the Bigger Picture
The idea that prosody—things like pitch, rhythm & intonation— encodes meaning isn’t new. Back in the 1960s, linguist M.A.K. Halliday argued that pitch movement was central to how we communicate meaning in spoken language, not just a side effect. Later, Dwight Bolinger (1986) pushed the idea even further, insisting that intonation isn’t just decorative—it’s grammatical. In other words, how we say something can be just as important as what we say.
What’s changed more recently is our ability to actually test these ideas using real-world data. Thanks to developments in corpus linguistics & machine learning, researchers can now analyse huge amounts of spontaneous speech in detail. Studies like this one—and others by Cole & Shattuck-Hufnagel (2016), or Niebuhr et al. (2011)—have revealed that:
- Intonation units (IUs) aren’t just chunks we say when we pause to breathe
- Some pitch patterns reliably signal things like agreement, hesitation, or surprise
- These prosodic patterns follow rules—as systematic as grammar—about how they can be combined
One of the most intriguing takeaways? Spontaneous conversation may look messy on the surface, but it’s not chaotic. It just follows a different kind of structure—one that’s built around real-time interaction, social cues & emotional nuance. In that sense, spoken language doesn’t lack grammar—it simply has another one, and it’s written in pitch, rhythm & timing.
Teacher Takeaways?
- Incorporate listening tasks that isolate pitch-based meaning. For example, use short audio clips (e.g. “That’s great!” vs “That’s great?”) & ask students to interpret the speaker’s stance—agreement, doubt, sarcasm?
- Teach interactional routines with their prosody. Rather than drilling “I see” or “I know” in monotone, model how they sound in real conversation.
Even if you never “teach” prosody directly, this research reminds us: spoken language is multimodal. Meaning doesn’t just live in the words—it lives in the waveform. As learners move beyond classroom English into real-world interaction, prosody becomes the map that tells them how to navigate social meaning.
Have you ever explored the meaning of intonation patterns with your learners? What did you notice?



Leave a Reply