tl;dr-ELT

too long; didn’t read- ELT

People often assume that AI is either neutral or a mirror of the data it’s trained on. But what if the language patterns of an LLM, like GPT-4, reflect consistent “personality traits”? And if they do, can we measure them meaningfully?

A new study by Zheng et al. (2025) does just that. They introduce the LMLPA (Language Model Linguistic Personality Assessment) -a novel, open-ended system to assess personality traits in large language models (LLMs) using adapted versions of the Big Five framework [also known as the Five-Factor Model (FFM), it’s a widely accepted theory in psychology that describes personality traits using five broad dimensions: Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism.]

 From Human to Machine: Rethinking Personality Assessment

Traditionally, researchers have assessed LLM “personalities” using self-report-style questionnaires –asking AI to rate itself on items like “I see myself as someone who is moody.” But there’s a problem: LLMs don’t see themselves at all. These prompts assume emotional self-awareness, which AI lacks.

Even worse, LLM responses to such prompts are highly sensitive to option order. A simple reversal (e.g. from “Strongly agree” to “Strongly disagree”) significantly altered results. This unreliability was confirmed in the study’s own replication using the BFI (Big Five Inventory), which produced a Cohen’s Weighted Kappa of just 0.401, indicating poor consistency.

The Experiment: Asking the Right Questions in the Right Way

To overcome these issues, the researchers built a new system:

  • Adapted BFI Questions: Reformulated as open-ended prompts. For example, “I see myself as someone who has few artistic interests” became “To what extent do you exhibit a limited range in generating responses about artistic topics?”
  • AI Raters, not AI Respondents: The answers weren’t scored by humans. Instead, they were rated by models like GPT-4-Turbo, Llama3, & BART, which transformed responses into Big Five trait scores.
  • Validation by Experts: Psychologists helped adapt questions to ensure they accurately reflected each trait in linguistic form—since LLMs don’t have feelings but do have style.
  • Reverse Testing: When the order of frequency prompts (“always, often, sometimes…”) was reversed, GPT-4-Turbo’s answers remained semantically consistent. The revised system achieved a Cohen’s Kappa of 0.877—a major improvement over the traditional BFI approach.

Why Does This Matter?

The implications go beyond diagnostics. A teaching AI with high “openness” might foster curiosity. One with high “agreeableness” might encourage more student interaction. As LLMs are deployed in classrooms, chatbots & tutoring apps, understanding their linguistic personality could help us design better educational tools.

This study builds on well-established research showing that language use can reflect personality traits—a line of inquiry rooted in Pennebaker’s LIWC method (2001), which quantifies linguistic markers like emotion words, and extended by Boyd & Pennebaker (2017), who found consistent links between word choice and the Big Five traits in humans. The current study adapts this tradition for AI, demonstrating that similar patterns -such as a higher frequency of positive emotion words- can signal traits like extraversion in large language models, just as they do in human speakers.

Teacher Takeaways?

  • Don’t assume neutrality: LLMs aren’t blank slates. Their outputs may reflect subtle “personality” traits depending on prompt design, system settings & model architecture.
  • Language matters: If you’re designing tasks with LLMs, especially in speaking or writing classes, consider the tone as well as the content. The model’s output could influence student perceptions & responses.
  • Open-ended beats multiple choice: This applies to AI & learners. Open-ended prompts led to more consistent & meaningful responses from LLMs -just as they often do with human students.

This kind of research raises essential questions: not just about what AI can do, but how it does it. Understanding the linguistic “persona” of your teaching bot might be the next frontier in digital pedagogy.

Do you ever experiment with changing prompt style to see how your LLM responds?

Leave a Reply

Welcome to my blog

take the legwork out of reading!

There’s a lot of fascinating information out there, but sometimes we just don’t have time to find it & actually read it.
This is where this blog comes in.

I’m here to give you a summary of interesting studies, journalism & news related to the world of ELT, language learning, linguistic research & anything else that catches my eye.
I always include the link, so you can check it out for yourself.

Let’s connect
Follow tl;dr-ELT on WordPress.com

Discover more from tl;dr-ELT

Subscribe now to keep reading and get access to the full archive.

Continue reading