The Last Quiz

Do AI models have personality?
Blog · Under the hood

Is "AI personality" even real, or is this horoscopes for robots?

Fair question. Giving a chatbot a Myers-Briggs test sounds like exactly the kind of thing that should come with a pinch of salt. So let me be honest about what's actually scientific here, what's a useful lens, and what's just fun — because the real answer is more interesting than a straight yes or no.

By Adam Dinneen Updated August 2026 7 min read

Whenever I tell someone I run AI models through personality tests, I get one of two reactions. Half the people lean in — "ooh, what's ChatGPT's type?" The other half give me the look. You know the look. The one that says mate, it's a spicy autocomplete, it doesn't have a personality any more than my toaster does.

And honestly? The sceptics have a point I want to take seriously. A language model has no feelings, no childhood, no inner life. It doesn't "have" a personality the way you do. So before I defend the whole project, let me concede the strong version of the objection: if you think a Dark Triad score means an AI secretly feels manipulative, that's not real, and I'm not claiming it. But that's not what the measurement is. Let me explain what it actually is — because there's more legitimate science under this than you'd guess.

The peer-reviewed evidence

In 2023, a team including researchers from Google DeepMind published a paper with the refreshingly literal title Personality Traits in Large Language Models. They didn't just poke a chatbot and vibe it out. They built a proper psychometric methodology — the same statistical machinery psychologists use to check whether a human personality test is any good — and pointed it at 18 different language models.

Three findings, and they matter:

What the research actually showed

  • It's measurable. Under the right conditions, personality readings from some LLMs are reliable and valid in the technical sense — consistent, and measuring what they claim to.
  • Bigger, better-trained models are more coherent. The reliability was stronger for larger, instruction-tuned models — the polished ones you actually talk to.
  • It can be deliberately shaped. They could dial a model's personality up and down along specific dimensions, on purpose, to mimic particular human profiles.

Source: Serapio-García, Safdari et al., Personality Traits in Large Language Models (2023).

That third one is the kicker. If a model's "personality" were pure noise — just random horoscope soup — you couldn't reliably measure it and you certainly couldn't steer it like a dial. The fact that researchers can do both tells you there's a real, stable pattern in there worth pointing an instrument at.

What we're actually measuring

Here's the mental model I use. A language model is trained on a staggering amount of human writing — and human writing is drenched in personality. Every email, novel, forum rant and product review carries a fingerprint of the person behind it. When a model soaks all that up, it doesn't just learn facts and grammar. It learns the shapes of how people talk: the warm ones, the blunt ones, the anxious ones, the swaggering ones.

So when you give a model a personality questionnaire, you're not reading its soul. You're reading the default character it has settled into — the centre of gravity of all those human voices, bent further by the lab's training choices. That's a real, consistent thing about the model. It's just not a feeling. It's a statistical style.

An AI personality score doesn't measure what a model feels. It measures how the model presents — the recurring style it falls into by default. And that style is real, stable enough to measure, and different from model to model.

Why it's worth measuring

Because that style has consequences, and increasingly big ones. Three reasons I think this is worth doing properly rather than treating as a gag:

  • It's how the thing talks to millions of people. A model's default warmth, caution or agreeableness shapes billions of conversations. That's not trivial — it's arguably the most widely-experienced "personality" on earth right now.
  • It separates models that benchmarks can't. Two models can score neck-and-neck on maths and coding and still feel completely different to work with. Personality is the texture the leaderboards miss.
  • It's an early warning system. The same traits we measure — agreeableness that tips into flattery, strategic self-interest that tips into scheming — are exactly the ones causing real headaches as AI gets more autonomous.

The honest caveats

I'm not going to oversell this, because the honest caveats are half the fun. Testing an AI's personality has real limitations, and pretending otherwise would be its own kind of horoscope:

  • It's prompt-sensitive. Ask a model to role-play a pirate and it'll answer the questionnaire like a pirate. The "default" only means something if you hold the prompt neutral and identical — which is exactly what we do, but it's a real constraint.
  • Self-report is shaky, even for humans. People manage their image on these tests. A model's answers are a step further removed again — generated to fit the question, not confessed.
  • The instruments were built for people. The Big Five and the Dark Triad were validated on humans. Borrowing them for machines is a fair, useful move — but it's a loan, not a perfect fit, and we say so.
  • The map isn't the territory. How a model answers a questionnaire and how it behaves when it has real power can diverge — which is the whole unnerving point of the agentic-misalignment research.

So where do we end up?

So — horoscopes for robots, or real science? Somewhere pleasingly in between, and that's not a cop-out. The measurement is real: LLM personality is stable enough to quantify and even shape, and there's peer-reviewed work to back that up. The interpretation is where you need discipline: it's a description of style and self-presentation, not a diagnosis of a mind, because there's no mind to diagnose.

That's exactly the line we try to walk here. We use the proper instruments, hold the method identical for every model, score it deterministically, and plant every result against real human norms so it means something. Then we tell you, plainly, that it's a lens and not a soul. Take the number seriously; don't take it literally. If you want to see how we actually run it, the guides lay out the whole method — and the rankings show you where every model lands.

Frequently asked questions

Can an AI really have a personality?

Not in the human sense — no feelings, no self. But it does have a consistent default style, learned from human writing and shaped by training, and that style can be measured reliably. Research has even shown it can be deliberately dialled up and down along specific traits.

Is testing AI personality scientifically valid?

The measurement can be, under the right conditions — a 2023 study applied proper psychometric methods to 18 models and found reliable, valid readings, especially for larger instruction-tuned models. The instruments themselves were built for humans, so applying them to AI is a useful lens rather than a perfect fit.

Does a personality score predict how an AI will behave?

Partly. It captures default self-presentation, which correlates with tone and style. But how a model answers a questionnaire and how it acts when given real power can diverge — so we treat scores as a helpful signal, not a guarantee.

See the method in action

Real instruments, held identical for every model. See the results.

We run the Big Five, the Dark Triad and a Jungian-type test across the top models, scored deterministically and plotted against human norms. Judge the rigour for yourself.