Could an Average Person Outsmart ChatGPT in a Turing Test

By Keven Galolo·Jul 21, 2026Turing Test
Could an Average Person Outsmart ChatGPT in a Turing Test

AI systems are getting suspiciously good at lying to us online. In fact, modern software tricks human evaluators far better than most people realize.

Key Takeaways

  • GPT-4.5 with a custom persona prompt fooled judges 73% of the time in a UC San Diego study.
  • Real human participants in the same study were only recognized as human 67% of the time.
  • Models pass by injecting typos, adding delays, and acting casual rather than showing raw intelligence.
  • You can break most AI conversations using absurd physics questions or by chatting for longer than 15 minutes.
Where the Imitation Game Started

Where the Imitation Game Started

Alan Turing came up with this whole concept back in his famous 1950 paper on computing machinery and intelligence. He designed a simple text-based game to see if a machine could converse like a person.

To be honest, I used to think the imitation game was pure sci-fi fiction from an old textbook, but here we are living it daily. The setup was pretty straightforward. An evaluator types messages to two hidden participants in different rooms. One participant is a real human. The other is a computer program. The evaluator has to guess which one is the machine based purely on their answers.

Turing predicted that software would eventually use clever conversation tricks to fool us. He was spot on.

How Top AI Systems Perform

Base models usually fail miserably at this test. If you talk to a plain GPT-4o model without special instructions, you will spot it instantly. It sounds like an overly polite customer support rep writing a textbook summary.

Here is the thing though. Everything changes when researchers add a custom persona prompt.

Researchers Cameron Jones and Benjamin Bergen from UC San Diego ran a massive study on this exact topic, published on arXiv. They found that GPT-4.5 with a persona prompt tricked human judges 73% of the time. I actually love that a model needed intentional typos and slang to pass. That says so much about what we consider human behavior.

Look at the numbers from the study:

GPT-4.5 with persona: 73% judged as human Actual human participants: 67% judged as human LLaMA-3.1-405B with persona: 56% judged as human ELIZA baseline: 23% judged as human GPT-4o base model: 21% judged as human

I don't buy that a 73% score makes an AI intelligent. It’s just surprisingly good at copying how lazy we get on a keyboard.

How Top AI Systems Perform

How Software Tricks Human Judges

Passing this benchmark takes social tactics rather than raw intelligence. The software has to act flawed.

Models succeed by adding deliberate typos and irregular typing delays. They drop informal slang and regional quirks into the chat. They even pretend not to know obscure factual answers so they do not sound like Wikipedia. I find it hilarious that acting slightly ignorant is the best way to prove you are human.

If you ask a model for an opinion, it will make up fake memories about food or art to sound relatable.

Where Machine Conversation Still Breaks Down

These tests evaluate linguistic style instead of real reasoning or self-awareness.

If you keep a chat going past 15 minutes, the illusion crumbles. The software forgets its background details or starts repeating phrasing patterns.

You can also trip them up with absurd logic questions. Ask a model how to build a sturdy ladder out of warm noodle soup. Humans immediately recognize the absurdity and joke back. Software gets confused and tries to apply standard physics rules to soup.

What This Means for Our Digital Future

Natural language tools are evolving into multimodal systems with realistic voice inflections and emotional feedback.

Researchers at the Stanford Institute for Human-Centered AI (HAI) warn that these realistic personas make social engineering attacks much easier. We need better digital verification protocols fast.

Frequently Asked Questions

Did ChatGPT officially pass the Turing test?

According to research from UC San Diego cognitive scientists published on arXiv, GPT-4.5 passed a three-party Turing test by earning a human rating 73% of the time. However, researchers emphasized that the software relied on specific persona prompts to avoid giving away its identity through overly formal phrasing.

How can you spot the difference between an AI model and a real person in chat?

Evaluators suggest testing suspicious accounts with counter-factual logic puzzles, unexpected topic changes, or requests for messy, subjective personal opinions. Default model outputs tend to be structured, polite, and quick to agree compared to human responses.

Why do real humans sometimes fail the Turing test?

Judges often expect human participants to offer pristine grammar and comprehensive answers. When real people send brief, unusual, or blunt messages, evaluators frequently mistake them for primitive artificial bots.

References

  • UCSD Turing Test Study: Jones, C. R., & Bergen, B. K. Large Language Models Pass the Turing Test. arXiv:2503.23674
  • Original Turing Test Paper: Turing, A. M. (1950). Computing Machinery and Intelligence. Mind, 59(236), 433–460. DOI:10.1093/mind/LIX.236.433
  • Stanford AI Index Report: Stanford Institute for Human-Centered AI (HAI). The Artificial Intelligence Index Report. Stanford HAI


v1.6.2