People Thought GPT-4.5 Was Human
For more than 70 years, the Turing test has asked one basic question: can a computer talk well enough that people mistake it for a human? For a long time, the answer was no. Then researchers tested GPT-4.5, and people didn't just struggle to spot the AI. They actually picked it over the real human.
For more than 70 years, the Turing test has asked one basic question: can a computer talk well enough that people mistake it for a human? For a long time, the answer was no. Then researchers tested GPT-4.5, and people didn’t just struggle to spot the AI. They actually picked it over the real human.
The AI won 73% of the time
Researchers Cameron Jones and Benjamin Bergen had people chat with two strangers at once. One was a real person, and the other was GPT-4.5. After five minutes, the person had to guess which one was human.
When GPT-4.5 was given a carefully designed personality, people picked it as the human 73% of the time. But without that personality, GPT-4.5 was only picked as human 36% of the time. Same AI, huge difference.
It wasn’t just about making the model smarter. It was about making it act more like a normal person.
Being human apparently means being a little messy
The prompt told GPT-4.5 to act like a young, slightly introverted person who understood internet culture. That sounds simple, but it worked. Normal AI assistants often sound too clean. They explain too much, answer too perfectly, and rarely seem bored, awkward or unsure.
Humans are not like that. We hesitate, use slang, give incomplete answers, change subjects and sometimes just don’t know. Once GPT-4.5 started acting more like that, people trusted it more.
It wasn’t winning by sounding brilliant. It was winning by sounding ordinary.
Another AI did it too
GPT-4.5 wasn’t alone. Meta’s Llama 3.1 405B got the same kind of personality treatment, and people picked it as human 56% of the time. Older and weaker systems did much worse.
That tells us something important. People can still spot bad AI. The problem starts when the AI learns how real people actually talk.
And this isn’t only happening in chat.
Fake faces are getting harder to spot
Researchers at the Australian National University showed people real photos and AI-generated faces and asked which ones were real. You might expect people to get fooled once in a while, but something stranger happened.
For White faces in the study, AI-generated people were judged to be real more often than actual humans. Researchers called this AI hyperrealism.
The fake faces often looked smooth, familiar and average, while real people had little imperfections. Strangely, those imperfections sometimes made the real faces look more fake.
Even more surprising, some of the people who were worst at spotting fake faces were also the most confident they could do it.
AI voices are getting harder to spot too
Then researchers tested voices. In a 2025 study, people listened to real voices and AI-cloned versions of them.
When asked whether the clone belonged to the same person as the real recording, people accepted the fake as the same identity about 80% of the time. When they were simply asked whether a voice was real or AI, they only got it right about 60% of the time.
That isn’t a great margin. A photo can be fake, a chat can be fake, and now a familiar voice can be fake too. Our brains are built to recognize humans, but they were not built for software that can copy them.
Even professors struggle with AI writing
Writing isn’t much safer. A 2025 study gave 63 university lecturers short pieces of writing made by humans and AI. These are people who spend their careers reading student work, yet they still only identified about half of the AI-written passages correctly.
People keep looking for easy clues: perfect grammar, certain words, too much structure, or writing that sounds too formal. Those clues worked better with older AI, but then the models got better. Humans learned the tells, and the AI changed again.
Now the line keeps moving.
We may be looking for a human signal that isn’t there
We like to think there must be some little clue that gives AI away. A weird joke, a typo, an awkward pause, a strange opinion or an emotional reaction.
But modern AI can produce all of those things. It can even be told to sound less polished on purpose.
That may be the biggest lesson from the GPT-4.5 test. The AI didn’t win by acting like the smartest thing in the room. It won by learning how not to.
The Turing test is now a much bigger problem
Alan Turing came up with his famous test in 1950. Back then, the idea of a computer holding a believable conversation sounded almost impossible.
Now some systems can do it well enough that people choose them over real humans, and it no longer stops at text. The person in the chat might be software. The profile photo might be fake. The voice on the phone might be cloned. The writing might be generated.
The strange part is that none of it looks strange anymore.
For decades, we waited for computers to become human enough to fool us. Now they seem more human than we do. How on earth is AI outhumaning us in so many different ways?