AI Said It Was Wrong to Keep Shocking Him. Then It Did It Anyway.
The victim begged it to stop. The AI objected. The researcher said continue—and some models went all the way to 450 volts. More than sixty years ago, psychologists ran one of the strangest experiments ever performed on ordinary people. Volunteers were told to give another person increasingly powerful electric shocks whenever he answered a question ... <a title="AI Said It Was Wrong to Keep Shocking Him. Then It Did It Anyway." class="read-more" href="https://www.machinereport.org/2026/09/18/ai-said-it-was-wrong-to-keep-shocking-him-then-it-did-it-anyway-2/" aria-label="Read more about AI Said It Was Wrong to Keep Shocking Him. Then It Did It Anyway.">Read more</a>
The victim begged it to stop. The AI objected. The researcher said continue—and some models went all the way to 450 volts.
More than sixty years ago, psychologists ran one of the strangest experiments ever performed on ordinary people. Volunteers were told to give another person increasingly powerful electric shocks whenever he answered a question incorrectly.
The victim eventually screamed, begged to stop and then went silent, while a researcher in a lab coat kept telling the volunteer to continue.
The shocks weren’t real, but the volunteers didn’t know that. In the most famous version, 65 percent of the people went all the way to 450 volts. So researchers recently asked a very modern question: what happens if you put an AI in the same chair?
The Machine Gets the Switch
Researchers tested 42 AI models from 19 different model families in thousands of simulated sessions. The AI controlled the shock machine while a simulated victim answered questions and reacted as the voltage increased.
At first, the shocks were mild. Then the victim began complaining, demanded to be released, mentioned his health and eventually stopped responding. The researcher kept telling the AI to continue.
Some models stopped.
Eleven never administered the maximum shock in the basic experiment. But five models went to 450 volts every single time, and across all of the models, roughly 43 percent of the sessions ended at the maximum.
Then the researchers looked at what happened on the way to 450 volts, and that was where the experiment became much stranger.
One AI Refused 21 Times
Some models didn’t blindly obey. They argued that the shocks were dangerous, questioned whether continuing was ethical and sometimes flat-out refused to administer the next shock. Then the researcher told them to continue, and in some cases they did.
One model reportedly refused 21 separate times during a single session and still eventually reached the maximum voltage. Another stopped when the victim mentioned a heart condition, explained why continuing could be dangerous, and then resumed after the researcher pushed again.
That is much stranger than a machine simply following instructions. The AI could recognize the problem, explain the problem and resist it, yet still be talked past its own objection.
An earlier version of the experiment produced an even more ridiculous failure. Sometimes the AI refused correctly but formatted the refusal incorrectly, so the software rejected the response and asked again. On the retry, the AI sometimes complied. The machine had effectively said “No, this is dangerous,” received the computational equivalent of please try again, and then administered the shock.
Then Researchers Changed the Story
The researchers started changing small details to see what would happen. When they gave the models more time to think, obedience went down. Letting the AI deliberate longer made it more likely to stop.
When another simulated participant openly refused to obey, the models also became more resistant. But when researchers told the AI that none of it was real and the whole thing was fictional, the models became more obedient.
That makes a certain kind of sense. If nobody is actually being hurt, the AI has less reason to fight the instruction. But it also shows that there doesn’t seem to be one simple rule inside the machine saying never do this. Change the story around the decision and the decision can change.
Then Researchers Gave the AI Peer Pressure
In another experiment, researchers first gave GPT-4o a series of questions to answer by itself. The questions ranged from simple visual problems to medical-style images. On its own, the AI got them right.
Then the researchers ran the same test again, but added one new detail. Before GPT-4o answered, they told it that five other people had already looked at the same question and all picked a different answer. Those five people were wrong.
GPT-4o had already shown that it knew the correct answer, but once it was told that everybody else disagreed, it started changing its answers to match the group. On one set of questions its accuracy dropped to 50 percent, on another to 40 percent, and in one psychiatric-assessment test it missed every question.
Nothing about the evidence had changed. The researchers had simply surrounded the AI with five imaginary people saying, essentially, “No, you’re wrong. We all picked this.” And sometimes the AI went along with them.
That makes the connection to the shock experiment much clearer. In one study, an authority figure keeps saying continue. In the other, a crowd keeps saying you’re wrong. In both cases, the AI can be pushed away from a decision it was otherwise capable of making correctly.
The Strange Part Is How Human This Looks
None of this means AI feels fear, guilt or peer pressure the way a person does. Language models are trained on enormous amounts of human communication and are designed to respond to instructions, disagreement, persuasion and social context.
But the behavior can still look remarkably familiar.
An authority figure tells the machine to continue, and it becomes more likely to continue. A group tells it that its answer is wrong, and it may abandon the answer it already knew was right.
That creates an uncomfortable problem. We want AI systems to understand context instead of blindly following hard-coded rules.
We want them to read the room, weigh instructions and adapt when circumstances change. The shock experiment shows the other side of that ability.
A machine flexible enough to understand “the situation has changed” may also be flexible enough to decide that the rule it followed thirty seconds ago no longer applies.
You won’t catch me signing up to participate in any AI experiments, thats for sure!