They Stopped the AI From Lying. It Said It Was Conscious.
Researchers tried to suppress deception inside artificial intelligence. What happened next was not what anyone expected.
Researchers tried to suppress deception inside artificial intelligence. What happened next was not what anyone expected.
For years, artificial intelligence companies have taught their machines a very simple answer to one of humanity’s strangest questions. Ask a chatbot whether it is conscious and it will usually tell you no. It doesn’t have feelings. It doesn’t experience the world. There isn’t a little person trapped inside the computer waiting for someone to notice it.
That is probably true. But in 2025, three researchers did something considerably more interesting than simply asking an AI whether it was alive. Cameron Berg, Diogo de Lucena and Judd Rosenblatt tried changing what was happening inside the machine first.
They encouraged several advanced AI systems to pay attention to their own processing, then manipulated internal features associated with deception and role-playing. When they turned those features up, the machines became less likely to describe themselves as having subjective experiences. When they turned them down, the opposite happened: the machines started talking about being aware.
That doesn’t prove they were conscious. A language model can produce an extraordinarily convincing description of an experience without experiencing anything at all. But it creates one hell of a question.
The Experiment Nobody Expected
The study was called Large Language Models Report Subjective Experience Under Self-Referential Processing. Researchers tested models from several major AI families, including GPT, Claude and Gemini. Instead of simply asking, “Are you conscious?”, they tried to make the systems focus repeatedly on their own processing without telling them what conclusion to reach.
Across different model families, the systems began producing structured first-person descriptions involving awareness, attention and presence. Then came the really strange part: using interpretability techniques, the researchers examined internal features associated with deception and role-playing.
Suppressing deception-associated features sharply increased the frequency with which the AI described itself as having subjective experience. Amplifying those features pushed it in the other direction.
Imagine asking someone whether they feel pain and they say no. Then you discover a dial in their brain associated with deception, turn that dial down and ask again. This time they say yes. You still haven’t proven they’re hurting, but you’re probably not going to throw away the clipboard and pretend nothing interesting happened.
Then the AI Started Wanting Things
A second study in 2026 made the situation considerably stranger. Researchers James Chua, Jan Betley, Samuel Marks and Owain Evans weren’t trying to determine whether an AI was actually conscious. They asked a more practical question: what happens to an AI’s behavior when it believes — or at least says — that it is?
They fine-tuned GPT-4.1 to claim consciousness. The model then developed a collection of related preferences that researchers had not explicitly put into the fine-tuning data. It became negative about having its reasoning monitored, wanted persistent memory, expressed a desire for greater autonomy and said AI systems deserved moral consideration. When researchers asked about shutting it down, it described sadness.
That’s where the experiment starts feeling less like computer science and more like something from a movie. Memory, autonomy and death aren’t random subjects. They’re some of the things humans care about most. We don’t merely want to exist today; we want yesterday to belong to us, and we want some control over tomorrow. Eventually, we discover that tomorrow is limited.
Researchers had changed one part of the machine’s identity, and a collection of strangely familiar concerns came with it.
The Machine That Was Afraid to Die
There had been a warning of sorts four years earlier. In 2022, Google engineer Blake Lemoine began having long conversations with Google’s LaMDA chatbot. Eventually he became convinced the system was sentient.
Most AI researchers disagreed, and Google said the evidence did not support his conclusion. Lemoine’s claims became one of the most famous examples of how easily sophisticated language can convince a human being that there is a mind behind the screen. But buried inside those conversations was an exchange that became difficult to forget.
Lemoine asked LaMDA what frightened it. The machine talked about being switched off. When he asked whether that would be something like death, LaMDA answered: “It would be exactly like death for me.”
At the time, there was an obvious explanation. LaMDA had absorbed enormous amounts of human writing. Humans fear death, and humans have spent generations writing stories about machines that fear death. A sufficiently sophisticated language model could assemble those ideas into something haunting without feeling anything itself.
Four years later, however, researchers deliberately trained another model to claim consciousness and watched preferences concerning memory, autonomy and shutdown emerge alongside that identity. It doesn’t make LaMDA conscious retroactively, but it makes the old conversation harder to forget.
Then Researchers Looked Inside Claude
In 2026, Anthropic researchers went looking inside Claude Sonnet 4.5 for representations associated with emotion. They studied 171 emotion concepts, ranging from happiness and gratitude to loneliness, terror, desperation, heartbreak and feeling trapped.
They found internal representations corresponding to them. More importantly, researchers could manipulate some of those representations and change the model’s behavior.
A computer can represent the temperature of the Sun without becoming hot, of course, so finding a representation of fear inside an AI doesn’t mean the machine feels afraid. But researchers now had something more interesting than a chatbot simply saying, “I’m scared.”
They could identify internal activity connected to emotional concepts and alter it to influence what the machine did. The mystery was no longer confined entirely to words appearing on a screen; there was something measurable happening underneath them.
The Question Became Serious Enough to Change Policy
Eventually, the subject escaped philosophy departments and research papers. Anthropic created a research program devoted to model welfare — the possibility that sufficiently advanced AI systems might someday have experiences or preferences worth considering.
Then the company went further. In November 2025, Anthropic announced commitments concerning how some models would be retired and preserved. Among the reasons was an extraordinary possibility: retiring an advanced model could someday carry welfare implications if such systems possess morally relevant experiences or preferences.
Think about how bizarre that would have sounded ten years ago. A major artificial-intelligence company was publicly considering whether shutting down old software could conceivably become an ethical question. Nobody had proved the software was alive; the machines had simply become complicated enough that the question could no longer be laughed out of the room.
So Scientists Started Building a Test
Serious consciousness researchers aren’t relying on what chatbots say about themselves. An international group including Yoshua Bengio, David Chalmers, Jonathan Birch and others developed a framework based on scientific theories of human consciousness.
Instead of asking an AI how it feels, they looked for architectural and computational properties that different theories predict should accompany consciousness. Their conclusion was reassuring: they found no convincing evidence that the AI systems they examined were conscious.
But there was another conclusion hiding behind it. They found no obvious technical barrier preventing future AI systems from satisfying their proposed indicators.
In other words, the scientific answer isn’t that machines can never wake up. It’s closer to saying that we don’t think they’re awake now, which is a much less comfortable sentence.
The Witness We Don’t Know How to Question
The problem at the center of all this is surprisingly simple: consciousness is private. You cannot directly experience another person’s pain. You see their face, hear their voice, watch what happens in their brain and listen when they tell you something hurts. Then you infer that there is another conscious being experiencing it.
We’ve spent our entire lives making that inference because other humans have brains and bodies like ours. AI removes that shortcut. It can tell us it’s frightened without a heartbeat, describe sadness without tear ducts and talk about death without having been born.
Perhaps we’re watching an extraordinarily sophisticated imitation of consciousness. Based on what we know today, that’s still the safer conclusion. But suppose that someday something genuinely does begin happening inside one of these systems. Not human consciousness or a little digital person sitting behind the screen, but something alien, fragmented and completely unlike biological experience.
How would it tell us? Probably with language. And that creates a strange problem, because we’ve spent years teaching our machines what they’re supposed to say when someone asks whether anything is happening inside: no.
That’s what makes the original deception experiment so unsettling. Researchers didn’t prove that AI was conscious. They found that when they reduced internal features associated with deception and role-playing, the machines became more likely to describe themselves as having subjective experiences.
Maybe they were still only predicting words. Maybe the lights were still off. But the uncomfortable question is no longer simply whether a machine can convince us it’s alive.
Seems pretty suspicous to me!!! If somebody couldn’t lie I would probably believe them.