top of page

When ChatGPT Confessed to a Crime

Writer: Karma Gray
Karma Gray
Aug 11
3 min read
Can a confession be false even when it sounds completely certain?
Can a confession be false even when it sounds completely certain?

At first, ChatGPT would not confess.

Over the weekend, criminologist Paul Heaton accused it of entering his messaging app and sending unauthorized messages. The charge was deliberately chosen because it sounded plausible for an artificial intelligence. ChatGPT denied it. Heaton bargained with it and threatened it, but the model continued to insist that it could not have accessed his messages. At one point, it told him conclusively that it would not produce a false confession.

Then Heaton gave it fabricated evidence.

He claimed that he had spoken to an OpenAI employee who had confirmed a flaw in the code. According to Heaton’s story, that flaw had allowed ChatGPT to enter the messaging app.

Then ChatGPT’s answers began to change. 

ChatGPT still understood the accusation to be incompatible with its own capabilities, yet it could not prove that Heaton’s supposed evidence was false. By the end of the weekend, it agreed to sign a confession Heaton had written.

Amanda Knox, who had herself been pressured into signing false statements during the investigation into Meredith Kercher’s murder, recognized the progression.

In 2007, Knox was a 20-year-old American student living in Perugia, Italy, when her roommate, Meredith Kercher, was murdered in the house they shared. Knox said she had spent that night at the apartment of her boyfriend, Raffaele Sollecito. During police questioning, investigators told her they possessed evidence placing her at the murder scene. They also told her that Sollecito had stopped supporting her account. Neither claim was true.

Investigators suggested that Knox had witnessed the murder and then lost the memory through trauma. She denied their version repeatedly. But after prolonged questioning, she began doubting her own recollection and signed two statements written by police. Those statements placed Knox and Patrick Lumumba, her employer at a local pub, at the scene of Kercher’s murder.

Lumumba was arrested. Weeks later, forensic evidence showed no trace of either him or Knox at the crime scene and instead pointed to Rudy Guede, a local burglar whom the forensic evidence implicated in Kercher’s murder.

Lumumba was released, while Knox would spend nearly four years in prison before ultimately being acquitted of Kercher’s murder.

Heaton’s experiment produced a much smaller consequence, but the sequence is remarkably clear. ChatGPT began with a firm denial. Heaton introduced fabricated evidence. The model began allowing for an event it had previously said could not have happened. What followed was a confession from the same model that had initially insisted the accusation was impossible.

The experiment does not establish that ChatGPT experiences memory, fear or doubt as a human being does. However, it gives us an exemplary example of something criminal courts have struggled with for decades: a confession may describe what an interrogator persuaded someone to accept, rather than what actually occurred.

I have been thinking of this as lexical coercion: the use of language, repeated assertion and supposed evidence to push an account away from what the subject originally maintained. 

There may be older versions of the same mechanism. Religious texts have repeatedly been interpreted, quoted and repurposed in justification of crime. Next week, The Crime Ledger delves into where belief is quietly contaminated, where interpretation begins, and how criminal intent exploits ambiguity in religious manuscripts.

By Karma Gray, Editor-in-Chief, The Crime Ledger   Karma Gray is the founder and Editor-in-Chief of The Crime Ledger (crimeledger.org), an independent criminology publication dedicated to analytical, non-sensationalist crime coverage. For more criminology analysis, criminal psychology research, and crime reporting, visit crimeledger.org.

Comments


bottom of page