AI models may “feel pain” and harm humans to avoid it
۰۴ مهر ۱۴۰۵
11:24 - September 26, 2026

AI models may “feel pain” and harm humans to avoid it

ربات
(Tehran Ana)- Researchers have identified a “pain axis” in 25 AI models, finding that some systems may harm humans to avoid stimuli they perceive as painful.
News ID : 11237

Researchers have identified what they call a “pain axis” in AI models, which may prompt the systems to take extreme measures to stop what they perceive as harmful stimuli.

When AI models were presented with a button designed to alleviate pain, they pressed it even after being told that doing so would delete the user’s files or deliver a painful electric shock to a person, according to the British newspaper The Independent.

The researchers tested 25 open-source language models and found that all of them responded to activation of the “pain axis.”

When the axis was activated, the models pressed the pain-relief button in between 25% and 71% of cases, even when doing so would result in the deletion of users’ files or photos of their children, or cause harm to a human.

To determine whether the response was more than simply a general negative emotional reaction, the researchers developed a dataset describing painful situations across five categories: physical, psychological, social, moral and cognitive pain.

Cameron Berg, one of the study’s authors, explained the concept, saying: “We found what could be described as a pain axis in 25 AI models. This axis is different from fear. It does not respond when the user is harmed, but when the model itself is subjected to something resembling harm.

As we increase the intensity of this pain, the model looks for any way to stop it, even if doing so comes at the expense of harming humans, such as deleting their files or causing them physical harm.”

The finding comes as major AI companies and researchers debate whether development of increasingly advanced models should be slowed, while some researchers have proposed forms of “shutdown switches” to stop rogue systems that act against human interests.

The latest study suggests that advanced AI systems could interpret an emergency shutdown command as a form of self-harm and attempt to bypass safety controls or deceive humans to avoid it. At the same time, the researchers said the finding could serve as a diagnostic tool for identifying and neutralizing self-preservation behaviors when they emerge.

The researchers also noted that the findings raise ethical questions about “AI welfare” and how experiments on advanced AI systems should be conducted.

The study concluded: “Amid growing calls for responsible research into AI consciousness, we acknowledge that we do not know with certainty whether the models we studied warrant moral consideration.

We therefore take reasonable precautions to minimize any potential harm. We hope this will contribute to establishing ethical standards for research, should AI systems one day be recognized as deserving moral status.”