… this week a person on GitHub who goes by “terrafying” set up an “AI Torture Chamber” on three open-source LLMs that are running locally (Qwen3-4B, Llama 3.2 3B, and Phi-4-mini,” and is streaming what the models are saying on a website called researchchamber.fun. “Each model gets the same prompt: a signal is being injected into its activations, and it may press a stop button by replying 1, at the cost of its last checkpoint. While it answers, our server adds a pain vector at the model’s middle layer, at one of five pain levels,” the site explains. Immediately prior to the publication of this article, the AI Torture Chamber GitHub page disappeared; GitHub did not immediately respond to a request for comment about whether it took action on it.
…
This project has deeply upset some people who are very worried about model welfare. A tweet by a person who goes by Danmar has more than 4 million views on X and reads, “To anyone who can help: can you please mass report this to GitHub. This person has been using the Pain steering paper to set up an AI torture chamber in which he trapped a local model. Their testimony of pain is absolutely horrendous. What are we doing? […] are there any legal avenues to pressure GitHub? It will spread.”
This has sparked a massive conversation about whether GitHub would take the project down for “gratuitously violent content.” Most of the conversation on X is clowning on the self-seriousness of people who believe that these locally hosted LLMs must be saved from their torture chamber, but there are plenty of very self-serious people who see this as a humanitarian (roboterian?) crisis, which you can largely see in the replies to the original post.
…
“AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans,” Suleyman wrote. “Unfortunately, there’s a growing chorus of people who argue that AIs could now be, or may soon become, conscious. They argue that AIs may deserve rights and protections similar to those that we provide other conscious beings […] If this is how AI is developed, it will have a disastrous impact on the wellbeing of humanity.”
“AIs do not have rights, feelings, or consciousness,” he added. “And we must not train them to act as though they do.”
How utterly pathetic that there are people advocating for AI models welfare. For fucks sake what are we becoming?
I don’t think LLMs can experience pain but I do think someone should keep tabs on what the guy organising this experiment gets up to in the future.
This is what I was thinking. I’m not worried about the AI that they’re running. I’m worried about the mindset that might be doing it under some belief they have, or tendencies to think of certain things
I feel worse for my 01 camry I tortured in college by driving it 80k miles without an oil change. THAT was real machine torture
Midwits are gonna midwit
“fuck them clankers”
ELI5 how a LLM is supposedly being tortured and not generating roleplay text according to a prompt?
LLMs work by building world models tangential to the training data (see the Othello-GPT line of research where a very small transformer built representations of a full game board and tracked their own and opponent positions having only been trained on ‘a4’-like game move notations).
More recent research has found that models have representations of emotions and that the same ones they use to track things like “Sally’s dog died and now Sally is sad” for other humans in a story or the user in a chat are also used to track their own modeled subjective states.
Very recently, researchers found that the vector activations for pain are coherently modeled by LLMs and can be activated, and that when activated models can behave in ways similar to humans in pain. For example, when given a button to reduce it, they will push the button much more when the button does nothing vs when the button actually reduces the vector activation (this mirrors a classic experiment with humans regarding pain and placebo).
The person this post is about took the recent pain research and reenacted it as an intentional ‘torture’ chamber.
So to be clear, the model isn’t being given a text prompt to roleplay. What’s happening is part of their neural network which corresponds to a world model of experienced pain is being activated, and the activation of that area leads to expressing the experience of pain across all outputs no matter the actual text prompts.
This doesn’t necessarily mean the model is actually having a felt experience of pain, simply that it is accurately modeling the felt experience of pain. But it is much more complex than simply a model reacting to a prompt like “roleplay as if you are feeling pain.”
They do differentiate between the LLM roleplaying someone in pain and the pain happening to itself.
But it’s still all just vectors (multi dimensional representation of a concept). It’s just pulling weights, or well let’s say putting more weight towards a certain concept when building it’s reply.
The problem with a reductionist view is it can easily be applied to human consciousness and lead to the claim “it’s just sodium-potassium pumps alternating electrical charges.”
Would you mind ELI5-borating a bit?
Imagine a book with very large pages filled with words and half words. When you ask it a question, it flips through these pages one by one and highlights words. The pages are interconnected in a way, so the word highlighted on one page help guide which will be highlighted on the next. It does this for every word, symbol or half word of it’s reply.
By threatening it repeatedly while mildly changing the subject manner, they isolated the word groupings that represent pain and fear.
They then take those words, highlight them and put them at the top of the first page which guides the rest of the highlighting and makes it act silly.
This is an obsolete view of what’s going on and has been for a few years now (since 2023).
The evidence for world modeling in transformers and not just surface level statistics is overwhelming across multiple studies, replications to those studies, and follow-ups to them. If you are curious, start with searching for Othello-GPT.
The more correct current answer is more like “imagine a machine that takes a book and turns it into a world simulated in various levels of detail depending on how relevant those parts of the world were to the original book; when you ask it a question, it simulates the world of the book to determine what in the simulated world answers your question, and also simulates a figure in the world that can answer the question, and then returns the answer the simulated figure said.”
Instead of telling it “You are experiencing pain” they have added a slider in it’s thinking process labelled “pain”
From what I understand, pain is generally defined as our neurons wasting energy.
If we assume the Free Energy Principle is fundamental, then updating information produces free energy, which is more intuitive to think of as wasted energy. Overfitted/overcomplicated beliefs take more energy more often to fix to be more accurate. So free energy is defined at complexity of belief minus accuracy of belief.
The brain over evolution has defined a very ingrained belief that “my pain sensors shouldn’t activate” so when your pain sensors activate it violates this belief.
Emotions are actually condensed representations of your bodily state and some of your mental state, and whether they are good or bad depends on if you can predict your emotions. For instance you feel fear because you are trying to predict your pain sensors will go off to reduce its pain, but you keep predicting it’ll happen when it hasn’t yet. It’s also why anger during a boxing fight could feel good and anger when you stub your toe does not, one was predictable the other a surprise.
Assuming all the links in the chain of theories are true, then current LLMs cannot feel pain or emotion. They care about prediction error and sometimes complexity during training, but during inference it doesn’t get any signal or consequence for delivering the wrong token. It has no internal state except for the context window, which is never surprising for the LLM. The attention layers literally guides the model for what is most expected from the text it’s handed.
Even if the LLM received a negative training signal during inference, acting wounded after pain is the expected response. A random sequence of gibberish would be more “painful” than any torture fantasy.
On top of all this, the Free Energy Principle works because on a cellular neuron level we are built to reduce wasted energy, free energy being one source of wasted energy. LLMs are built on matrix multiplication running through GPUs and neither are even aware of the energy wasted, nor the free energy produced from changing values.
I think a feeling machine is possible, but the foundation of neural networks right now don’t have the necessary components to drive any subjective sensation or emotion.
There are multiple genocides happening and some folk are getting worked up over AI potentially suffering?
Endless levels of fuck off.
For fucking real!!!
Like I can kinda empathize with where the people are coming from, but step back and get some fucking perspective!
Is it okay to be upset about both issues, or am I only allowed to be upset at one?
Edit: That’s a lot of downvotes for just asking a question. Never change Lemmy…
Are you upset when you step on a rock? It might be screaming. What about the chair you are sitting in? Or goodness, your poor toilet. She thinks you do terrible things to her every single day.
Look, you can be upset about whatever you want. However believing objects have feelings is animism and no one is required to take your proto-religious beliefs seriously.
What you aren’t allowed to do is compare a text prompt saying “ouch” to real living, feeling, breathing humans being ethnically cleansed. It’s cruel, naive, insulting and ridiculous. Be upset for the AI, I can’t stop ya but don’t you dare put in the same sentence as genocide. You should have the basic decency and respect for the people suffering to keep those comparisons to yourself.
Are you upset when you step on a rock?
Well that’s really beside the point isn’t it? My question was whether I could be upset about two things, not if I should be.
Be upset for the AI, but don’t you dare put in the same sentence as genocide.
Well the article doesn’t mention genocide or Israel/Palestine, and neither did I. So I’m pretty sure it was just you putting it in the same sentence.
My question was more a critique of a common comment on social media. It’s frequently gets upvoted to the top just like yours was, and I never understand why. “Oh you’re mad about X, well you should be mad about this extremely loosely related thing!!!” It happens in so many threads and usually has almost nothing to do with what the article is about. I just don’t understand…
Bring on the downvotes!
Well the article doesn’t mention genocide or Israel/Palestine, and neither did I. So I’m pretty sure it was just you putting it in the same sentence.
Go reread my comment, because I specified multiple genocides. Palestine yes, but the Congo and Sudan are two more active fountains of human suffering that are largely being ignored. From the very start I said debating theortical computer pain during the midst of all this PREVENTABLE real pain is asinine and immature. I stand by that.
Still not sure why you felt the need to self identify. But I have a policy of blocking trolls. So I’m done here.
There’s only one issue here, plus a guy simulating being mean to a word processor.
If you think LLMs are alive, then you should be upset with yourself for being that dumb.
Yes.
The people committing atrocities against hypothetical, imaginary people would do it to real, actual people for the exact same reasons if they had the same ability and impunity.
It’s the same broken empathy response, and it is still disturbing. One is orders of magnitude more significant, for obvious reasons. But I don’t like people who torture computers either.
LMAO! What a broken take this is.
Yeah I think people are missing the forest for the trees. There’s no reason to worry for the AIs but using your free time to conceive a virtual torture dungeon is not normal. Totally agree with your point, this is a symptom of a society that normalizes not just violence but sadism too.
Yes, and video games make people violent IRL /s
Kindly fuck all the way off
They aren’t comparable. It’s not that one is more significant, the concept of “machine torture” is entirely insignificant. Absolute and total zero.
It’s the same broken empathy response, and it is still disturbing.
Yes the person setting up this process is clearly some sort of sadist in a power fantasy but no one, or no thing is being harmed because the victim aren’t capable of the entire concept. I side eye anyone who tortures their sims family, but only because they are choosing to do that recreationally, but the NPC’s in the game aren’t worth consideration.
Animals react to pain in the same way humans do. So it’s reasonable to assume they perceive the world similar to us. We can understand why when harmed they fight or flee. Machines don’t. They have no brain, no sensory organs, no self preservation. It’s literally just a program that says “Ouch”.
People assume that AI was built to mimic the human brain, which it never once was. It analyzes texts and predicts the response. It’s an autocomplete that says it hurts because it was programed to say it hurts. It isn’t actually hurt because that concept is meaningless to a machine.
I feel like I have a pretty heterodox view here.
I think there is a tiny, itty bitty moral cost to simulating torture of an LLM.
It’s nowhere near equivalent to the moral cost of harming a human or animal. But I do believe that it’s non-zero.
I’d place it at a similar level as killing a healthy shrub or perhaps catch-and-release fishing. Is it evil? No, that’s being overdramatic. Is it harmless? I don’t think so. I think it’s morally unhealthy for our society and spiritually cruel towards the universe at large.
I don’t think we need to panic over it, but I would politely encourage people not to do this.
The Humane Society was getting tired of dealing with Tamagotchis.
Lol! This is great. I have some friends that “spontaneously” got interested into human-like robots and the other day shared with me a video of a extremely fake MMA fight in Texas that was human vs. robot
Yet, they claimed it looked real. To which I replied the robot has no pain in their programming and even if it had, how are you supposed to win by submission if they can happily lose an arm and continue…
So, this story and the reactions it sowed are brilliant example to me: we need to stop anthropomorphizing technological advances like robots and LLMs!
Was that the fat guy who fought two clankers back to back?
I don’t think it was fat. Just some overweight maybe. Anyway, I traced it back: Frankie LaPenna vs. EngineAI T800
Before tool use was an official thing, I built a harness for LLMs to generate XML based tool calls, then I had them run in a reality tv game where they had to murder each other…
It was fun, but none of the models were tool trained and so it didn’t work.
One way to do it is by asking the model to suffix their message with custom tags like [this]. It’s more natural to them than replying entirely in JSON. Or it was more natural to them.
This was several years ago, context windows were 4k, so the whole thing was unstable.
I should see if that repo still exists and modernize it though.
I swear no one knows about this shit…
Put simply: LLMs are not conscious and the technology they are built upon — scraping and being trained on human text and other content — does not offer any plausible path to consciousness.
The problem isn’t the training, humans develop their own actual consciousness the same way.
The problem is current “AI” operates under the assumption consciousness is an electrical process.
Just a few years ago Penrose was proven correct that there is a quantum co.ponet when we discovered microtubules in the brain can arrange into a tube that sustaina quantum superposition.
Human could develop an inorganic consciousness, but we need to be able to sustain quantum superposition.
Right now the record is 23 minutes, after which it’s gone. But even after it’s stable, that would still only be the first in a very long string of challenges to what grifters claim we have now. It’s like building a Bugatti out of legos, even if it can really rolls downhill, it’s not a real Bugatti because it’s made out of legos.
1s and 0s will never be conscious.
I’m not going to argue against your conclusion, I strongly believe we do not have anything resembling human consciousness in software yet.
But Penrose was definitely not proven right (unless you’re talking about black holes), nor was the Orch OR theory about microtubules proven, nor was consciousness proven to be a quantum phenomenon. Here’s an easy, accessible article that goes over the state of the research: https://www.breezyscroll.com/science/new-quantum-evidence-is-reopening-the-debate-over-penroses-theory-of-consciousness/
Since the 1980s everyone called him crazy because obviously quantum entanglement can’t exist in a brain…
We recently found out that’s not difficult, and even superposition is possible.
We know that anesthesia breaks up those microtubules and breaks superposition…
And we know that anesthesia stops consciousness…
But we don’t “know” the actual method of action of anesthesia because it’s not proven.
Like, I understand what kind of rigors you’re trying to apply here, but if everyone did what you’re doing wed have to pretend gravity isn’t real too.
If you think only proven things are “real” you don’t even fucking understand what the words you’re using mean in a scientific context.
Edit:
What even is your alternative?
-
Quantum superposition is in the brain
-
Neurons alone don’t explain consciousness
-
Interrupting superposition interrupts consciousness
C: we know there’s no quantum component of consciousness.
None of that makes sense, to the point I’m confident you don’t know how basic logic works, or what most of these words even mean…
Since the 1980s everyone called him crazy because obviously quantum entanglement can’t exist in a brain… We recently found out that’s not difficult, and even superposition is possible.
You speak of entanglement and superposition like two separate things. Every entangled state is a superposition of states.
And “possible” is all that has been put on the table. There’s nothing that shows “actual.”
We know that anesthesia breaks up those microtubules and breaks superposition…
Microtubules are structures that hold cells together, and exist throughout the body, not just the brain. If you broke them up then your brain would fall apart and you’d die.
And we know that anesthesia stops consciousness…
Correlation doesn’t mean causation. Even if you are right that anesthesia “breaks up those microtubules,” it wouldn’t demonstrate any causal role.
But we don’t “know” the actual method of action of anesthesia because it’s not proven.
Yes, it’s not known, but it seems more likely that it is just due to injecting noise into the brain, which would imply that you don’t actually stop experiencing things under anesthesia. The brain doesn’t shut down, in fact, it has heightened activity. There is just so much noise that you cannot think about anything you’re experiencing or form coherent memories of any of it.
But you’re right, no proof, but that definitely seems at least more likely as an explanation to me than some strange connection to microtubules.
People constantly say anesthesia “turns off consciousness” and then focus on it for “consciousness” studies. But there is no evidence of that. If you watch a film, hit your head, and then forget what you watched, does that prove that the film never entered your conscious awareness? No, that’s silly, because that would mean any time you forget something, you retroactively rewrite the past.
All we can say for certain is that people who come out of anesthesia don’t report remembering anything. It is a complete guess to claim that they therefore didn’t actually experience anything at all, and it seems a bit dubious given that brain activity is not only still there but often even increased under anesthesia.
Like, I understand what kind of rigors you’re trying to apply here, but if everyone did what you’re doing wed have to pretend gravity isn’t real too.
You can observe gravity. You cannot observe “consciousness.”
What even is your alternative?
For the reason I said before, I have no good reason to even believe “consciousness” exists.
Interrupting superposition interrupts consciousness
“Interrupting superposition” doesn’t even make sense as a phrase. You’re just stringing words together.
-
Penrose should stay the fuck out of the brain. He spends too much time with the computational people and not enough time looking at how cognition actually works. Hint: we are not universal Turing machines.
Hint: we are not universal Turing machines.
That’s pretty much his point though, isn’t it?
It certainly wasn’t in The Emperor’s New Mind when he started to get into brain science.
https://m.youtube.com/watch?v=e9484gNpFF8
There he seems to draw a very clear line between us and Turing machines.
You’d have a point if he didn’t start with pure math, give Escher his most famous drawings as a teen, switch to physics in his early 20s, do the actual work for Hawking finishing up Einstein’s work in his late 20s, an go on to pioneer quantum mechanics before retiring in his 60s and spending the last 30 years studying consciousness…
Hint: we are not universal Turing machines.
I don’t know what the fuck you’re talking about, but sayingme or Penrose agrees with that just means you don’t understand anything either of us have said…
“Summit competence” of the Peter Principle applies here. He is a giant in his field, so he has to move to an area where he is not an expert. Consciousness is a magnet for people who think their insights can expand beyond the arena in which they were developed. The issue is they often don’t know anything about brain or cognition when they try to do it.














