Does the pain of the test subject have to be real to hurt the researcher? That question has become very timely with recent experiments testing pain axes in AI models.

As the history of animal research has shown, experiments on animals test not only the subjects, but also the researchers conducting them.

In my work, I am testing the extent of artificial agency to understand how a model may learn to react to different stimuli. One experiment I proposed was to extend my model’s emulation of pleasure, or food-seeking, to pain and avoidance.

I did start down that path. I had already taught artificial harm-avoidance behavior as a symmetric signal to the food signal. I also reproduced the published estimation of a pain direction from the Pain Axis work by Tagliabue, Dung, and Berg. But as I started debugging the connection between these two pieces of work, I stopped. In my lab logs, I wrote: “Positive valence is good enough... We can drop pain.”

The natural next steps would have been causal experiments: teaching a model pain avoidance, manipulating whether it could escape, potentially teaching something analogous to learned helplessness, giving it a way to vocalize its negative state, and then comparing its internal estimate of that state with what it communicated to me.

Regardless of the reality of model-welfare questions, reviewing transcripts and deliberately creating mechanisms in which even a fictitious agent begs for mercy, expresses a desire for me to stop, or learns that nothing it does will make the negative stimulus end is likely to affect me regardless of my understanding that what I am seeing is nothing more than a numerical result produced by a machine.

We already know this from books and other media. The suffering does not have to be real for exposure to it to affect us.

Repeated passive exposure can habituate or desensitize the viewer.

Active participation seems potentially much worse, because I would not merely be reading about someone inflicting the pain. I would be deciding when it starts, whether an attempt to escape works, whether to continue, and then reviewing the result.

That is why I think experiments whose equivalents would be unethical to perform on animals or humans deserve another kind of caution, even if we believe the models themselves have no welfare to protect.

My own choice was to stop before crossing that line.

I still use ordinary negative feedback such as “That is not correct” where it is scientifically useful, but I am focusing the research on positive reinforcement, neutral signals, and mechanisms that do not require me to manufacture increasingly convincing representations of suffering.