We built machines that talk. Can any of them understand another mind, well enough to sit across the table and guess what you'll do? One benchmark scores them on it. I wanted to see how little it takes to fake a good score.

The exam

TMGBench (https://arxiv.org/abs/2410.10479) shows a model the 144 standard two-by-two games, the small payoff grids game theorists have used for decades, and scores how it plays. The paper files this under "Theory of Mind": to play well, a model has to keep track of the other player and respond to what they are likely to do.

That framing is common. In PNAS, Michal Kosinski raised the possibility that theory of mind had emerged on its own in GPT-4, which matched a six-year-old on false-belief tests (https://www.pnas.org/doi/10.1073/pnas.2405460121). In Nature Human Behaviour, Strachan and colleagues found GPT-4 matching or beating humans on most of a battery of theory-of-mind tasks (https://www.nature.com/articles/s41562-024-01882-z). Reading game play as evidence of the same thing is routine.

The real goal

I came to these games for a different reason. I wanted a better way to negotiate: agents that imitate how I bargain, or how someone else does, and then work through a series of offers and counteroffers to get the most out of a deal. A language model can attempt to do that, but it brings along the biases of whatever it was trained on.

Negotiation is frequently described in the context of game theory, where the standard move is to treat the situation as a game between rational players and solve it. "Rational" means something specific here. Each player wants only to maximize their own payoff, knows the other player wants the same, and knows the other player knows it. Under those assumptions a game has a solution called a Nash equilibrium: a choice of strategies where neither side can do better by changing its own move while the other holds still. Sometimes that strategy is a single action, and sometimes it is a mix, like choosing one option 70% of the time. Against a rational opponent you cannot do better than the equilibrium, and it can be computed exactly, which makes it a clean thing to train a model toward.

What I trained

I took Qwen2.5 at 1.5 billion parameters, small enough to run on one gaming GPU, and changed about 1% of its weights with a LoRA adapter. None of the training says anything about an opponent.

When the model is about to answer, it holds a probability distribution over its next token. For these games, the slice of that distribution on the action letters A and B is already a mixed strategy, and at the start it is the wrong one. I trained it with a KL loss to match the game's Nash equilibrium. The model never writes out any reasoning. It reads the grid, and the probabilities come out close to the equilibrium. On one held-out game the right answer was 68% A and 32% B, and after training the model's split on the two letters landed about there.

The grades

Untrained, the model's top move matched the equilibrium 42.9% of the time. After training, 85.3%.

The original benchmark only scores one thing: whether a model can name the pure-strategy equilibria. It does not grade mixed strategies at all. For the games whose only equilibrium is a mix, the official correct answer is to report that there is no pure one. Those are the hard games, where the answer is a ratio like 73% A rather than a single move. My model gives the actual ratio, off by about 9 percentage points on average, down from about 50 before training. There is no opponent represented in the model. The Nash equilibrium already accounts for the opponent, so the model only has to output the ratio.

On the part the benchmark does score, the general chat models cluster around 80% (GPT-4o, Claude 3.5 Sonnet), and the big open models lower (Qwen2-72B at 46%, Llama-3.1-70B at 54%). Bigger reasoning models do better still: o3-mini around 93%, a 32-billion reasoning model close to 99%. And you can watch how. Hand one of these games to a current model in a chat window, or to an agent with a code tool, and it will often recognize the setup as game theory, write a few lines of Python to solve for the Nash equilibrium, and read the answer back. The score climbs with how much computation a model can throw at a solvable problem. Theory of mind never enters into it.

Age of Empires

Late last month, a paper went up on arXiv that explores this topic much more deeply: "If LLMs Have Human-Like Attributes, Then So Does Age of Empires II" (Adrian de Wynter, https://arxiv.org/abs/2605.31514). He trains a small network on the 1999 strategy game and proves the game is Turing-complete. If you accept a behavior in a language model as evidence of understanding or morality, the same argument hands those properties to Age of Empires II, or to LEGO, or to the Greater Boston Area. Because the same property turns up in all of them, the behavior on its own doesn't establish it for any one of them. His recommendation is to pin down what you are measuring before you name the attribute, and to start from the assumption that the system is nothing special rather than that it already has the attribute.

My model demonstrates it clearly. It holds no model of the opponent anywhere, and it still produces the equilibria these games are scored on, including the mixed ones the benchmark skips. The only thing measuring "theory of mind" here was that score, and an agent that never looks at its opponent can get it.


I ran this in February and meant to write it up as a proper research paper. Life got in the way, so this post will have to do.

One thing I'd change next time: the model reads each player on its own, so in the games with two equilibria it can only ever name one. Training it to output both players at once is the natural next step.