This game theory problem will change the way you see the world
Source: This game theory problem will change the way you see the world, Veritasium, 27:18, uploaded 2023-12-23, Watch Later position 1037.
On 3 September 1949, an American weather plane collected air samples over Japan and found traces of radioactive material. The Navy found the same isotopes in rainwater gathered around the world, including Cerium-141 and Yttrium-91. Their short half-lives meant that the material came from a recent nuclear explosion, although the United States had conducted no test that year. The Soviet Union had built a bomb.
The news ended the American monopoly created by the Manhattan Project. Some officials argued that the United States should attack whilst it still held the advantage. Navy Secretary Francis Matthews called the idea a plan to become “aggressors for peace”. John von Neumann, one of the founders of game theory, pressed the logic further by asking why a first strike should happen tomorrow rather than today, or at five o’clock rather than one.
The one-shot prisoner’s dilemma
RAND, the American research organisation, studied the nuclear problem in 1950. That same year, two mathematicians at RAND devised a game that resembled the emerging US–Soviet conflict. It became known as the prisoner’s dilemma.
The game gives two players a choice between cooperation and defection. Mutual cooperation pays three coins each. A player who defects against a cooperator receives five coins whilst the cooperator receives nothing. Mutual defection pays one coin each. Since defection gives the higher return against either possible move, a rational player should always defect. The other player reaches the same conclusion, so both defect and receive one coin even though cooperation would have given each three.
The structure describes a conflict in which each side can protect itself by preparing for the other’s worst move. The video applies it to the nuclear arms race, which left both superpowers with tens of thousands of weapons and a combined cost that it places at around $10 trillion. Each side had enough weapons to destroy the other many times over. Possessing them made their use harder, while the fear of falling behind kept the build-up going. Mutual restraint would have produced the better result. The individual incentive inside the game keeps pointing towards defection.
The same problem appears at a smaller scale. An impala can groom parts of its own body, yet it needs another impala to reach the remaining ticks. Grooming costs saliva, electrolytes, time and attention, and it exposes the animal to predators. A single encounter gives each animal an incentive to accept help whilst withholding its own effort. Impalas meet repeatedly, however, so one act changes what the other animal can expect next time.
Axelrod’s tournament
Robert Axelrod, a political scientist, turned the repeated problem into a computer tournament in 1980. He invited game theorists to submit programmes that would play one another. Each programme faced every other programme and a copy of itself. Axelrod described the games as lasting 200 rounds on average, with a random endpoint that kept the players from knowing the final move with certainty. The tournament ran five times so that one accidental result would carry less weight.
Axelrod received 14 entries and added a fifteenth programme called Random, which cooperated or defected half the time. The entries ranged from a rule that cooperated until the opponent defected twice in a row to 77 lines of code called Name Withheld. Friedman cooperated until the first defection and then defected for the rest of the game. Joss copied the opponent’s last move whilst defecting at random around ten per cent of the time. Graaskamp copied Joss while defecting on the fiftieth round to probe for weaknesses.
The winner was the simplest entry, Tit for Tat. It cooperates first and then copies the opponent’s previous move. A defection receives one defection in reply. Cooperation resumes as soon as the opponent returns to cooperation.
Tit for Tat and Friedman can cooperate for an entire match because neither begins with a defection. Its match with Joss develops differently. Joss defects on the sixth move, which starts an alternating series of retaliations. A later unprovoked defection turns the exchange into mutual defection for the rest of the game. Tit for Tat and Joss both score badly in that match, yet Tit for Tat cooperates with enough other entries to win the tournament.
Axelrod found four qualities among the strongest programmes. A nice programme gives the first cooperative move and waits for the other player to defect. A forgiving programme retaliates once and allows cooperation to resume when the other player changes course. A retaliatory programme responds immediately to defection, which stops a cooperative player from becoming an easy target. A clear programme makes its rule legible enough for the opponent to learn how cooperation can continue.
The first two qualities surprised the experts who had submitted intricate ways to gain an advantage. Eight of the 15 entries were nice, and the top eight all belonged to that group. The weakest nice programme still outscored the strongest nasty programme. Tit for Tat therefore combines a willingness to cooperate with a response to betrayal and a short memory after the response has done its work.
The second tournament and the missing endpoint
Axelrod circulated his analysis and then held a second tournament. This time he received 62 entries, and contestants knew what had worked in the first round. Some submitted nice and forgiving programmes, including Tit for Two Tats, which waits for two consecutive defections before retaliating. Others expected excessive forgiveness and entered programmes designed to exploit it. Tester defected on the first move, apologised if the opponent retaliated, and then followed Tit for Tat. If the opponent accepted the first attack, Tester defected on alternate moves.
The nasty programmes failed again. Tit for Tat remained the most effective entry. In the top 15, only one programme was nasty, whilst almost every programme at the bottom belonged to that group. Axelrod’s remaining qualities became clearer in this tournament: a strategy needs to retaliate without becoming a permanent enemy, and it needs a rule that the other player can understand.
The second tournament also exposed a limit. Tit for Two Tats would have won the first tournament, yet it came 24th in the second. The best strategy depends on the population it meets. Tit for Tat comes last if every opponent always defects. A strategy that works among cooperative players can become a liability in a population of bullies.
Axelrod tested that dependence with an ecological simulation. Successful strategies became more common across generations, whilst unsuccessful strategies shrank. Harrington, the only nasty entry in the top 15, grew quickly by exploiting weaker strategies. Once those strategies disappeared, Harrington’s population fell with them. After 1,000 generations, nice strategies made up the surviving population, and Tit for Tat held 14.5 per cent of it.
The simulation contains no mutations, so the source calls it ecological rather than evolutionary. Axelrod then imagined a hostile world with a small, geographically isolated cluster of Tit for Tat players. Because they mostly met one another, they could earn the benefits of cooperation inside the cluster and produce more offspring. The island could expand into the surrounding population. Cooperation could therefore begin as a local arrangement among self-interested players and spread through its own success.
Cooperation, biology, and noise
The repeated dilemma gives biologists a possible route from selfish organisms to cooperative behaviour. Impalas can groom one another, and cleaner fish can remove parasites from sharks, because repeated interaction changes the return on an act that costs something in the present. A strategy can also enter DNA. The animals need no conscious idea of trust if the inherited rule produces more successful descendants than its alternatives.
The original tournaments assumed that players correctly perceived each move. Real systems contain noise. A player may cooperate whilst the other player reads the action as defection. In 1983, the Soviet early-warning system mistook sunlight reflected from high-altitude clouds for an American missile launch. Stanislav Petrov, the officer on duty, rejected the alarm. The episode shows how one signal error can turn a cooperative action into an apparent attack when the players respond mechanically.
Noise damages Tit for Tat. Two copies begin by cooperating, yet one mistaken defection triggers retaliation. The other programme answers in kind, and the pair can fall into alternating defections. A second misread cooperation can leave them defecting for the rest of the game. The source says that each programme then earns roughly one third of the points it would have received in a perfect environment.
One answer is a more generous form of Tit for Tat that retaliates around nine times out of ten. The occasional forgiveness breaks the chain of retaliation, whilst the remaining responses still make exploitation costly. The resulting strategy keeps the basic pattern of cooperation, response and return without treating every apparent attack as proof of permanent hostility.
The world as banker
Tit for Tat can never score more than the player it faces in one match. By design, it can lose or draw. Always Defect can never lose a match, since it can only draw or win. Their overall tournament positions reverse that local picture. Tit for Tat does well because it creates high-scoring relationships across many encounters. Always Defect performs badly because it leaves cooperative rewards unused.
This distinction separates the prisoner’s dilemma from chess and poker. Those games are zero-sum because one player’s gain comes from the other’s loss. Many situations outside games work differently. The banker in the coin example supplies the reward. In life, the world supplies it. Two players can find a way to create a benefit that neither can obtain alone, then share it through continued cooperation.
The US and Soviet Union eventually used that structure for nuclear disarmament. From 1950 to 1986, both countries kept developing weapons. From the late 1980s, they began reducing their stockpiles through a sequence of smaller agreements. Each side removed a limited number of weapons, inspected the other’s actions, and continued the process after the other side had cooperated. A single promise to abolish every weapon would have concentrated the risk in one final dilemma. Repeated, verifiable steps made cooperation easier to sustain.
Later research changed the payoff structures, strategies, errors and rules of the game. Tit for Tat and generous Tit for Tat do not win in every environment. Axelrod’s four qualities remain a useful account of the conditions that let cooperation survive: begin with cooperation, answer exploitation, forgive a return, and make the rule clear.
The source ends with a qualification from Axelrod about Tit for Tat’s author, Anatol Rapoport. Axelrod asked Rapoport to submit the programme, although Rapoport wrote back that he was unsure whether it was a good idea. As a peace researcher, he may have preferred a more forgiving strategy that was less easy to provoke. The choice of Tit for Tat came from a practical experiment, and the source leaves its moral status open.
The final movement returns to agency. A player’s environment shapes which strategies succeed in the short term. Across a longer span, the players reshape the environment through the choices they make. The source treats that feedback as a reason to choose carefully when the consequences reach beyond the next move.
Limits of the account
This note follows Veritasium’s reconstruction of the prisoner’s dilemma, Axelrod’s tournaments and their applications. The figures for nuclear spending, stockpiles, simulation results and later disarmament remain claims made in the video. Its description supplies a substantial bibliography, although several journal links resolve to JSTOR or publisher pages that did not allow a full unauthenticated reading during acquisition.
The tournament results also depend on designed payoffs, a chosen population and repeated encounters. Tit for Tat performs poorly against universal defection, and noise can turn its clarity into a retaliation loop. The biological examples show how repeated interaction can support cooperation. They do not by themselves establish that every cooperative behaviour evolved through this one model. The nuclear cases carry their own history and institutions, which the game abstracts into moves and payoffs.
Further reading / references
- Robert Axelrod, A Passion for Cooperation, the interview and background page linked by Veritasium.
- The Evolution of Trust, Nicky Case’s interactive game, which the description calls a major inspiration for the video.
- Robert Axelrod, “Effective Choice in the Prisoner’s Dilemma” (1980), Journal of Conflict Resolution, and “More Effective Choice in the Prisoner’s Dilemma” (1980), Journal of Conflict Resolution.
- Robert Axelrod and William D. Hamilton, “The Evolution of Cooperation” (1981), Science.
- Robert Axelrod, The Evolution of Cooperation (1984), Richard Dawkins, The Selfish Gene (2016), William Poundstone, Prisoner’s Dilemma (1992), Martin A. Nowak and Roger Highfield, SuperCooperators (2011), and Ken Binmore, Game Theory: A Very Short Introduction (2007). These books appear in the video’s reference list without direct links.
- Prisoner’s Dilemma, Stanford Encyclopedia of Philosophy.
- J. Wu and R. Axelrod, “How to Cope with Noise in the Iterated Prisoner’s Dilemma” (1995), Journal of Conflict Resolution.
- Michael M. Flood, “Some Experimental Games” (1952), RAND research memorandum, and the video’s linked accounts of the INF Treaty, START treaties, and Stanislav Petrov.