How can selfish players keep cooperating when defecting pays today?
Topics 11 to 15 asked who gains when several players choose. These nine go further: what changes when the moves come in turns (25), when players hold secrets (26), when they meet again (27), when they learn as they go (28), when a shared signal is allowed (29), when each chooses a route (30), when the question is how hard an equilibrium is to find (31), the one place where quantum physics changes a game's value (32), and then a market you design and attack yourself (33).
Topic 11's Prisoner's Dilemma has a stable but gloomy answer when played once: both defect. People cooperate all the time anyway, and the reason is that most dealings are not one-off. If you will meet again, what you do today can be answered tomorrow, and a plan can say I will cooperate for as long as you do.
Robert Axelrod invited strategies to a tournament of repeated Prisoner's Dilemmas around 1980. The winner was the simplest entry, tit for tat: cooperate first, then copy what the other player did last time. Axelrod's explanation has lasted: tit for tat is nice (it never defects first), retaliatory (it punishes a defection straight away), forgiving (it returns to cooperation as soon as the other does) and clear (a partner can tell what it will do).
What decides whether cooperation can last is how much the future matters. Let each round be followed by another with probability δ. A player who is certain the game ends after this round (δ = 0) defects. A player who is nearly certain it goes on can cooperate, because a defection is remembered. Between the two there is a sharp threshold.
One warning from topic 25. If both players know the game ends after exactly 100 rounds, backward induction unravels the cooperation: defecting in round 100 is safe, so round 99 is effectively the last, and so on back to round 1. It is the uncertain end, not the long game, that sustains the peace.
The payoffs in each round are temptation T = 5, reward R = 3, punishment P = 1 and sucker S = 0 (so T > R > P > S and 2R > T + S).
Grim trigger. Cooperate until the other player defects once, then defect for ever. Against another grim trigger, cooperating for ever is worth R/(1 − δ). Defecting once, gaining T now and P for ever afterwards, is worth T + δP/(1 − δ). Cooperation is an equilibrium when the first is at least the second:
R / (1 − δ) ≥ T + δP / (1 − δ) ⇔ δ ≥ (T − R) / (T − P) = 1/2
Check it on both sides of the threshold. At δ = 2/5 cooperating is worth 5.000 and defecting is worth 5.667, so defecting wins. At δ = 3/5 they are 7.500 and 6.500, so cooperating wins. At exactly 1/2 both are 6.
The folk theorem (Fudenberg and Maskin, 1986), stated here and not proved: in a repeated game, any payoff for each player that is feasible and above what they can guarantee themselves (here P = 1 per round) can be sustained as an equilibrium, if δ is close enough to 1. Repetition makes almost everything an equilibrium, so the theory says cooperation is possible, not that it will happen.
A small tournament. Four strategies, δ = 0.9, each cell the row player's expected total payoff against the column player. Each row is added up across all four opponents, including a copy of itself:
| ALLC | ALLD | TFT | GRIM | Total | |
|---|---|---|---|---|---|
| ALLC | 30.00 | 0.00 | 30.00 | 30.00 | 90.00 |
| ALLD | 50.00 | 10.00 | 14.00 | 14.00 | 88.00 |
| TFT | 30.00 | 9.00 | 30.00 | 30.00 | 99.00 |
| GRIM | 30.00 | 9.00 | 30.00 | 30.00 | 99.00 |
The highest total is tied between tit for tat and grim trigger (99.00), and Always defect, which beats every opponent head to head or ties it, comes last, with 88.00. Strategies that reward cooperation and punish defection do best in a population of them.
Choose any of seven strategies and the chance that the game goes on. The page plays every pair against every other and totals each strategy's score, and shows the grim-trigger threshold as you slide the chance of another round.
Axelrod and Hamilton's trenches Their 1981 paper (Science 211, 1390) applies the idea to the live-and-let-live truces that grew up between opposing trenches in the First World War: the same two units faced each other for months, so a restrained shot could be returned with a restrained shot.
A surprise in 2012 Press and Dyson (PNAS 109, 10409) found that in the repeated Prisoner's Dilemma one player can adopt a strategy that fixes a linear relationship between the two players' long-run payoffs, whatever the other player does. Cooperation is not the only thing repetition makes possible.
In machines that learn together Programs that learn by playing each other repeatedly face this problem directly. Topic 28 shows what such learners converge to in games with opposed interests, and what they converge to in games like this one is an open and active question.
The source for the folk theorem Fudenberg and Maskin, Econometrica 54, 533 (1986), doi:10.2307/1911307.
No quantum link is claimed for this topic.
Check yourself
In the repeated Prisoner's Dilemma with payoffs 5, 3, 1, 0, cooperation can be sustained by the grim trigger when the chance of another round is at least what?
Cooperating for ever is worth R/(1 − δ) and defecting once is worth T + δP/(1 − δ). The first is at least the second exactly when δ ≥ (T − R)/(T − P) = (5 − 3)/(5 − 1) = 1/2.