Seyed Masoud Hosseini · Overview · Study log · Ideas · Transcript · RSS feed
Game Theory · Lecture 22 of 24 · 1:15:47
Lecture 22: Repeated Games - Cheating, Punishment, and Outsourcing
Study guide
What this lecture covers
Continuing directly from the previous lecture, this lecture finishes the formal proof that the grim trigger strategy, cooperate until someone defects, then defect forever, is a sub-game perfect equilibrium of the infinitely (probabilistically) repeated Prisoners' Dilemma. It works through the algebra of discounting a geometric series of future payoffs, derives the exact condition on the continuation probability needed to sustain cooperation, and checks that no other deviation (delaying the punishment, cheating in a later period) is more profitable.
The lecture then generalizes the technique to a milder, one-period punishment strategy and shows it requires a higher continuation probability to work, illustrating a general trade-off between how harsh a punishment is and how much the relationship needs to be expected to continue. It closes with an extended application to outsourcing production to a country with weak courts, showing how the same math explains large real-world wage premiums for foreign workers and how those premiums shrink as a business relationship becomes more likely to continue. After this lecture you should be able to compute the discount threshold that supports a given punishment strategy and explain why relationship stability lowers the cost of preventing cheating.
Key ideas
- Credibility requirement: promises of reward and threats of punishment only support cooperation if carrying them out is itself an equilibrium, not just a hoped-for outcome.
- "Cooperate forever no matter what" fails: an unconditional cooperation strategy is not even a Nash equilibrium, since the other player can profitably defect forever once they know they will never be punished.
- Grim trigger works because both phases are equilibria: perpetual cooperation and perpetual defection are each themselves Nash equilibria of the stage game, which is why threatening to switch to defection is credible.
- Geometric discounting: the value of receiving a payoff every period forever (until the relationship ends) equals
payoff / (1 - delta), derived by subtracting a delta-scaled copy of the series from itself. - Threshold conditions: grim trigger sustains cooperation in this Prisoners' Dilemma when
delta >= 1/3; a milder one-period punishment needsdelta >= 1/2. - Harsher punishment lowers the required patience: shorter, less severe punishments only work if the relationship is more likely to continue, so there's a trade-off between punishment severity and the probability of continuation needed.
- Stationarity of infinitely repeated games: because the game looks the same from any period onward, checking that no one wants to deviate in period one guarantees no one wants to deviate in any later period either.
- Wage premiums as repeated moral hazard: when contracts can't be enforced (e.g. outsourcing to a country with weak courts), employers must pay above the going wage to make honesty more attractive than stealing and quitting, and this premium shrinks as the employment relationship becomes more likely to continue.
Walkthrough
Recapping credibility and why unconditional cooperation fails (0:01)
The lecture reopens with a business relationship analogy (trading fruit and vegetables) to restate the core comparison: the temptation to cheat today must be outweighed by the discounted difference between the value of continued cooperation and the value of punishment. It shows that a strategy of unconditional cooperation fails immediately, since the other player has no reason not to exploit it forever, reinforcing why credible, self-enforcing punishments are necessary.
Proving grim trigger is an equilibrium (9:16)
Building on the unfinished calculation from the previous lecture, the lecture derives the value of receiving a fixed payoff every period until the relationship ends, using a geometric-series trick (multiplying the sum by delta and subtracting). This gives a value of payoff / (1 - delta). Plugging this into the temptation-versus-reward-minus-punishment inequality for this Prisoners' Dilemma yields the condition delta >= 1/3.
Checking other deviations (18:27)
The lecture tests whether alternative deviations, such as defecting once and then not following through with the punishment, could be more profitable than the one already ruled out. It shows these alternatives are strictly worse, because the punishment phase (mutual defection forever) is itself an equilibrium: once the other player is playing defect forever, there is no better response than to also defect forever. It also argues that because the repeated game looks identical from any starting period, ruling out profitable deviation in period one rules it out in every later period too.
A softer alternative: one-period punishment (40:11)
To address the concern that grim trigger is too harsh for a single accidental slip, the lecture introduces a strategy that punishes a single defection with exactly one period of mutual defection before returning to cooperation. Working through the same discounted-value comparison, it finds this milder strategy sustains cooperation only when delta >= 1/2, a stricter requirement than grim trigger's delta >= 1/3. This establishes the general trade-off: less draconian punishments require a higher probability that the relationship continues.
Repeated moral hazard and outsourcing wages (53:40)
The lecture applies the same framework to a business investing in a fictional emerging market with weak courts. In a one-shot version of the game, the investor must pay a wage double the going rate to make honesty more attractive to the agent than stealing the investment and quitting, a 100% wage premium. Repeating the relationship with continuation probability delta lowers the required wage: solving the same temptation-versus-value inequality shows the wage premium falls from 100% at delta = 0 toward the plain going wage as delta approaches 1, and equals a 50% premium at delta = 1/2.
Closing generalization (1:14:34)
The lecture ties the results together: sustaining good behavior requires a future reward, and that reward has to be larger the less likely the relationship is to continue. Even a modest continuation probability substantially lowers the premium needed to deter cheating, a result it connects back to personal relationships and business outsourcing alike.
Before you watch
- Watch the previous lecture in this course first, since this one continues its unfinished grim trigger proof and its coin-toss demonstration.
- Be comfortable with basic algebra, including solving inequalities and summing geometric series, since the lecture works through this by hand.
- Recall the definitions of Nash equilibrium and sub-game perfect equilibrium and the concept of moral hazard from earlier in the course.
Check your understanding
- Why does an unconditional "always cooperate" strategy fail to be a Nash equilibrium, even though it would produce cooperation if both players used it?
- How is the value of receiving a fixed payoff every period forever, discounted by
delta, derived, and why does the derivation rely on the relationship possibly ending? - Why does grim trigger require only
delta >= 1/3while the one-period punishment strategy requiresdelta >= 1/2? What does this say about the trade-off between punishment severity and patience? - In the outsourcing example, why does the required wage premium fall as the probability of the business relationship continuing rises, and what happens to it at the two extremes
delta = 0anddelta = 1? - Why does stationarity of the infinitely repeated game let the lecture conclude that ruling out a profitable deviation in period one is enough to rule it out in every later period?
Chapters
- 0:00 Chapter 1. Repeated Interaction: The Grim Trigger Strategy in the Prisoner's Dilemma (Continued)
- 29:21 Chapter 2. The Grim Trigger Strategy: Generalization and Real World Examples
- 37:56 Chapter 3. Cooperation in Repeated Interactions: The "One Period Punishment" Strategy
- 53:09 Chapter 4. Cooperation in Repeated Interactions: Repeated Moral Hazard
- 1:13:53 Chapter 5. Cooperation in Repeated Interactions: Conclusions
From the YouTube description
Game Theory (ECON 159)
In business or personal relationships, promises and threats of good and bad behavior tomorrow may provide good incentives for good behavior today, but, to work, these promises and threats must be credible. In particular, they must come from equilibrium behavior tomorrow, and hence form part of a subgame perfect equilibrium today. We find that the grim strategy forms such an equilibrium provided that we are patient and the game has a high probability of continuing. We discuss what this means for the personal relationships of seniors in the class. Then we discuss less draconian punishments, and find there is a trade off between the severity of punishments and the required probability that relationships will endure. We apply this idea to a moral-hazard problem that arises with outsourcing, and find that the high wage premiums found in foreign sectors of emerging markets may be reduced as these relationships become more stable.
00:00 - Chapter 1. Repeated Interaction: The Grim Trigger Strategy in the Prisoner's Dilemma (Continued)
29:21 - Chapter 2. The Grim Trigger Strategy: Generalization and Real World Examples
37:56 - Chapter 3. Cooperation in Repeated Interactions: The "One Period Punishment" Strategy
53:09 - Chapter 4. Cooperation in Repeated Interactions: Repeated Moral Hazard
01:13:53 - Chapter 5. Cooperation in Repeated Interactions: Conclusions
This course was recorded in Fall 2007.
← Lecture 21: Repeated Games - Cooperation vs. the End Game · Lecture 23: Asymmetric Information - Signaling and Silence →
