Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 85 of 99 · 14:21
Lecture 20: Inverse Reinforcement Learning, Part 4
Study guide
What this lecture covers
This final part of the inverse RL lecture reframes the guided cost learning algorithm from the previous part as a game between a reward function and a policy, and shows this is formally equivalent to a generative adversarial network (GAN). It works out what discriminator structure recovers a genuine reward function versus what structure only recovers a policy, closing out the course's coverage of inverse reinforcement learning.
After watching, you can explain the correspondence between IRL and GANs, describe the special discriminator form used in adversarial IRL methods, and distinguish reward-recovering IRL from adversarial imitation learning that only recovers a policy.
Key ideas
- IRL as a game: the reward function tries to score human demonstrations high and policy samples low, while the policy tries to make its samples score as well as the demonstrations, mirroring a GAN's generator/discriminator dynamic.
- GAN basics: a generator produces samples to resemble a true data distribution, while a discriminator is trained to distinguish real data from generated samples.
- Optimal discriminator ratio: at convergence, the Bayes-optimal GAN discriminator represents the density ratio
p_star / (p_star + p_theta), which reaches 0.5 exactly when the generator matches the data distribution. - Adversarial IRL discriminator: by substituting the optimal policy's unnormalized density (
p(tau) * exp(reward) / Z) into the discriminator formula, the resulting discriminator recovers an actual reward function while training exactly like a GAN. - Partition function as a parameter:
Zcan be optimized jointly with the reward parameters using the same objective, removing the need for the importance-sampling correction used in guided cost learning. - Generative adversarial imitation learning (GAIL): using an ordinary binary classifier as the discriminator instead of the special reward-based form recovers a working policy but not a reusable reward function.
- Transfer benefit of reward recovery: because IRL decouples the goal from the dynamics, a recovered reward can be re-optimized in modified conditions (for example, a robot missing two legs) to produce meaningful new behavior, which pure imitation cannot do.
Walkthrough
IRL viewed as an adversarial game (0:01)
The lecture opens by pointing out that the guided cost learning procedure already resembles a game: a reward function scores demonstrations high and policy samples low, while the policy adapts to fool that reward function. This observation is made formal by connecting inverse RL to generative adversarial networks, after a quick refresher on how GANs train a generator and discriminator against each other.
Deriving the adversarial IRL discriminator (4:03)
Substituting the soft-optimal policy's unnormalized density in place of the true data distribution in the optimal discriminator formula produces a discriminator with a specific form: the exponentiated reward divided by itself plus the policy's trajectory probability. Training this discriminator with the standard GAN objective trains the reward parameters directly, and optimizing the partition function Z alongside the reward removes the need for importance weights used in guided cost learning.
From reward recovery to pure imitation (9:06)
Adversarial IRL methods that use this special discriminator can recover rewards that generalize, demonstrated with an ant robot whose recovered reward still produces sensible behavior after two of its legs are disabled. The lecture then asks what happens if a regular binary classifier is used as the discriminator instead: this gives generative adversarial imitation learning (GAIL), which recovers a working policy but no longer yields a meaningful reward function, since the discriminator converges to a constant 0.5 with no informative structure left inside it.
Comparing methods and further reading (12:07)
The lecture closes by summarizing that guided cost learning and GAIL are structurally the same procedure, differing only in the discriminator's parametrization, one recovers a reward, the other only a policy, and points to suggested readings covering the classic feature-matching and maximum entropy methods alongside the modern deep IRL and adversarial imitation learning papers discussed across the lecture.
Before you watch
- Watch the earlier parts of this lecture on maximum entropy IRL and guided cost learning, since this part directly reframes guided cost learning as a GAN.
- Basic familiarity with how generative adversarial networks train a generator against a discriminator will help the derivation land.
Check your understanding
- What does the optimal discriminator in a GAN converge to, and why isn't it simply 1 for all real samples?
- How does the discriminator used in adversarial IRL differ from a standard GAN discriminator, and what does that difference buy you?
- Why can a reward recovered through IRL generalize to a modified robot body while a cloned policy cannot?
- What is the key structural difference between guided cost learning and generative adversarial imitation learning?
Vocabulary
- generative adversarial network (noun)
- A model made of two competing networks: one generates fake data, the other tries to detect it.
IRL is shown to be equivalent to a generative adversarial network. - generator (noun)
- The part of a GAN that creates new samples meant to look real.
The generator tries to produce samples resembling real data. - discriminator (noun)
- The part of a GAN that tries to tell real data apart from generated data.
The discriminator scores demonstrations high and fake samples low. - converge (verb)
- To settle toward a stable final value or state.
At convergence, the discriminator represents a specific density ratio. - density ratio (noun)
- The ratio of two probability densities compared at the same point.
The optimal discriminator represents a particular density ratio. - unnormalized (adjective)
- Not yet scaled so its total adds up to one.
The optimal policy's unnormalized density appears in the formula. - generalize (verb)
- To still work correctly in new, different situations.
A recovered reward can generalize to a modified robot body. - reusable (adjective)
- Able to be used again in a different context.
IRL produces a reusable reward function, not just a policy. - structural (adjective)
- Relating to the basic design or organization of something.
The structural difference is the discriminator's form. - partition function (noun)
- A normalizing constant that makes a set of probabilities add up to one.
The partition function Z is optimized jointly with the reward. - importance sampling (noun)
- A technique for estimating a value using samples drawn from a different distribution, with correction weights.
Optimizing Z removes the need for the importance-sampling correction. - binary classifier (noun)
- A model that sorts inputs into one of two categories.
GAIL uses an ordinary binary classifier as its discriminator. - decouple (verb)
- To separate two things so each can be handled independently.
IRL decouples the goal from the dynamics of the environment. - substitute (verb)
- To put one thing in place of another.
The derivation substitutes the policy's density into the discriminator formula. - refresher (noun)
- A short review of something already learned.
The lecture gives a quick refresher on how GANs work. - feature-matching (adjective)
- Describes an approach that tries to match statistics computed from features, rather than raw data.
Classic feature-matching methods are among the suggested readings. - maximum entropy (noun)
- A principle that picks the most spread-out, least assuming distribution consistent with known constraints.
Maximum entropy IRL was covered in an earlier part of the lecture. - reframe (verb)
- To describe or think about something from a new point of view.
The lecture reframes guided cost learning as a GAN. - trajectory (noun)
- A full sequence of states and actions taken over time.
The discriminator formula includes the policy's trajectory probability. - sensible (adjective)
- Reasonable and showing good judgment.
The disabled-leg ant still shows sensible behavior with the recovered reward.
Chapters
- 0:00 Intro
- 0:24 It looks a bit like a game...
- 1:42 Generative Adversarial Networks
- 4:19 Inverse RL as a GAN
- 9:31 Generalization via inverse RL
- 10:41 Can we just use a regular discriminator?
- 12:07 IRL as adversarial optimization
- 13:37 Suggested Reading on Inverse RL
← Lecture 20: Inverse Reinforcement Learning, Part 3 · Guest Lecture: Eric Mitchell on RLHF, Algorithms and Applications →
