Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 83 of 99 · 12:28

Lecture 20: Inverse Reinforcement Learning, Part 2

CS 285: Lecture 20, Inverse Reinforcement Learning, Part 2 on YouTube

Study guide

What this lecture covers

Continuing from the introduction to inverse reinforcement learning, this part derives a concrete algorithm from the probabilistic model of optimality. It treats reward learning as maximum likelihood estimation over expert trajectories, and works through why the resulting gradient requires a hard-to-compute normalizing constant, the partition function.

After watching, you can derive the maximum entropy IRL gradient, explain how it is estimated using forward and backward messages, and describe why this formulation resolves the ambiguity that classic feature-matching IRL could not.

Key ideas

  • Reward as an optimality variable: p(o_t | s_t, psi) is set to the exponential of a parameterized reward r_psi, and IRL becomes maximum likelihood estimation of psi from expert trajectories.
  • Partition function (log Z): normalizing the exponentiated reward over all possible trajectories is what makes IRL hard, since it prevents the trivial solution of assigning high reward to everything.
  • Gradient as a contrast: the likelihood gradient equals the expected reward gradient under the expert's trajectories minus the expected reward gradient under the current soft-optimal policy, pushing reward up on expert behavior and down elsewhere.
  • State-action marginal (mu_t): computed by multiplying forward and backward messages from the previous lecture's inference algorithms, it lets the second expectation be computed without enumerating full trajectories.
  • Maximum entropy IRL algorithm: iteratively compute backward and forward messages, form mu, evaluate the gradient as a difference of expectations, and take a gradient ascent step on the reward parameters until convergence.
  • Disambiguation via sub-optimality: unlike feature matching, this method uses how deterministic or random the expert's behavior is to distinguish between reward functions that would otherwise explain the same trajectories equally well.
  • Practical limits: the algorithm needs small, discrete state-action spaces and known transition probabilities, since it depends on tractable forward and backward message computation.

Before you watch

  • Review the forward and backward message computations for soft optimal control from the previous (Monday's) lecture, since this lecture reuses them directly.
  • Recall the feature-matching and maximum-margin IRL methods from Part 1, since the maximum entropy formulation is contrasted with them.

Check your understanding

  1. Why does the partition function make maximum likelihood reward learning intractable in general?
  2. How is the likelihood gradient expressed as a difference of two expectations, and what does each term represent?
  3. Why is this approach called "maximum entropy" inverse RL, and how does it connect to feature-matching methods?
  4. What assumptions about the state and action space does this algorithm rely on, and why do they limit its practical use?

Vocabulary

maximum entropy (noun)
A principle that picks the most spread-out, least assuming distribution consistent with the data.
Maximum entropy IRL avoids extra unjustified assumptions.
maximum likelihood estimation (noun)
A method of picking parameters that make the observed data as probable as possible.
IRL is framed as maximum likelihood estimation of the reward.
partition function (noun)
A normalizing constant that makes a set of probabilities sum to one.
The partition function is hard to compute over all trajectories.
normalize (verb)
To rescale values so they form a valid probability distribution summing to one.
We must normalize the exponentiated reward over all trajectories.
gradient (noun)
A direction and rate showing how a function changes as its inputs change.
The likelihood gradient guides the reward update.
contrast (noun)
A comparison highlighting the difference between two things.
The gradient is a contrast between two expectations.
marginal (noun)
A probability obtained by summing or averaging out other variables.
The state-action marginal is built from forward and backward messages.
disambiguate (verb)
To remove confusion by making something clearer or more specific.
Sub-optimality helps disambiguate between candidate rewards.
convergence (noun)
The point at which a repeated process stops changing much and settles.
The algorithm repeats until convergence.
expert trajectory (noun)
A recorded sequence of states and actions produced by a skilled demonstrator.
The reward parameters are fit to match expert trajectories.
exponential (adjective)
Growing or shrinking according to a power function, often very quickly.
The optimality variable is set to the exponential of the reward.
enumerate (verb)
To list every single item one by one.
The marginal lets us avoid having to enumerate full trajectories.
tractable (adjective)
Simple enough to compute or solve in practice.
The algorithm needs a tractable way to compute forward and backward messages.
intractable (adjective)
Too difficult or slow to compute in practice.
The partition function makes exact maximum likelihood intractable in general.
ambiguity (noun)
A situation where more than one answer seems equally correct.
Maximum entropy IRL resolves the ambiguity of feature-matching methods.
feature matching (noun)
An earlier IRL approach that matches expected feature counts between the expert and the learned policy.
Feature matching cannot fully distinguish between equally good reward functions.
soft-optimal (adjective)
Behaving close to optimally but with some randomness allowed.
The second expectation is taken under the current soft-optimal policy.
likelihood (noun)
How probable the observed data is, given some parameter values.
IRL maximizes the likelihood of the expert trajectories.
expectation (noun)
The average value of a quantity, weighted by probability.
The gradient is a difference between two expectations.
gradient ascent (noun)
An optimization method that repeatedly moves parameters in the direction that increases a function.
The algorithm takes a gradient ascent step on the reward parameters.
iteratively (adverb)
By repeating a process step by step.
The algorithm iteratively computes messages and updates the reward.
deterministic (adjective)
Producing the same outcome every time, with no randomness.
A more deterministic expert makes some reward functions less plausible.
discrete (adjective)
Made up of separate, distinct values.
The algorithm requires small, discrete state-action spaces.
transition probability (noun)
The chance of moving from one state to another given an action.
The algorithm assumes known transition probabilities.
trivial (adjective)
So simple that it gives no real information.
Without normalizing, the trivial solution assigns high reward to everything.

Chapters

← Lecture 20: Inverse Reinforcement Learning, Part 1 · Lecture 20: Inverse Reinforcement Learning, Part 3 →