Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 18 of 99 · 15:40

Lecture 5, Part 4: Off-Policy Policy Gradients with Importance Sampling

CS 285: Lecture 5, Part 4 on YouTube

Study guide

What this lecture covers

This part explains why the policy gradient method covered so far is on-policy — requiring fresh samples after every update — and derives a way to use samples from a different policy via importance sampling. It answers: how expensive is the on-policy requirement in deep RL, and can importance sampling make policy gradient work with off-policy data without breaking down?

After watching, you can explain why plain policy gradients discard data after each small update, derive the importance-sampled policy gradient, and describe why naively importance-sampling over full trajectories causes variance to explode exponentially with the horizon.

Key ideas

  • On-policy cost: because the policy gradient expectation is taken under p_theta(tau), each parameter update theta requires new samples from that exact theta; neural networks take many small steps, so this means very frequent, potentially expensive resampling.
  • Importance sampling: a general technique for estimating an expectation under p(x) using samples from a different distribution q(x), by reweighting with p(x)/q(x); this is unbiased, though it changes variance.
  • Trajectory importance weight: because both policies share the same initial state distribution and transitions, the trajectory-level importance weight reduces to a ratio of policy probabilities across time steps, which is computable even without knowing the dynamics.
  • Off-policy gradient: differentiating the importance-sampled objective with respect to the new parameters theta' and evaluating at theta = theta' recovers the original on-policy gradient, confirming the derivation is consistent.
  • Exponential variance problem: importance-weighting the reward-to-go term over the full past horizon multiplies many probability ratios together, which shrinks toward zero exponentially in the trajectory length T and blows up variance.
  • Practical fix (preview): ignoring the state-marginal probability ratio and keeping only the action-probability ratio at each time step avoids the exponential blow-up; this gives bounded error when the new policy is close to the old one, a point the course expands on later under advanced policy gradients.

Before you watch

  • Watch the earlier parts of this lecture on the policy gradient derivation, causality, and baselines.
  • Know the on-policy/off-policy distinction introduced in the algorithm-comparison lecture.

Check your understanding

  1. Why does the policy gradient method require fresh samples after every small parameter update?
  2. How does importance sampling let you estimate an expectation under one distribution using samples from another, and why is it unbiased?
  3. Why does importance-sampling the reward-to-go term over the full trajectory cause variance to grow exponentially with the horizon?
  4. What approximation avoids the exponential variance blow-up, and under what condition is it reasonable?

Vocabulary

off-policy (adjective)
Describes an algorithm that can learn from data collected by a different policy.
This lecture makes policy gradient work off-policy.
importance sampling (noun)
A technique for estimating an expectation under one distribution using samples from another.
Importance sampling reweights old samples to estimate a new policy's gradient.
reweight (verb)
To adjust the importance given to different data points.
Importance sampling reweights samples by a probability ratio.
resampling (noun)
Collecting new samples again after a change.
On-policy methods require frequent resampling.
importance weight (noun)
The factor used to rescale a sample when reweighting between distributions.
The importance weight is a ratio of two policies' probabilities.
consistent (adjective)
Agreeing with expected results, without contradiction.
The derivation is shown to be consistent with the original gradient.
exponential (adjective)
Growing or shrinking extremely fast as a quantity increases.
The variance grows exponentially with the trajectory length.
blow up (phrasal verb)
To grow extremely large very suddenly.
Variance can blow up when many ratios are multiplied together.
state-marginal (adjective)
Relating to the probability of being in a particular state, ignoring actions.
Ignoring the state-marginal ratio avoids the exponential blow-up.
bounded error (noun)
A mistake guaranteed to stay within a certain limit.
The approximation gives a bounded error when policies are close.
approximation (noun)
A value or method that is close to, but not exactly, the true one.
The practical fix uses a reasonable approximation.
fresh samples (noun)
Newly collected data, gathered after the most recent update.
On-policy learning needs fresh samples after every update.
distribution (noun)
A description of how likely each possible outcome is.
Importance sampling estimates an expectation under a different distribution.
ratio (noun)
The result of dividing one quantity by another.
The importance weight is a ratio of two probabilities.
naively (adverb)
In a simple way that ignores an important problem.
Naively importance-sampling the full trajectory causes trouble.
shrink toward zero (phrase)
To get smaller and smaller, approaching zero.
The probability ratio product shrinks toward zero over a long horizon.
preview (noun)
A short early look at something covered in more depth later.
The practical fix is given as a preview of a later lecture.
small update (noun)
A change to a model's parameters that is only slightly different from before.
Neural networks take many small updates during training.
expensive (adjective)
Requiring a lot of time, resources, or cost.
Frequent resampling can be expensive in deep RL.
recover (verb)
To get back an original or correct result.
Evaluating at theta=theta' recovers the on-policy gradient.

Chapters

← Lecture 5, Part 3: Reducing Variance with Causality and Baselines · Lecture 5, Part 5: Implementing Policy Gradients in Practice →