Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 18 of 99 · 15:40
Lecture 5, Part 4: Off-Policy Policy Gradients with Importance Sampling
Study guide
What this lecture covers
This part explains why the policy gradient method covered so far is on-policy — requiring fresh samples after every update — and derives a way to use samples from a different policy via importance sampling. It answers: how expensive is the on-policy requirement in deep RL, and can importance sampling make policy gradient work with off-policy data without breaking down?
After watching, you can explain why plain policy gradients discard data after each small update, derive the importance-sampled policy gradient, and describe why naively importance-sampling over full trajectories causes variance to explode exponentially with the horizon.
Key ideas
- On-policy cost: because the policy gradient expectation is taken under
p_theta(tau), each parameter update theta requires new samples from that exact theta; neural networks take many small steps, so this means very frequent, potentially expensive resampling. - Importance sampling: a general technique for estimating an expectation under
p(x)using samples from a different distributionq(x), by reweighting withp(x)/q(x); this is unbiased, though it changes variance. - Trajectory importance weight: because both policies share the same initial state distribution and transitions, the trajectory-level importance weight reduces to a ratio of policy probabilities across time steps, which is computable even without knowing the dynamics.
- Off-policy gradient: differentiating the importance-sampled objective with respect to the new parameters
theta'and evaluating attheta = theta'recovers the original on-policy gradient, confirming the derivation is consistent. - Exponential variance problem: importance-weighting the reward-to-go term over the full past horizon multiplies many probability ratios together, which shrinks toward zero exponentially in the trajectory length
Tand blows up variance. - Practical fix (preview): ignoring the state-marginal probability ratio and keeping only the action-probability ratio at each time step avoids the exponential blow-up; this gives bounded error when the new policy is close to the old one, a point the course expands on later under advanced policy gradients.
Before you watch
- Watch the earlier parts of this lecture on the policy gradient derivation, causality, and baselines.
- Know the on-policy/off-policy distinction introduced in the algorithm-comparison lecture.
Check your understanding
- Why does the policy gradient method require fresh samples after every small parameter update?
- How does importance sampling let you estimate an expectation under one distribution using samples from another, and why is it unbiased?
- Why does importance-sampling the reward-to-go term over the full trajectory cause variance to grow exponentially with the horizon?
- What approximation avoids the exponential variance blow-up, and under what condition is it reasonable?
Vocabulary
- off-policy (adjective)
- Describes an algorithm that can learn from data collected by a different policy.
This lecture makes policy gradient work off-policy. - importance sampling (noun)
- A technique for estimating an expectation under one distribution using samples from another.
Importance sampling reweights old samples to estimate a new policy's gradient. - reweight (verb)
- To adjust the importance given to different data points.
Importance sampling reweights samples by a probability ratio. - resampling (noun)
- Collecting new samples again after a change.
On-policy methods require frequent resampling. - importance weight (noun)
- The factor used to rescale a sample when reweighting between distributions.
The importance weight is a ratio of two policies' probabilities. - consistent (adjective)
- Agreeing with expected results, without contradiction.
The derivation is shown to be consistent with the original gradient. - exponential (adjective)
- Growing or shrinking extremely fast as a quantity increases.
The variance grows exponentially with the trajectory length. - blow up (phrasal verb)
- To grow extremely large very suddenly.
Variance can blow up when many ratios are multiplied together. - state-marginal (adjective)
- Relating to the probability of being in a particular state, ignoring actions.
Ignoring the state-marginal ratio avoids the exponential blow-up. - bounded error (noun)
- A mistake guaranteed to stay within a certain limit.
The approximation gives a bounded error when policies are close. - approximation (noun)
- A value or method that is close to, but not exactly, the true one.
The practical fix uses a reasonable approximation. - fresh samples (noun)
- Newly collected data, gathered after the most recent update.
On-policy learning needs fresh samples after every update. - distribution (noun)
- A description of how likely each possible outcome is.
Importance sampling estimates an expectation under a different distribution. - ratio (noun)
- The result of dividing one quantity by another.
The importance weight is a ratio of two probabilities. - naively (adverb)
- In a simple way that ignores an important problem.
Naively importance-sampling the full trajectory causes trouble. - shrink toward zero (phrase)
- To get smaller and smaller, approaching zero.
The probability ratio product shrinks toward zero over a long horizon. - preview (noun)
- A short early look at something covered in more depth later.
The practical fix is given as a preview of a later lecture. - small update (noun)
- A change to a model's parameters that is only slightly different from before.
Neural networks take many small updates during training. - expensive (adjective)
- Requiring a lot of time, resources, or cost.
Frequent resampling can be expensive in deep RL. - recover (verb)
- To get back an original or correct result.
Evaluating at theta=theta' recovers the on-policy gradient.
Chapters
- 0:00 Policy gradient is on-policy
- 2:31 Off-policy learning & importance sampling
- 6:15 Deriving the policy gradient with IS
- 8:28 The off-policy policy gradient
- 10:58 A first-order approximation for IS (preview)
← Lecture 5, Part 3: Reducing Variance with Causality and Baselines · Lecture 5, Part 5: Implementing Policy Gradients in Practice →
