Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 51 of 99 · 15:04
Lecture 12, Part 2: Model-Based RL with Policies
Study guide
What this lecture covers
This part continues the model-based reinforcement learning discussion by asking why you would use a model-free algorithm, like the policy gradient, on top of a learned model instead of backpropagating through it directly. It builds on the previous part's derivation of the backpropagation (pathwise) gradient and contrasts it with the likelihood-ratio (policy gradient) estimator.
The lecture then turns to a core practical problem: rollouts through a learned dynamics model accumulate error the longer they run, for the same distributional-shift reasons seen earlier in the course with imitation learning. This sets up the motivation for the short-rollout designs used by modern model-based RL algorithms, covered in the next part.
Key ideas
- Likelihood ratio (policy gradient) estimator: does not require differentiating through the transition dynamics, only through sampling; it is valid for stochastic policies and transitions.
- Pathwise (backpropagation) gradient: computes exact derivatives via the chain rule but requires multiplying a long product of Jacobians across time steps, which can explode or vanish.
- Sampling trade-off: the policy gradient avoids the Jacobian product but needs many samples; with a model, samples are cheap compute rather than costly real-world interaction.
- Model-based RL version 2.5: fit a model, generate many trajectories from it, and improve the policy with the policy gradient instead of backpropagation.
- Distributional shift in rollouts: as in imitation learning, small model errors push the policy into unfamiliar states where errors compound, growing roughly with the square of the rollout length.
- Branched short rollouts: sampling states from real trajectories and rolling out only briefly from each reduces compounding error while still covering later time steps.
- On-policy vs. off-policy: because branched rollouts mix data from the collecting policy and the latest policy, off-policy algorithms like Q-learning handle this state-distribution mismatch better than on-policy policy gradient methods.
Before you watch
- Watch the first part of Lecture 12, which introduces the backpropagation gradient for model-based RL.
- Review the distributional shift argument from the imitation learning lectures, since it is reused here.
Check your understanding
- Why does the pathwise gradient become numerically unstable for long horizons, while the policy gradient does not?
- What trade-off do you accept when switching from backpropagation to the policy gradient for training with a learned model?
- Why does branching short rollouts from real-world states reduce accumulated error compared to full-length model rollouts?
- Why are off-policy methods generally preferred over on-policy policy gradient methods for training on these branched rollouts?
Vocabulary
- likelihood ratio (phrase)
- A ratio comparing how likely an outcome is under two different distributions.
The policy gradient is a likelihood ratio estimator. - pathwise gradient (phrase)
- A gradient computed by directly differentiating through the whole computation, sample by sample.
The pathwise gradient uses the chain rule through the dynamics. - stochastic (adjective)
- Involving randomness rather than a fixed outcome.
The estimator is valid for stochastic policies and dynamics. - differentiate through (phrase)
- To compute a gradient that passes through a certain part of a computation.
The pathwise gradient differentiates through the dynamics. - trade-off (noun)
- A balance where gaining one benefit costs you another.
There's a trade-off between the two gradient estimators. - compute (resource) (noun)
- Processing power or machine resources used to run calculations.
With a model, extra samples cost cheap compute. - distributional shift (phrase)
- A mismatch between the data used for training and the data actually encountered.
Rollouts through the model suffer from distributional shift. - accumulate error (phrase)
- To build up mistakes over time, becoming larger with each step.
Small model errors accumulate error over a long rollout. - branch (v) (verb)
- To start a new short path from a specific point.
We branch short rollouts from real states. - cover (later steps) (verb)
- To include or account for something within a range.
Branched rollouts still cover later time steps. - on-policy (adjective)
- Only valid when data comes from the exact same policy being trained.
On-policy methods struggle with mixed rollout data. - off-policy (adjective)
- Able to learn from data collected under a different policy.
Off-policy methods like Q-learning handle mixed data better. - mismatch (noun)
- A difference between two things that should ideally match.
There's a state-distribution mismatch in branched rollouts. - numerically unstable (phrase)
- Prone to producing wildly incorrect results due to calculation errors.
The pathwise gradient becomes numerically unstable for long horizons. - generate (trajectories) (verb)
- To produce new sequences of states and actions from a model.
We generate many trajectories from the fitted model. - roughly (with the square) (adverb)
- Approximately, without being an exact figure.
Error grows roughly with the square of rollout length. - core (practical problem) (adjective)
- Central or most important part of an issue.
This is a core practical problem for model rollouts. - contrast (v) (verb)
- To compare two things by highlighting their differences.
The lecture contrasts backprop with the policy gradient. - motivate (verb)
- To give a reason that leads to a certain design choice.
This motivates the short-rollout designs used later. - point (v) (verb)
- To indicate or lead toward something as the reason.
This points toward using short branched rollouts.
Chapters
- 0:00 <Untitled Chapter 1>
- 0:09 Model-free optimization with a model
- 7:14 The curse of long model-based rollouts
- 13:31 Model-based RL with short rollouts
← Lecture 12, Part 1: Model-Based RL with Policies · Lecture 12, Part 3: Model-Based RL with Policies →
