Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 51 of 99 · 15:04

Lecture 12, Part 2: Model-Based RL with Policies

CS 285: Lecture 12, Part 2: Model-Based RL with Policies on YouTube

Study guide

What this lecture covers

This part continues the model-based reinforcement learning discussion by asking why you would use a model-free algorithm, like the policy gradient, on top of a learned model instead of backpropagating through it directly. It builds on the previous part's derivation of the backpropagation (pathwise) gradient and contrasts it with the likelihood-ratio (policy gradient) estimator.

The lecture then turns to a core practical problem: rollouts through a learned dynamics model accumulate error the longer they run, for the same distributional-shift reasons seen earlier in the course with imitation learning. This sets up the motivation for the short-rollout designs used by modern model-based RL algorithms, covered in the next part.

Key ideas

  • Likelihood ratio (policy gradient) estimator: does not require differentiating through the transition dynamics, only through sampling; it is valid for stochastic policies and transitions.
  • Pathwise (backpropagation) gradient: computes exact derivatives via the chain rule but requires multiplying a long product of Jacobians across time steps, which can explode or vanish.
  • Sampling trade-off: the policy gradient avoids the Jacobian product but needs many samples; with a model, samples are cheap compute rather than costly real-world interaction.
  • Model-based RL version 2.5: fit a model, generate many trajectories from it, and improve the policy with the policy gradient instead of backpropagation.
  • Distributional shift in rollouts: as in imitation learning, small model errors push the policy into unfamiliar states where errors compound, growing roughly with the square of the rollout length.
  • Branched short rollouts: sampling states from real trajectories and rolling out only briefly from each reduces compounding error while still covering later time steps.
  • On-policy vs. off-policy: because branched rollouts mix data from the collecting policy and the latest policy, off-policy algorithms like Q-learning handle this state-distribution mismatch better than on-policy policy gradient methods.

Before you watch

  • Watch the first part of Lecture 12, which introduces the backpropagation gradient for model-based RL.
  • Review the distributional shift argument from the imitation learning lectures, since it is reused here.

Check your understanding

  1. Why does the pathwise gradient become numerically unstable for long horizons, while the policy gradient does not?
  2. What trade-off do you accept when switching from backpropagation to the policy gradient for training with a learned model?
  3. Why does branching short rollouts from real-world states reduce accumulated error compared to full-length model rollouts?
  4. Why are off-policy methods generally preferred over on-policy policy gradient methods for training on these branched rollouts?

Vocabulary

likelihood ratio (phrase)
A ratio comparing how likely an outcome is under two different distributions.
The policy gradient is a likelihood ratio estimator.
pathwise gradient (phrase)
A gradient computed by directly differentiating through the whole computation, sample by sample.
The pathwise gradient uses the chain rule through the dynamics.
stochastic (adjective)
Involving randomness rather than a fixed outcome.
The estimator is valid for stochastic policies and dynamics.
differentiate through (phrase)
To compute a gradient that passes through a certain part of a computation.
The pathwise gradient differentiates through the dynamics.
trade-off (noun)
A balance where gaining one benefit costs you another.
There's a trade-off between the two gradient estimators.
compute (resource) (noun)
Processing power or machine resources used to run calculations.
With a model, extra samples cost cheap compute.
distributional shift (phrase)
A mismatch between the data used for training and the data actually encountered.
Rollouts through the model suffer from distributional shift.
accumulate error (phrase)
To build up mistakes over time, becoming larger with each step.
Small model errors accumulate error over a long rollout.
branch (v) (verb)
To start a new short path from a specific point.
We branch short rollouts from real states.
cover (later steps) (verb)
To include or account for something within a range.
Branched rollouts still cover later time steps.
on-policy (adjective)
Only valid when data comes from the exact same policy being trained.
On-policy methods struggle with mixed rollout data.
off-policy (adjective)
Able to learn from data collected under a different policy.
Off-policy methods like Q-learning handle mixed data better.
mismatch (noun)
A difference between two things that should ideally match.
There's a state-distribution mismatch in branched rollouts.
numerically unstable (phrase)
Prone to producing wildly incorrect results due to calculation errors.
The pathwise gradient becomes numerically unstable for long horizons.
generate (trajectories) (verb)
To produce new sequences of states and actions from a model.
We generate many trajectories from the fitted model.
roughly (with the square) (adverb)
Approximately, without being an exact figure.
Error grows roughly with the square of rollout length.
core (practical problem) (adjective)
Central or most important part of an issue.
This is a core practical problem for model rollouts.
contrast (v) (verb)
To compare two things by highlighting their differences.
The lecture contrasts backprop with the policy gradient.
motivate (verb)
To give a reason that leads to a certain design choice.
This motivates the short-rollout designs used later.
point (v) (verb)
To indicate or lead toward something as the reason.
This points toward using short branched rollouts.

Chapters

← Lecture 12, Part 1: Model-Based RL with Policies · Lecture 12, Part 3: Model-Based RL with Policies →