Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 53 of 99 · 29:20

Lecture 12, Part 4: Model-Based RL with Policies

CS 285: Lecture 12, Part 4: Model-Based RL with Policies on YouTube

Study guide

What this lecture covers

This closing part of Lecture 12 steps away from one-step dynamics models and introduces successor representations, an advanced and more research-oriented idea about what a "model" could be. Instead of predicting the very next state, a successor representation predicts a discounted distribution over all future states, which is enough to evaluate a policy's value for any reward function without knowing the reward in advance.

The lecture builds this idea from first principles: it rewrites the value function as an inner product between a state-visitation vector and a reward vector, generalizes that vector into lower-dimensional "successor features" for tractability in large state spaces, and shows how to use them to recover Q-functions and improve policies. It closes by reframing the whole problem as classifier training so the idea extends to continuous state spaces, pointing toward the C-learning algorithm for further reading.

Key ideas

  • Successor representation (mu^pi): a vector giving, for each state, the discounted probability of visiting that state at some point in the future under policy pi; it depends on the policy but not on the reward.
  • Value as an inner product: the value function can be written as mu^pi(s_t) dotted with the reward vector, scaled by 1/(1-gamma), separating "where you'll end up" from "what you get there."
  • Bellman equation for mu: successor representations satisfy their own Bellman backup, using a pseudo-reward that is 1 for the current state and 0 elsewhere, so they can be learned with value-iteration-style updates.
  • Successor features (psi): projections of the successor representation onto a smaller feature basis phi, making the representation tractable when the state space is huge (for example, images).
  • Reward-as-linear-combination: if the reward can be written as phi(s) dotted with weights w, the value function is psi^pi(s) dotted with those same w, so one trained psi can support many rewards via a quick least-squares fit for w.
  • One step of policy iteration: choosing actions greedily from a recovered Q-function only improves the policy once; it is not generally optimal, though using many candidate policies and taking the best one per state does better.
  • Classifier trick for continuous states: framing "is this a future state of this state-action" as a binary classification problem avoids needing an ill-defined density for a specific future state, and the Bayes-optimal classifier recovers the successor density up to a constant that does not affect action selection.

Walkthrough

What kind of model is needed to evaluate a policy? (1:01)

The lecture reframes model-based RL: the model's real job is to let you evaluate a policy's expected return, not necessarily to predict the immediate next state. It rewrites the standard value function definition and rearranges the summation to expose a distribution over future states weighted by the discount factor.

Defining the successor representation (6:02)

The lecture formally defines the discounted future-state distribution P^pi(s_future = s | s_t), gives two equivalent interpretations (sampling a geometric time step, or a per-step termination probability), and shows the value function reduces to an inner product between this vector (mu^pi) and the reward vector. It also derives the Bellman backup for mu and notes the practical problems: it may not be easier to learn than model-free RL, and it does not scale to large or continuous state spaces.

Scaling with successor features (12:08)

To handle large state spaces, the lecture projects mu onto a feature basis phi to get successor features psi. It proves that if the reward is a linear combination of the same features, the value function is recovered from psi using the same weights, and extends the idea to Q-function versions of successor features.

Using successor features (19:14)

The lecture describes a two-step procedure: train psi^pi for a policy via Bellman backups, then fit reward weights w by least-squares regression on sampled (state, reward) pairs, and recover an approximate Q-function as psi dotted with w. It explains why this yields only one step of policy improvement, and how using many candidate policies and taking the per-state best can do better than any single one.

Extending to continuous states (24:17)

Because probability densities make a specific future state's probability zero in continuous spaces, the lecture reframes learning the successor distribution as training a binary classifier that distinguishes true future states from randomly sampled background states. It derives the Bayes-optimal classifier, shows how its ratio recovers the successor density up to a constant, and outlines how to train it on-policy with a cross-entropy loss, pointing to the C-learning paper for the off-policy extension.

Before you watch

  • Watch the earlier parts of Lecture 12 covering one-step dynamics models and short model-based rollouts, since this part explicitly moves away from that approach.
  • Be comfortable with value functions, Bellman backups, and Q-functions from the earlier value-based RL lectures.
  • Some familiarity with basic linear algebra (inner products, projections) will help with the successor-feature derivations.

Check your understanding

  1. Why does the successor representation depend on the policy but not on the reward, and how is that useful?
  2. How does the value function reduce to an inner product once you have the successor representation or successor features?
  3. Why does greedily acting on a Q-function recovered from successor features give only one step of policy improvement rather than an optimal policy?
  4. Why doesn't the classical successor representation extend directly to continuous state spaces, and how does the classifier-based reformulation get around this?
  5. What is the practical benefit of training successor features once and reusing them for multiple reward functions?

Vocabulary

successor representation (phrase)
A vector giving how likely a policy is to visit each state at some point in the future.
The successor representation depends on the policy, not the reward.
inner product (phrase)
A single number computed by multiplying matching entries of two vectors and adding the results.
The value function is an inner product of visitation and reward vectors.
discounted (adjective)
Reduced in weight the further away in time something is.
This is a discounted probability of future visitation.
geometric (adjective)
Related to a pattern where each step shrinks by a constant ratio.
One interpretation samples a geometric time step.
termination probability (phrase)
The chance that a process stops at each step.
Discounting can be modeled as a per-step termination probability.
Bellman equation (phrase)
An equation relating a value to the reward plus the value of the next state.
The successor representation satisfies its own Bellman equation.
pseudo-reward (phrase)
An artificial reward signal used to make an unrelated quantity learnable like a value function.
A pseudo-reward of 1 marks the current state.
successor features (phrase)
A compressed, lower-dimensional version of the successor representation using a feature basis.
Successor features scale to large state spaces.
feature basis (phrase)
A set of chosen functions used to summarize a state compactly.
We project the state onto a feature basis.
linear combination (phrase)
A sum of terms, each multiplied by its own weight.
The reward is a linear combination of features.
least-squares (phrase)
A method for fitting a model by minimizing the sum of squared errors.
We fit reward weights by least-squares regression.
policy improvement (phrase)
The step of updating a policy to perform better based on a value estimate.
Acting greedily gives one step of policy improvement.
candidate policy (phrase)
One of several possible policies being compared or tried.
Using many candidate policies can do better than one.
density (probability) (noun)
How concentrated a continuous probability distribution is at a point.
A specific future state has zero density in continuous spaces.
classifier (noun)
A model trained to sort inputs into categories.
We train a classifier to distinguish future states from others.
Bayes-optimal (adjective)
Making the mathematically best possible decision given the available probabilities.
The Bayes-optimal classifier recovers the successor density.
recover (a quantity) (verb)
To derive or obtain a value from related information.
The classifier's ratio helps recover the successor density.
advanced (adjective)
More complex, building on basic ideas already covered.
This is an advanced, research-oriented idea.
tractable (adjective)
Feasible to compute within reasonable time and resources.
Successor features make the idea tractable for large state spaces.
cross-entropy loss (phrase)
A loss function commonly used for training classifiers.
The classifier is trained with a cross-entropy loss.

Chapters

← Lecture 12, Part 3: Model-Based RL with Policies · Lecture 13, Part 1: Exploration →