Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 29 of 99 · 17:01

Lecture 7, Part 4: Why Fitted Q-Iteration Doesn't Converge

CS 285: Lecture 7, Part 4 on YouTube

Study guide

What this lecture covers

This closing part of the value-based methods lecture gives the theoretical explanation for a claim made earlier: fitted value iteration and fitted Q-iteration are not guaranteed to converge when using neural networks, even though tabular value iteration is. It builds this argument using the Bellman backup operator and a projection operator from supervised learning.

You come away understanding why each operator individually is a contraction, but why their composition is not, and why Q-learning's update does not correspond to gradient descent on any well-defined objective.

Key ideas

  • Bellman backup operator (B): applies the value iteration update to a whole value function vector at once; it is a contraction in the infinity norm, meaning it always moves any two value functions closer together by a factor of gamma.
  • Fixed point: the optimal value function V* is the unique fixed point of B, so repeatedly applying B to any starting vector converges to V* in the tabular case.
  • Projection operator (Pi): represents the supervised learning step of fitted value iteration, projecting the backed-up value onto the space of representable neural network value functions; it is a contraction in the L2 norm.
  • Why fitted value iteration can fail to converge: B contracts in the infinity norm and Pi contracts in the L2 norm, but the composition Pi * B is not guaranteed to be a contraction under any norm, so each fitted value iteration step can move away from the optimum even though both operators individually contract.
  • Same argument applies to fitted Q-iteration and Q-learning: the analogous Bellman operator with a max over actions is still a contraction in the infinity norm, but combined with a projection step, convergence is not guaranteed with function approximation.
  • Q-learning is not gradient descent: because the target value itself depends on the current Q function but the gradient is not taken through that dependence, the update does not follow the gradient of a well-defined objective; taking the true gradient (residual gradient) is possible but tends to have poor numerical behavior in practice.
  • Actor-critic has the same issue: fitted bootstrapped policy evaluation in actor-critic is subject to the same non-convergence argument, since it also combines a Bellman backup with a projection.

Before you watch

  • Watch Parts 1 through 3 of this lecture on policy iteration, fitted Q-iteration, and Q-learning, since this part explains why those algorithms lack convergence guarantees.

Check your understanding

  1. Why is the Bellman backup operator B a contraction, and in which norm?
  2. Why is the projection operator Pi a contraction in the L2 norm?
  3. Why does composing two contractions, B and Pi, not guarantee that Pi * B is a contraction?
  4. Why is the Q-learning update not the gradient of a well-defined objective?
  5. Why does this non-convergence argument also apply to actor-critic's fitted value function updates?

Vocabulary

contraction (noun)
An operation that always brings two things closer together each time it is applied.
The Bellman backup is a contraction in the infinity norm.
operator (noun)
A rule or function that transforms one mathematical object into another.
The Bellman backup operator updates a whole value function at once.
fixed point (phrase)
A value that stays the same after an operation is applied to it.
The optimal value function is the fixed point of the backup operator.
infinity norm (phrase)
A way of measuring the largest single difference between two vectors.
The Bellman operator contracts in the infinity norm.
L2 norm (phrase)
A way of measuring overall distance between two vectors using squared differences.
The projection operator is a contraction in the L2 norm.
projection (noun)
Mapping a value onto the closest point within a limited set of allowed values.
Projection fits the backed-up value onto representable functions.
composition (noun)
Applying one operation after another, treating the pair as a single combined step.
The composition of the two operators is not guaranteed to be a contraction.
guarantee (noun)
A firm promise that something will happen or hold true.
Fitted value iteration lacks a convergence guarantee.
vector (noun)
A list of numbers treated as a single mathematical object.
The Bellman backup updates a value function vector at once.
supervised learning (phrase)
Training a model to match given input-output example pairs.
The projection step is like a supervised learning fit.
represent (verb)
To be able to express or capture a certain form or pattern.
A neural network can only represent a limited set of functions.
analogous (adjective)
Similar in structure or role to something else already discussed.
The Q-learning operator is analogous to the value function operator.
max over actions (phrase)
Choosing the highest value among all possible actions.
The Q-learning operator uses a max over actions.
gradient descent (noun)
An optimization method that repeatedly moves parameters to reduce error.
Q-learning's update does not follow gradient descent on a fixed objective.
objective (noun)
The quantity a training process tries to minimize or maximize.
There is no well-defined objective behind the Q-learning update.
residual gradient (phrase)
A method that takes the true gradient of the Bellman error, including through the target.
Residual gradient methods have poor numerical behavior in practice.
numerical behavior (phrase)
How stable or well-behaved calculations are during training.
Residual gradient tends to have poor numerical behavior.
bootstrapped (adjective)
Based partly on the model's own current predictions rather than only real data.
Fitted bootstrapped policy evaluation has the same issue.
theoretical (adjective)
Based on formal reasoning rather than experiment.
This part gives the theoretical explanation for non-convergence.
argument (proof) (noun)
A logical explanation used to support a conclusion.
The same argument applies to actor-critic's value updates.

Chapters

← Lecture 7, Part 3: Q-Learning and Exploration · Lecture 8, Part 1: Replay Buffers and the Correlation Problem →