Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 29 of 99 · 17:01
Lecture 7, Part 4: Why Fitted Q-Iteration Doesn't Converge
Study guide
What this lecture covers
This closing part of the value-based methods lecture gives the theoretical explanation for a claim made earlier: fitted value iteration and fitted Q-iteration are not guaranteed to converge when using neural networks, even though tabular value iteration is. It builds this argument using the Bellman backup operator and a projection operator from supervised learning.
You come away understanding why each operator individually is a contraction, but why their composition is not, and why Q-learning's update does not correspond to gradient descent on any well-defined objective.
Key ideas
- Bellman backup operator (
B): applies the value iteration update to a whole value function vector at once; it is a contraction in the infinity norm, meaning it always moves any two value functions closer together by a factor ofgamma. - Fixed point: the optimal value function
V*is the unique fixed point ofB, so repeatedly applyingBto any starting vector converges toV*in the tabular case. - Projection operator (
Pi): represents the supervised learning step of fitted value iteration, projecting the backed-up value onto the space of representable neural network value functions; it is a contraction in the L2 norm. - Why fitted value iteration can fail to converge:
Bcontracts in the infinity norm andPicontracts in the L2 norm, but the compositionPi * Bis not guaranteed to be a contraction under any norm, so each fitted value iteration step can move away from the optimum even though both operators individually contract. - Same argument applies to fitted Q-iteration and Q-learning: the analogous Bellman operator with a max over actions is still a contraction in the infinity norm, but combined with a projection step, convergence is not guaranteed with function approximation.
- Q-learning is not gradient descent: because the target value itself depends on the current Q function but the gradient is not taken through that dependence, the update does not follow the gradient of a well-defined objective; taking the true gradient (residual gradient) is possible but tends to have poor numerical behavior in practice.
- Actor-critic has the same issue: fitted bootstrapped policy evaluation in actor-critic is subject to the same non-convergence argument, since it also combines a Bellman backup with a projection.
Before you watch
- Watch Parts 1 through 3 of this lecture on policy iteration, fitted Q-iteration, and Q-learning, since this part explains why those algorithms lack convergence guarantees.
Check your understanding
- Why is the Bellman backup operator
Ba contraction, and in which norm? - Why is the projection operator
Pia contraction in the L2 norm? - Why does composing two contractions,
BandPi, not guarantee thatPi * Bis a contraction? - Why is the Q-learning update not the gradient of a well-defined objective?
- Why does this non-convergence argument also apply to actor-critic's fitted value function updates?
Vocabulary
- contraction (noun)
- An operation that always brings two things closer together each time it is applied.
The Bellman backup is a contraction in the infinity norm. - operator (noun)
- A rule or function that transforms one mathematical object into another.
The Bellman backup operator updates a whole value function at once. - fixed point (phrase)
- A value that stays the same after an operation is applied to it.
The optimal value function is the fixed point of the backup operator. - infinity norm (phrase)
- A way of measuring the largest single difference between two vectors.
The Bellman operator contracts in the infinity norm. - L2 norm (phrase)
- A way of measuring overall distance between two vectors using squared differences.
The projection operator is a contraction in the L2 norm. - projection (noun)
- Mapping a value onto the closest point within a limited set of allowed values.
Projection fits the backed-up value onto representable functions. - composition (noun)
- Applying one operation after another, treating the pair as a single combined step.
The composition of the two operators is not guaranteed to be a contraction. - guarantee (noun)
- A firm promise that something will happen or hold true.
Fitted value iteration lacks a convergence guarantee. - vector (noun)
- A list of numbers treated as a single mathematical object.
The Bellman backup updates a value function vector at once. - supervised learning (phrase)
- Training a model to match given input-output example pairs.
The projection step is like a supervised learning fit. - represent (verb)
- To be able to express or capture a certain form or pattern.
A neural network can only represent a limited set of functions. - analogous (adjective)
- Similar in structure or role to something else already discussed.
The Q-learning operator is analogous to the value function operator. - max over actions (phrase)
- Choosing the highest value among all possible actions.
The Q-learning operator uses a max over actions. - gradient descent (noun)
- An optimization method that repeatedly moves parameters to reduce error.
Q-learning's update does not follow gradient descent on a fixed objective. - objective (noun)
- The quantity a training process tries to minimize or maximize.
There is no well-defined objective behind the Q-learning update. - residual gradient (phrase)
- A method that takes the true gradient of the Bellman error, including through the target.
Residual gradient methods have poor numerical behavior in practice. - numerical behavior (phrase)
- How stable or well-behaved calculations are during training.
Residual gradient tends to have poor numerical behavior. - bootstrapped (adjective)
- Based partly on the model's own current predictions rather than only real data.
Fitted bootstrapped policy evaluation has the same issue. - theoretical (adjective)
- Based on formal reasoning rather than experiment.
This part gives the theoretical explanation for non-convergence. - argument (proof) (noun)
- A logical explanation used to support a conclusion.
The same argument applies to actor-critic's value updates.
Chapters
- 0:00 Intro
- 0:15 Value iteration
- 6:42 Fitted value iteration
- 12:31 Fitted q iteration
- 13:32 Gradient descent
- 15:12 Not convergent
- 16:01 Review
← Lecture 7, Part 3: Q-Learning and Exploration · Lecture 8, Part 1: Replay Buffers and the Correlation Problem →
