Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 27 of 99 · 15:08
Lecture 7, Part 2: Fitted Value and Fitted Q-Iteration
Study guide
What this lecture covers
This part moves value-based methods beyond the tabular setting into function approximation. It first explains why storing a value function as a table is infeasible for large or continuous state spaces (the curse of dimensionality), then derives fitted value iteration, which fits a neural network to target values computed with the max-over-actions trick.
The key move is switching from a value function to a Q function: because a Q function recurrence conditions on both state and action, the sample (s, a, s') needed to fit it does not depend on the current policy, which removes the need for known transition dynamics or the ability to try multiple actions from the same state. This leads to fitted Q-iteration, the basis of most model-free value-based RL algorithms.
Key ideas
- Curse of dimensionality: a tabular value function needs one entry per state, which becomes infeasible or infinite for high-dimensional or continuous state spaces such as images.
- Fitted value iteration: fits a neural network
V_phi(s)to target values computed asmax_a [r(s,a) + gamma * E[V(s')]], but still requires known transition dynamics to compute the expectation and to try multiple actions from the same state. - Switching to a Q function recurrence: because
Q(s,a)conditions the next state on a fixed(s,a), samples(s, a, s')remain valid regardless of what the current policy is, removing the dependence on known dynamics. - Fitted Q-iteration: alternates between computing target values
y_i = r(s_i,a_i) + gamma * max_a' Q_phi(s_i', a')using sampled transitions, and regressingQ_phionto those targets. - Off-policy by construction: fitted Q-iteration works with samples gathered by any policy, since it never requires actions from the current greedy policy, unlike actor-critic.
- No convergence guarantees with neural networks: fitted Q-iteration is guaranteed to converge for tabular representations but not in general for nonlinear function approximators.
- Algorithm structure: collect transitions with some policy, compute target values from the previous Q function, then run some number of gradient steps to fit a new Q function, optionally repeating the fit step multiple times before collecting more data.
Before you watch
- Watch Part 1 of this lecture on policy and value iteration, since fitted Q-iteration is presented as its function-approximation, model-free extension.
- Recall the Q function and Bellman backup definitions from the actor-critic lecture.
Check your understanding
- Why is a tabular value function infeasible for a high-dimensional or continuous state space?
- Why does fitted value iteration still require knowing the transition dynamics?
- What change makes the Q function recurrence usable with samples from any policy?
- What are the two steps of fitted Q-iteration, and what does each compute?
- Why does fitted Q-iteration have convergence guarantees in the tabular case but not with neural network function approximation?
Vocabulary
- function approximation (phrase)
- Using a model, such as a neural network, to estimate values instead of storing every one exactly.
Function approximation is needed once states no longer fit in a table. - curse of dimensionality (phrase)
- The problem that the number of possible states explodes as the state description grows larger.
The curse of dimensionality makes a table of image states impossible. - infeasible (adjective)
- Not realistically possible to do.
Storing a table for every possible image is infeasible. - neural network (noun)
- A model made of connected layers that learns patterns from data.
Fitted value iteration fits a neural network to target values. - target value (phrase)
- The number a model is trained to predict, computed from rewards and estimates.
The target value combines the reward and the next state's value. - expectation (noun)
- The average outcome, weighted by how likely each outcome is.
Computing the expectation requires knowing the transition probabilities. - recurrence (noun)
- A rule that defines a value in terms of earlier or related values.
The Q function recurrence relates Q(s,a) to the next state's value. - regress (verb)
- To fit a model so its output matches given target numbers as closely as possible.
We regress the Q function onto the sampled target values. - off-policy (adjective)
- Able to learn from data collected by a different policy than the one being improved.
Fitted Q-iteration is off-policy by construction. - by construction (phrase)
- True automatically because of how something is built or defined.
The method is off-policy by construction, not by accident. - gradient step (phrase)
- One update to a model's parameters that moves them slightly to reduce error.
We take several gradient steps to fit the new Q function. - convergence guarantee (phrase)
- A mathematical promise that a method will eventually reach the correct answer.
There is no convergence guarantee with neural networks. - nonlinear (adjective)
- Not following a straight-line relationship between input and output.
Neural networks are nonlinear function approximators. - algorithm structure (phrase)
- The overall sequence of steps that make up a method.
The algorithm structure alternates data collection and fitting. - high-dimensional (adjective)
- Described by many numbers or variables at once.
Images are a high-dimensional kind of state. - continuous (adjective)
- Able to take any value within a range, not just a fixed set.
A continuous state space has infinitely many possible states. - sampled transition (phrase)
- One recorded example of moving from a state to a next state after an action.
Target values are computed using a sampled transition. - iterate (verb)
- To repeat a process again and again, each time using the previous result.
We iterate between computing targets and fitting the network. - extension (noun)
- A version of a method that adds new capability to an earlier one.
Fitted Q-iteration is a function-approximation extension of value iteration. - gather (verb)
- To collect data from the environment.
We gather transitions with some exploration policy.
Chapters
- 0:00 Intro
- 0:21 Fitted value iteration
- 3:49 What if we don't know the transition dynamics?
- 7:08 Can we do the "max" trick again?
- 11:05 Fitted Q-iteration
- 14:35 Review
← Lecture 7, Part 1: From Actor-Critic to Policy Iteration · Lecture 7, Part 3: Q-Learning and Exploration →
