Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 28 of 99 · 11:57
Lecture 7, Part 3: Q-Learning and Exploration
Study guide
What this lecture covers
This part explains why fitted Q-iteration counts as an off-policy algorithm: the policy only enters through the max inside the Q-function target, which acts like a simulated lookup rather than requiring a new rollout, so old transitions stay valid for training. It also clarifies what fitted Q-iteration actually optimizes, the Bellman error, and what that error implies about policy quality.
The lecture then derives online Q-learning as a special case of fitted Q-iteration with a single transition per update, and addresses a practical problem: using a purely greedy policy to collect data can get the agent stuck taking a bad action forever. This motivates exploration strategies, in particular epsilon-greedy and Boltzmann (softmax) exploration.
Key ideas
- Why fitted Q-iteration is off-policy: the current policy only appears as an argument to the Q function inside a max, so changing the policy does not require generating new rollouts, since the transition
(s,a,s')does not depend on it. - Bellman error: fitted Q-iteration minimizes the difference between
Q_phi(s,a)and target valuesy; zero error implies the optimal Q function and policy, but nonzero error gives no guarantees about policy quality. - Online Q-learning: a special case of fitted Q-iteration using one transition at a time: take an action, compute one target value, and take one gradient step on the temporal difference error
Q_phi(s,a) - y. - Temporal difference error: the difference between a Q function's current prediction and its bootstrapped target value.
- Problem with a purely greedy policy during learning: because the argmax policy is deterministic, a poor initial Q function can cause the agent to repeatedly take the same suboptimal action and never discover better ones.
- Epsilon-greedy exploration: take the greedy action with probability
1 - epsilonand a uniformly random other action with probabilityepsilon, often decreasing epsilon over training as the Q function improves. - Boltzmann (softmax) exploration: choose actions in proportion to a positive transformation (typically exponentiation) of their Q values, so similarly good actions get similar probability and very bad actions are rarely chosen.
Before you watch
- Watch Part 2 of this lecture on fitted Q-iteration, since this part builds directly on its algorithm and target value computation.
Check your understanding
- Why doesn't fitted Q-iteration need new rollouts when the policy changes?
- What does zero Bellman error imply about the resulting policy, and what can be said when the error is not zero?
- How is online Q-learning derived as a special case of fitted Q-iteration?
- Why can using a purely greedy policy during learning prevent the agent from discovering better actions?
- How does Boltzmann exploration differ from epsilon-greedy in how it treats two similarly good actions?
Vocabulary
- off-policy (adjective)
- Able to learn from data collected by a policy different from the one currently being improved.
Fitted Q-iteration is off-policy because samples stay valid regardless of the policy. - rollout (noun)
- One full run of taking actions in the environment and observing the results.
Changing the policy does not require a new rollout. - Bellman error (phrase)
- The gap between a Q function's prediction and its target value.
Fitted Q-iteration minimizes the Bellman error. - optimal (adjective)
- The best possible option available.
Zero Bellman error implies the optimal Q function. - guarantee (noun)
- A firm promise that something will definitely happen.
Nonzero error gives no guarantees about policy quality. - temporal difference error (phrase)
- The difference between a value's current prediction and its bootstrapped target.
Online Q-learning takes a gradient step on the temporal difference error. - greedy policy (phrase)
- A policy that always chooses the action with the highest estimated value.
A purely greedy policy can get stuck on a bad action. - suboptimal (adjective)
- Worse than the best possible choice.
The agent keeps repeating a suboptimal action. - exploration (noun)
- Trying new or uncertain actions to discover potentially better outcomes.
Exploration helps the agent discover better actions. - epsilon-greedy (noun)
- A strategy that usually picks the best-known action but sometimes picks a random one.
Epsilon-greedy takes a random action with small probability epsilon. - anneal (verb)
- To gradually reduce a value over time.
We anneal epsilon as the Q function improves. - Boltzmann exploration (phrase)
- A strategy that picks actions with probability proportional to how good they seem.
Boltzmann exploration gives similar actions similar probabilities. - softmax (noun)
- A function that turns a set of numbers into probabilities that add up to one.
Softmax exponentiates Q values to get action probabilities. - exponentiation (noun)
- Raising a number to a power, such as e to the power of x.
Boltzmann exploration uses exponentiation of Q values. - proportion (noun)
- A share or fraction of a whole.
Actions are chosen in proportion to their Q values. - uniformly random (phrase)
- Chosen with equal chance among all options.
With probability epsilon we take a uniformly random action. - argument (noun)
- An input value given to a function.
The policy only appears as an argument to the Q function. - special case (phrase)
- A particular, simpler version of a more general idea.
Online Q-learning is a special case of fitted Q-iteration. - practical problem (phrase)
- A real difficulty that comes up when actually using a method.
Getting stuck on one action is a practical problem for Q-learning. - clarify (verb)
- To explain something more clearly.
The lecture clarifies what fitted Q-iteration optimizes.
Chapters
- 0:00 Intro
- 0:46 Q function
- 3:06 Optimization
- 4:53 Generic Fit a Q
- 5:19 Online Q Learning
- 6:45 greedy policy
- 8:14 exploration rules
- 11:16 review
← Lecture 7, Part 2: Fitted Value and Fitted Q-Iteration · Lecture 7, Part 4: Why Fitted Q-Iteration Doesn't Converge →
