Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 31 of 99 · 11:56
Lecture 8, Part 2: Target Networks for Q-Learning
Study guide
What this lecture covers
This part of Lecture 8 addresses a second core instability in Q-learning, after the replay buffer already fixed the problem of correlated samples. The lecture asks: why does Q-learning behave differently from ordinary regression, and how can training be stabilized when the regression targets themselves keep changing as the network updates?
By the end you can describe how a replay buffer combined with a periodically-updated target network produces the classic "deep Q-learning" (DQN) algorithm, and you understand the alternative of Polyak-averaged target updates. This builds directly on the replay buffer discussion earlier in Lecture 8 and sets up the DQN implementation used in homework 3.
Key ideas
- Moving target problem: in Q-learning the regression target depends on the same parameters being updated, so unlike supervised regression the target shifts on every gradient step, which destabilizes learning.
- Target network: a separate parameter vector
phi'is used to compute target values while the online parametersphiare updated, so targets stay fixed for a stretch of training. - General recipe: collect data with the current policy, sample a batch from the replay buffer, take
kgradient steps onphi, and everynsteps copyphiintophi'. - Classic DQN: the special case where
k = 1, matching the update rate of data collection and gradient steps whilephi'refreshes only everynsteps (often around 10,000). - Lag inconsistency: because
phi'is only replaced periodically, the target network's staleness varies from one step old tonsteps old depending on where you are in the cycle. - Polyak averaging alternative: instead of a hard copy, update
phi'every step as a weighted averagetau * phi' + (1 - tau) * phiwithtauclose to 1, giving a uniformly lagged target at every step.
Walkthrough
What's wrong with plain Q-learning (0:13)
The lecture recaps that a replay buffer solves the correlated-samples problem, but Q-learning still is not true gradient descent because its regression target moves as the parameters change. Running the regression step to full convergence on a moving target is not necessarily useful, which is why practice uses only a small number of gradient steps.
Q-learning as regression, and target networks (1:05)
The lecture connects fitted Q-iteration's regression step to this instability and introduces a hybrid: collect data continuously, but compute targets using a separate parameter vector phi' that lags behind the online parameters phi. Multiple gradient updates can be taken against the same fixed target before phi' is refreshed.
Q-learning with target networks (2:12)
The general procedure is laid out step by step: collect data, sample a batch, update phi for k steps using targets from phi', and every n steps set phi' = phi. Typical values are k between 1 and 4 and n around 10,000, chosen so target values stay stable long enough for the regression to behave like ordinary supervised learning.
The classic DQN algorithm (5:20)
Setting k = 1 in the general recipe recovers the standard deep Q-network algorithm: take a step and store it in the buffer, sample a mini-batch, compute targets with phi', update phi by regression, and every n steps synchronize phi' with phi. The lecture notes this is the version students implement for homework.
An alternative target update (8:02)
The lecture points out the oddity that periodic hard copies make the target's staleness inconsistent across steps, then introduces Polyak averaging as a smoother alternative: interpolate phi' toward phi a little every step using a large tau, which keeps the lag roughly constant and has some theoretical grounding as long as phi' stays close to phi.
Before you watch
- Know how the replay buffer removes sample correlation in Q-learning, covered earlier in Lecture 8.
- Be comfortable with fitted Q-iteration and its regression step, since target networks are framed as a middle ground between online Q-learning and full fitted Q-iteration.
Check your understanding
- Why is Q-learning not equivalent to standard gradient descent on a fixed objective?
- What roles do
phiandphi'play in the target-network version of Q-learning? - How do the choices of
kandntrade off data collection speed against target stability? - How does Polyak averaging differ from the periodic hard-copy update of the target network?
Vocabulary
- instability (noun)
- A tendency for training to become unreliable or diverge.
This part addresses a second core instability in Q-learning. - moving target (phrase)
- A goal that keeps changing while you try to reach it.
The regression target is a moving target in Q-learning. - regression (noun)
- Fitting a model to predict a continuous number from data.
Q-learning can be seen as a kind of regression. - destabilize (verb)
- To make something less stable or reliable.
A shifting target destabilizes training. - target network (phrase)
- A separate, slower-updating copy of the network used to compute stable targets.
The target network keeps targets fixed for a while. - online parameters (phrase)
- The network weights being actively updated during training right now.
The online parameters are updated every gradient step. - synchronize (verb)
- To make two things match or align at the same point.
Every n steps we synchronize the target network with the online one. - lag (noun)
- A delay between one thing and another.
The target network has a lag behind the online network. - staleness (noun)
- How out of date something is.
The target's staleness varies depending on the training step. - Polyak averaging (phrase)
- Smoothly blending old and new parameter values a little at a time.
Polyak averaging gives a uniformly lagged target network. - hard copy (phrase)
- Directly replacing one set of values with another all at once.
DQN uses a periodic hard copy to update the target network. - interpolate (verb)
- To calculate a value partway between two known values.
Polyak averaging interpolates the target toward the online parameters. - grounding (noun)
- A basis of support or justification for an idea.
Polyak averaging has some theoretical grounding. - mini-batch (noun)
- A small group of samples used together for one training update.
We sample a mini-batch from the replay buffer. - hybrid (noun)
- A combination of two different approaches.
Target networks are a hybrid between full and online updates. - stretch (of training) (noun)
- A continuous period of time during training.
Targets stay fixed for a stretch of training. - convergence (noun)
- The point where a process settles into a stable, final answer.
Target networks help convergence by stabilizing the targets. - oddity (noun)
- Something strange or unusual worth noticing.
The lecture points out an oddity in the hard-copy schedule. - classic (adjective)
- Well-known and widely used as a standard example.
This is the classic DQN algorithm used in the homework. - alternative (noun)
- A different option that can be used instead of the usual one.
Polyak averaging is an alternative to hard copies.
Chapters
- 0:00 <Untitled Chapter 1>
- 0:13 What's wrong?
- 1:05 Q-Learning and Regression
- 2:12 Q-Learning with target networks
- 5:20 "Classic" deep Q-learning algorithm (DQN)
- 8:02 Alternative target network
← Lecture 8, Part 1: Replay Buffers and the Correlation Problem · Lecture 8, Part 3: A Unified View of Q-Learning →
