Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 30 of 99 · 13:37

Lecture 8, Part 1: Replay Buffers and the Correlation Problem

CS 285: Lecture 8, Part 1 on YouTube

Study guide

What this lecture covers

Opening Lecture 8, this part recaps the general fitted Q-iteration recipe and its online special case, Q-learning, then digs into a practical problem beyond the non-convergence issue raised previously: taking gradient steps on sequential, highly correlated transitions makes the function approximator locally overfit to whatever part of the trajectory it just saw, then "forget" as the trajectory moves on.

You come away understanding two solutions: parallel data collection across multiple workers (as used in actor-critic), and the more common approach of a replay buffer, which stores past transitions and samples i.i.d. batches from them to decorrelate updates and provide broader data coverage.

Key ideas

  • Online Q-learning recap: take one action, compute one target value y = r + gamma * max_a' Q(s', a'), and take one gradient step toward it, repeating this every environment step.
  • Correlation problem: consecutive states in a trajectory are highly correlated, so training on them one at a time causes the function approximator to locally overfit to recent transitions and lose accuracy on earlier parts of the trajectory.
  • Parallel workers: running multiple simulators and collecting a batch of transitions from all of them at once reduces correlation, similar to the parallel actor-critic approach from an earlier lecture.
  • Replay buffer: a stored collection of past transitions from potentially many different policies; sampling a batch i.i.d. from the buffer decorrelates training samples and lowers gradient variance.
  • Buffer needs refreshing: because an early, poor policy may not visit useful parts of the state space, the buffer must periodically be replenished by deploying the latest policy (often with epsilon-greedy exploration) to collect new transitions.
  • Full deep Q-learning recipe: collect transitions into a buffer, sample a batch, compute target values and take a gradient step summed over the batch, and repeat some number of times before collecting more data.

Before you watch

  • Watch the earlier parts of this lecture sequence on fitted Q-iteration and Q-learning, since this part builds directly on that algorithm and target value computation.
  • Recall the discussion of synchronous and asynchronous parallel actor-critic, since a similar parallelism idea is reused here.

Check your understanding

  1. Why does training on sequential transitions one at a time cause the function approximator to overfit locally?
  2. How does using multiple parallel workers help address the correlation problem?
  3. Why does sampling i.i.d. batches from a replay buffer decorrelate training samples?
  4. Why is it still necessary to periodically collect new data with the latest policy rather than relying only on the initial buffer contents?

Vocabulary

recap (noun)
A short summary of ideas covered earlier.
The lecture opens with a recap of fitted Q-iteration.
sequential (adjective)
Happening one after another in a fixed order.
Sequential transitions from one trajectory are highly correlated.
correlated (adjective)
Related to each other, so knowing one tells you something about the other.
Nearby states in a trajectory are strongly correlated.
locally overfit (phrase)
Adjusted too closely to a small, recent set of examples.
The network can locally overfit to whatever part of the trajectory it just saw.
forget (verb)
To lose previously learned information.
The model can forget earlier parts of the trajectory as training moves on.
parallel workers (phrase)
Several separate processes running the same task at the same time.
Parallel workers collect less correlated data.
decorrelate (verb)
To make data less related to each other.
Sampling randomly helps decorrelate training data.
replay buffer (noun)
A stored collection of past experiences used to sample training data later.
The replay buffer holds transitions from many past policies.
i.i.d. (adjective)
Independent and identically distributed; each sample drawn separately and from the same distribution.
Sampling i.i.d. batches lowers gradient variance.
gradient variance (phrase)
How much the direction of a training update changes from one batch to another.
A larger buffer lowers gradient variance.
refresh (verb)
To update something with new, current information.
The buffer needs to be refreshed with new transitions.
replenish (verb)
To fill something back up after it has been used or emptied.
We replenish the buffer with the latest policy's data.
deploy (verb)
To put a policy into use in the real environment.
We deploy the latest policy to collect new transitions.
target value (phrase)
The number a model tries to match, computed from the reward and next-state estimate.
We compute the target value using the reward and max Q.
coverage (noun)
How much of the possible situations or states a dataset includes.
A large replay buffer gives broader state coverage.
trajectory (noun)
A full sequence of states and actions from start to end of an episode.
Consecutive states in a trajectory are highly correlated.
environment step (phrase)
One single action taken and its resulting outcome in the environment.
Online Q-learning updates after every environment step.
practical problem (phrase)
A real difficulty encountered when actually running a method.
Correlation is a practical problem beyond non-convergence.
recipe (noun)
A set sequence of steps used to solve a problem.
The full deep Q-learning recipe adds a replay buffer step.
simulator (noun)
A computer program that mimics a real environment for training.
Multiple simulators can run in parallel to gather data.

Chapters

← Lecture 7, Part 4: Why Fitted Q-Iteration Doesn't Converge · Lecture 8, Part 2: Target Networks for Q-Learning →