Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 23 of 99 · 18:31

Lecture 6, Part 3: Implementing Actor-Critic

CS 285: Lecture 6, Part 3 on YouTube

Study guide

What this lecture covers

This part turns the actor-critic algorithm from earlier parts into something that works in practice. It first compares two neural network designs, separate actor and critic networks versus a shared trunk with two heads, then addresses the fact that fully online actor-critic updates on a single transition have too much variance for stochastic gradient descent to work well.

You come away understanding why batching matters, how synchronous and asynchronous parallel actor-critic get larger batches, and how switching to a replay buffer forces the algorithm to learn a Q function instead of a value function so it can handle transitions from older policies.

Key ideas

  • Separate networks: simplest to implement and more stable to train, at the cost of not sharing learned features between actor and critic.
  • Shared network with two heads: can be more sample-efficient when representations transfer, such as with image inputs, but harder to stabilize because the actor and critic gradients differ in scale.
  • Synchronous parallel actor-critic: multiple simulator workers each take one step, and the resulting transitions form a batch used for a synchronized update.
  • Asynchronous parallel actor-critic: workers update without waiting for each other, introducing slight bias from using parameters that are marginally older than the latest ones, which in practice tends to be an acceptable trade-off.
  • Off-policy actor-critic with a replay buffer: instead of running multiple threads, transitions from many past policies are stored and sampled to form a batch, but naive reuse of old actions breaks both the value target and the policy gradient.
  • Fix 1, learn Q instead of V: because old actions in the buffer are valid inputs to a Q function (unlike a V function that assumes on-policy actions), the critic is switched to Q(s,a), and target values are computed by querying the latest policy for the action it would take at the next state.
  • Fix 2, resample actions for the policy gradient: the policy gradient uses an action sampled from the current policy at the buffer's state, not the action originally stored, which keeps the gradient estimate unbiased with respect to pi_theta.
  • Remaining bias: the states themselves still come from older policies' visitation distributions, which is accepted as a source of bias that tends to be benign in practice.

Before you watch

  • Watch Parts 1 and 2 of this lecture, since they introduce the value function, advantage estimate, and discounting used throughout.
  • Recall the Q function definition from earlier in the course, since this part relies on it to fix the off-policy update.

Check your understanding

  1. What is the trade-off between separate actor/critic networks and a shared-trunk design?
  2. Why does a fully online, single-sample actor-critic update perform poorly with deep neural networks?
  3. What is the key difference between synchronous and asynchronous parallel actor-critic?
  4. Why does off-policy actor-critic need to learn a Q function instead of a value function?
  5. How does resampling the action at the buffer's state keep the policy gradient estimator unbiased?

Vocabulary

architecture (noun)
The overall structural design of a neural network.
This lecture compares two network architecture choices.
trunk (noun)
A shared part of a network before it splits into separate outputs.
A shared trunk feeds into two separate output heads.
head (noun)
A separate output branch of a neural network built on a shared base.
The actor and critic each have their own head.
stochastic gradient descent (noun)
An optimization method that updates parameters using small random batches of data.
A single transition has too much variance for stochastic gradient descent.
synchronous (adjective)
Happening at the same coordinated time across multiple parts.
Synchronous parallel actor-critic waits for all workers each step.
asynchronous (adjective)
Happening independently, without waiting for other parts to finish.
Asynchronous actor-critic lets workers update without waiting.
replay buffer (noun)
A stored collection of past experiences used for later training.
A replay buffer holds transitions from many past policies.
naive reuse (phrase)
Using old data again without adjusting for how it no longer matches the current model.
Naive reuse of old actions breaks the value target.
query (verb)
To ask a model or system for a value or prediction.
We query the latest policy for its next action.
resample (verb)
To draw a new sample again, replacing an older one.
The policy gradient resamples the action at each state.
visitation distribution (noun)
The pattern of which states a policy tends to visit over time.
Old states still come from an outdated visitation distribution.
benign (adjective)
Not harmful; causing little or no serious problem.
This remaining bias tends to be benign in practice.
stable (adjective)
Not prone to sudden failure or divergence during training.
Separate networks are more stable to train.
sample-efficient (adjective)
Able to learn well using relatively little collected data.
A shared network can be more sample-efficient.
worker (noun)
A separate process or machine running part of a computation.
Multiple simulator workers collect data in parallel.
marginally (adverb)
By a small amount.
Asynchronous updates use parameters that are marginally older.
trade-off (noun)
A balance between two good things where you cannot fully have both.
There's a trade-off between network designs.
target value (noun)
The value a model is trained to match during learning.
The target value is computed from the latest policy.
unbiased (adjective)
Not systematically wrong on average.
Resampling the action keeps the gradient estimate unbiased.
share features (phrase)
To reuse the same learned internal representations across different tasks.
A shared trunk lets the actor and critic share features.

Chapters

← Lecture 6, Part 2: Discount Factors in Actor-Critic · Lecture 6, Part 4: Eligibility Traces and GAE →