Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 23 of 99 · 18:31
Lecture 6, Part 3: Implementing Actor-Critic
Study guide
What this lecture covers
This part turns the actor-critic algorithm from earlier parts into something that works in practice. It first compares two neural network designs, separate actor and critic networks versus a shared trunk with two heads, then addresses the fact that fully online actor-critic updates on a single transition have too much variance for stochastic gradient descent to work well.
You come away understanding why batching matters, how synchronous and asynchronous parallel actor-critic get larger batches, and how switching to a replay buffer forces the algorithm to learn a Q function instead of a value function so it can handle transitions from older policies.
Key ideas
- Separate networks: simplest to implement and more stable to train, at the cost of not sharing learned features between actor and critic.
- Shared network with two heads: can be more sample-efficient when representations transfer, such as with image inputs, but harder to stabilize because the actor and critic gradients differ in scale.
- Synchronous parallel actor-critic: multiple simulator workers each take one step, and the resulting transitions form a batch used for a synchronized update.
- Asynchronous parallel actor-critic: workers update without waiting for each other, introducing slight bias from using parameters that are marginally older than the latest ones, which in practice tends to be an acceptable trade-off.
- Off-policy actor-critic with a replay buffer: instead of running multiple threads, transitions from many past policies are stored and sampled to form a batch, but naive reuse of old actions breaks both the value target and the policy gradient.
- Fix 1, learn Q instead of V: because old actions in the buffer are valid inputs to a Q function (unlike a V function that assumes on-policy actions), the critic is switched to
Q(s,a), and target values are computed by querying the latest policy for the action it would take at the next state. - Fix 2, resample actions for the policy gradient: the policy gradient uses an action sampled from the current policy at the buffer's state, not the action originally stored, which keeps the gradient estimate unbiased with respect to
pi_theta. - Remaining bias: the states themselves still come from older policies' visitation distributions, which is accepted as a source of bias that tends to be benign in practice.
Before you watch
- Watch Parts 1 and 2 of this lecture, since they introduce the value function, advantage estimate, and discounting used throughout.
- Recall the Q function definition from earlier in the course, since this part relies on it to fix the off-policy update.
Check your understanding
- What is the trade-off between separate actor/critic networks and a shared-trunk design?
- Why does a fully online, single-sample actor-critic update perform poorly with deep neural networks?
- What is the key difference between synchronous and asynchronous parallel actor-critic?
- Why does off-policy actor-critic need to learn a Q function instead of a value function?
- How does resampling the action at the buffer's state keep the policy gradient estimator unbiased?
Vocabulary
- architecture (noun)
- The overall structural design of a neural network.
This lecture compares two network architecture choices. - trunk (noun)
- A shared part of a network before it splits into separate outputs.
A shared trunk feeds into two separate output heads. - head (noun)
- A separate output branch of a neural network built on a shared base.
The actor and critic each have their own head. - stochastic gradient descent (noun)
- An optimization method that updates parameters using small random batches of data.
A single transition has too much variance for stochastic gradient descent. - synchronous (adjective)
- Happening at the same coordinated time across multiple parts.
Synchronous parallel actor-critic waits for all workers each step. - asynchronous (adjective)
- Happening independently, without waiting for other parts to finish.
Asynchronous actor-critic lets workers update without waiting. - replay buffer (noun)
- A stored collection of past experiences used for later training.
A replay buffer holds transitions from many past policies. - naive reuse (phrase)
- Using old data again without adjusting for how it no longer matches the current model.
Naive reuse of old actions breaks the value target. - query (verb)
- To ask a model or system for a value or prediction.
We query the latest policy for its next action. - resample (verb)
- To draw a new sample again, replacing an older one.
The policy gradient resamples the action at each state. - visitation distribution (noun)
- The pattern of which states a policy tends to visit over time.
Old states still come from an outdated visitation distribution. - benign (adjective)
- Not harmful; causing little or no serious problem.
This remaining bias tends to be benign in practice. - stable (adjective)
- Not prone to sudden failure or divergence during training.
Separate networks are more stable to train. - sample-efficient (adjective)
- Able to learn well using relatively little collected data.
A shared network can be more sample-efficient. - worker (noun)
- A separate process or machine running part of a computation.
Multiple simulator workers collect data in parallel. - marginally (adverb)
- By a small amount.
Asynchronous updates use parameters that are marginally older. - trade-off (noun)
- A balance between two good things where you cannot fully have both.
There's a trade-off between network designs. - target value (noun)
- The value a model is trained to match during learning.
The target value is computed from the latest policy. - unbiased (adjective)
- Not systematically wrong on average.
Resampling the action keeps the gradient estimate unbiased. - share features (phrase)
- To reuse the same learned internal representations across different tasks.
A shared trunk lets the actor and critic share features.
Chapters
- 0:00 Architecture design
- 2:01 Online actor-critic in practice
- 5:27 Can we remove the on-policy assumption entirely?
- 7:29 Let's see what that looks like
- 10:33 Fixing the value function
- 13:59 Fixing the policy update
- 16:18 What else is left?
- 17:19 Some implementation details
← Lecture 6, Part 2: Discount Factors in Actor-Critic · Lecture 6, Part 4: Eligibility Traces and GAE →
