Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 25 of 99 · 3:36
Lecture 6, Part 5: Actor-Critic Summary and Examples
Study guide
What this lecture covers
This short closing part of the actor-critic lecture reviews the full set of ideas covered across the previous parts: the actor-critic algorithm, discount factors, network design choices, and the various baseline and n-step tricks used to control variance. It then points to specific papers that put these ideas into practice.
Key ideas
- Actor-critic recap: the actor is the policy, the critic is the value function, and the algorithm follows the sample, evaluate, and gradient-update structure with substantially lower variance than plain policy gradients.
- Discount factor: makes policy evaluation feasible over infinite horizons and can be interpreted either as a preference for near-term reward or as a variance-reduction trick.
- Architecture and batching choices: shared versus separate actor/critic networks, and batch versus online updates, including using parallelism to get larger effective batch sizes.
- Baselines and control variates: state-dependent baselines keep the gradient unbiased, and action-dependent control variates can reduce variance further at the cost of needing a correction term.
- GAE paper: "High-Dimensional Continuous Control with Generalized Advantage Estimation" applies the GAE trick in a batch-mode actor-critic setting on continuous control tasks such as a running humanoid.
- A3C paper: "Asynchronous Methods for Deep Reinforcement Learning" uses online, asynchronous, parallel actor-critic with image inputs, an n-step return with
n=4, and a shared network for actor and critic. - Foundational reading: the classic "Policy Gradient Methods for Reinforcement Learning with Function Approximation" paper underlies much of the lecture's material, including the causality trick.
Before you watch
- Watch Parts 1 through 4 of this lecture, since this part is a review and reference list rather than new material.
Check your understanding
- What are the three boxes in the general recipe for actor-critic algorithms, and what does each one do in this setting?
- How does the asynchronous methods (A3C) paper's use of n-step returns and a shared network relate to the design choices discussed earlier in the lecture?
- What does the GAE paper's estimator combine to control bias and variance?
Vocabulary
- recap (noun)
- A short summary of what was covered before.
This part is a recap of the whole actor-critic lecture. - actor-critic (noun)
- A method that learns a policy (actor) and a value function (critic) at the same time.
Actor-critic combines the strengths of value learning and policy gradients. - policy gradient (noun)
- A method that improves a policy by following the direction that increases expected reward.
Plain policy gradient has much higher variance than actor-critic. - discount factor (noun)
- A number less than 1 that lowers the value of rewards received later in time.
The discount factor makes evaluating an infinite horizon feasible. - variance reduction (phrase)
- Making an estimate less noisy without changing what it is estimating on average.
Baselines are a common variance reduction trick. - architecture (noun)
- The overall design of a neural network, including how its parts are connected.
Shared versus separate networks is an architecture choice. - batching (noun)
- Grouping several samples together before updating the model.
Batching multiple trajectories can lower gradient noise. - online update (phrase)
- Updating the model immediately after each single step, rather than waiting to collect a batch.
Online updates react faster but use noisier data. - parallelism (noun)
- Running many copies of a process at the same time.
Parallelism lets us gather a large effective batch quickly. - baseline (noun)
- A reference value subtracted from an estimate to reduce its noise.
A state-dependent baseline keeps the gradient unbiased. - control variate (noun)
- A quantity with a known expected effect, used to reduce variance in an estimate.
An action-dependent control variate reduces variance further. - generalized advantage estimation (GAE) (noun)
- A method that blends many n-step return estimates into one to balance bias and variance.
The GAE paper applies this trick to continuous control tasks. - continuous control (noun)
- Reinforcement learning tasks where actions are real numbers (like force or angle) rather than a fixed set of choices.
A running humanoid is a continuous control task. - asynchronous (adjective)
- Happening independently and not at the same time as other processes.
A3C uses asynchronous parallel workers to collect data. - shared network (phrase)
- One network whose lower layers are used by both the actor and the critic.
A3C uses a shared network for actor and critic. - foundational (adjective)
- Forming the basic starting point that later work builds on.
This is a foundational paper on policy gradient theory. - function approximation (phrase)
- Using a model, like a neural network, to estimate a value instead of storing it exactly.
The paper covers policy gradients with function approximation. - causality trick (phrase)
- The idea that an action can only affect rewards that come after it, not before.
The causality trick removes terms that don't depend on the action. - influential (adjective)
- Having a strong effect on later work or ideas.
These are some of the most influential actor-critic papers. - pointer (noun)
- A reference that directs you to more information elsewhere.
The lecture ends with pointers to key papers.
Chapters
- 0:00 <Untitled Chapter 1>
- 0:42 Policy Evaluation
- 0:48 Discount Factors
- 1:44 Active Critic Algorithms
- 1:54 Gae Estimator
- 2:19 Asynchronous Methods
- 2:32 Actor Critic Algorithm
- 2:59 Policy Gradients
← Lecture 6, Part 4: Eligibility Traces and GAE · Lecture 7, Part 1: From Actor-Critic to Policy Iteration →
