Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 25 of 99 · 3:36

Lecture 6, Part 5: Actor-Critic Summary and Examples

CS 285: Lecture 6, Part 5 on YouTube

Study guide

What this lecture covers

This short closing part of the actor-critic lecture reviews the full set of ideas covered across the previous parts: the actor-critic algorithm, discount factors, network design choices, and the various baseline and n-step tricks used to control variance. It then points to specific papers that put these ideas into practice.

Key ideas

  • Actor-critic recap: the actor is the policy, the critic is the value function, and the algorithm follows the sample, evaluate, and gradient-update structure with substantially lower variance than plain policy gradients.
  • Discount factor: makes policy evaluation feasible over infinite horizons and can be interpreted either as a preference for near-term reward or as a variance-reduction trick.
  • Architecture and batching choices: shared versus separate actor/critic networks, and batch versus online updates, including using parallelism to get larger effective batch sizes.
  • Baselines and control variates: state-dependent baselines keep the gradient unbiased, and action-dependent control variates can reduce variance further at the cost of needing a correction term.
  • GAE paper: "High-Dimensional Continuous Control with Generalized Advantage Estimation" applies the GAE trick in a batch-mode actor-critic setting on continuous control tasks such as a running humanoid.
  • A3C paper: "Asynchronous Methods for Deep Reinforcement Learning" uses online, asynchronous, parallel actor-critic with image inputs, an n-step return with n=4, and a shared network for actor and critic.
  • Foundational reading: the classic "Policy Gradient Methods for Reinforcement Learning with Function Approximation" paper underlies much of the lecture's material, including the causality trick.

Before you watch

  • Watch Parts 1 through 4 of this lecture, since this part is a review and reference list rather than new material.

Check your understanding

  1. What are the three boxes in the general recipe for actor-critic algorithms, and what does each one do in this setting?
  2. How does the asynchronous methods (A3C) paper's use of n-step returns and a shared network relate to the design choices discussed earlier in the lecture?
  3. What does the GAE paper's estimator combine to control bias and variance?

Vocabulary

recap (noun)
A short summary of what was covered before.
This part is a recap of the whole actor-critic lecture.
actor-critic (noun)
A method that learns a policy (actor) and a value function (critic) at the same time.
Actor-critic combines the strengths of value learning and policy gradients.
policy gradient (noun)
A method that improves a policy by following the direction that increases expected reward.
Plain policy gradient has much higher variance than actor-critic.
discount factor (noun)
A number less than 1 that lowers the value of rewards received later in time.
The discount factor makes evaluating an infinite horizon feasible.
variance reduction (phrase)
Making an estimate less noisy without changing what it is estimating on average.
Baselines are a common variance reduction trick.
architecture (noun)
The overall design of a neural network, including how its parts are connected.
Shared versus separate networks is an architecture choice.
batching (noun)
Grouping several samples together before updating the model.
Batching multiple trajectories can lower gradient noise.
online update (phrase)
Updating the model immediately after each single step, rather than waiting to collect a batch.
Online updates react faster but use noisier data.
parallelism (noun)
Running many copies of a process at the same time.
Parallelism lets us gather a large effective batch quickly.
baseline (noun)
A reference value subtracted from an estimate to reduce its noise.
A state-dependent baseline keeps the gradient unbiased.
control variate (noun)
A quantity with a known expected effect, used to reduce variance in an estimate.
An action-dependent control variate reduces variance further.
generalized advantage estimation (GAE) (noun)
A method that blends many n-step return estimates into one to balance bias and variance.
The GAE paper applies this trick to continuous control tasks.
continuous control (noun)
Reinforcement learning tasks where actions are real numbers (like force or angle) rather than a fixed set of choices.
A running humanoid is a continuous control task.
asynchronous (adjective)
Happening independently and not at the same time as other processes.
A3C uses asynchronous parallel workers to collect data.
shared network (phrase)
One network whose lower layers are used by both the actor and the critic.
A3C uses a shared network for actor and critic.
foundational (adjective)
Forming the basic starting point that later work builds on.
This is a foundational paper on policy gradient theory.
function approximation (phrase)
Using a model, like a neural network, to estimate a value instead of storing it exactly.
The paper covers policy gradients with function approximation.
causality trick (phrase)
The idea that an action can only affect rewards that come after it, not before.
The causality trick removes terms that don't depend on the action.
influential (adjective)
Having a strong effect on later work or ideas.
These are some of the most influential actor-critic papers.
pointer (noun)
A reference that directs you to more information elsewhere.
The lecture ends with pointers to key papers.

Chapters

← Lecture 6, Part 4: Eligibility Traces and GAE · Lecture 7, Part 1: From Actor-Critic to Policy Iteration →