Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 14 of 99 · 3:04

Lecture 4, Part 6: Examples of Deep RL Algorithms

CS 285: Lecture 4, Part 6 on YouTube

Study guide

What this lecture covers

This short closing part of the lecture rounds things out with concrete examples rather than new theory. It names specific algorithms within each family covered earlier (value-function fitting, policy gradient, actor-critic, model-based) and shows short video results so you can see what these methods actually produce.

After watching, you can recognize the names of common deep RL algorithms grouped by family, and connect each to a real result: playing Atari from pixels, robotic manipulation, and learning to walk.

Key ideas

  • Value function fitting methods: named examples include Q-learning, DQN, and temporal difference learning.
  • Policy gradient methods: named examples include REINFORCE, natural gradient, TRPO (trust region policy optimization), and PPO.
  • Actor-critic methods: named examples include A3C (asynchronous advantage actor-critic), soft actor-critic, and DDPG.
  • Model-based methods: named examples include Dyna, guided policy search, MPPI, and SVG.
  • Atari from pixels: a 2013 DQN-style result uses Q-learning with convolutional networks to output a Q-value per discrete action, then selects the action with the highest value.
  • Robot grasping and walking: a guided policy search result (model-based, using learned dynamics plus image-based convolutional networks) performs robotic skills, and a TRPO-based actor-critic result trains a simulated humanoid to walk; a continuous-action variant of Q-learning is used for a grasping robot.

Before you watch

  • Watch the earlier parts of this lecture on Q-functions/value functions and the four algorithm families, since this part assumes you already know what value-based, policy gradient, actor-critic, and model-based methods are.

Check your understanding

  1. Which family does DQN belong to, and how does it choose actions in a discrete-action game like Atari?
  2. What two components does the guided policy search example combine to perform robotic skills?
  3. Which algorithm type is used in the humanoid walking example, and why is it described as a hybrid?

Vocabulary

non-technical (adjective)
Explained in a simple way, without deep mathematical detail.
This is a quick, non-technical tour of RL algorithms.
temporal difference learning (noun)
A method that updates value estimates using the difference between predictions at consecutive time steps.
Temporal difference learning is a value-based method.
trust region (noun)
A limited area around the current parameters where an update is trusted to be safe.
TRPO limits updates to a trust region.
asynchronous (adjective)
Running independently at different times, not perfectly synchronized.
A3C runs asynchronous actor-critic workers.
advantage (noun)
How much better an action is compared to the average for that state.
Advantage actor-critic uses the advantage as its learning signal.
soft actor-critic (noun)
An actor-critic algorithm that encourages exploration by rewarding randomness in the policy.
Soft actor-critic balances reward and exploration.
guided policy search (noun)
A model-based method that uses a learned model to guide training a policy.
Guided policy search combines learned dynamics with a neural network policy.
discrete action (noun)
A choice from a fixed, countable set of options.
Atari games use a discrete action space.
continuous action (noun)
An action represented by a real number rather than a fixed category.
Robot grasping often uses continuous actions.
humanoid (noun)
A robot or simulated character shaped like a human.
A humanoid learns to walk using TRPO.
manipulation (noun)
The physical handling and moving of objects, especially by a robot.
Robotic manipulation is shown as an example result.
family (noun)
A group of related algorithms that share a common approach.
Each algorithm family gets a few named examples.
concrete (adjective)
Specific and real, not abstract or theoretical.
This part gives concrete examples rather than new theory.
named example (noun)
A specific, labeled instance used to illustrate a general category.
DQN is a named example of a value-based method.
pixels (noun)
The tiny individual points of color that make up a digital image.
DQN learns to play Atari directly from pixels.
convolutional network (noun)
A neural network built around convolution operations, often used for images.
A convolutional network processes the Atari screen image.
select (verb)
To choose one option from several available ones.
The agent selects the action with the highest Q-value.
learned dynamics (noun)
A model of how the environment changes, learned from data.
Guided policy search uses learned dynamics.
result (noun)
The outcome produced by running an experiment or method.
The video shows several deep RL results.
recognize (verb)
To identify something you have seen or learned before.
You can recognize common deep RL algorithm names.

Chapters

← Lecture 4, Part 5: Comparing RL Algorithms · Lecture 5, Part 1: Deriving the Policy Gradient →