Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 12 of 99 · 5:50
Lecture 4, Part 4: Types of RL Algorithms
Study guide
What this lecture covers
This part gives a map of the reinforcement learning landscape before the course dives into each family in detail. It answers: what are the main categories of RL algorithms, and how does each one fill in the "green box" (fitting/learning) and "blue box" (policy improvement) steps of the generic RL algorithm anatomy?
After watching, you can name the four main algorithm families, describe what each one learns and how it uses that to improve the policy, and recognize terms like rollout, value-based method, and model-based RL that recur throughout the rest of the course.
Key ideas
- Policy gradient methods: directly estimate the derivative of the RL objective with respect to the policy parameters
thetaand take a gradient ascent step; the green box just sums rewards along sampled trajectories (rollouts). - Value-based methods: fit
V(s)orQ(s, a), usually with a neural network, and often represent the policy only implicitly asarg max_a Q(s, a)rather than as an explicit distribution. - Actor-critic methods: a hybrid that fits a value function or Q-function like value-based methods, then uses it in the blue box to compute a more accurate policy gradient step.
- Model-based RL: learns a transition model
p(s_{t+1} | st, at), then either plans directly with it, backpropagates rewards through it, uses it to learn a value/Q-function via dynamic programming, or uses it to generate extra synthetic data for a model-free algorithm. - Rollout: sampling from the policy step by step in the environment; the term comes from "unrolling" the policy over time.
- Second-order tricks: backpropagating rewards through a learned model for policy improvement is simple in principle but often needs second-order optimization methods for numerical stability.
Before you watch
- Review the Q-function and value function definitions and the green-box/blue-box anatomy of an RL algorithm from the previous part of this lecture.
Check your understanding
- What does a policy gradient algorithm compute in the blue box that a value-based method does not?
- Name three different ways a learned transition model can be used to improve a policy.
- How does an actor-critic method combine ideas from value-based and policy gradient methods?
- In a pure value-based method, how is the policy represented?
Vocabulary
- landscape (noun)
- The overall range of approaches or options available in a field.
This lecture maps the reinforcement learning landscape. - family (noun)
- A group of related methods that share a common approach.
Policy gradient is one family of RL algorithms. - gradient ascent (noun)
- An optimization method that moves parameters to increase a value, like reward.
Policy gradient methods take a gradient ascent step. - value-based (adjective)
- Describes a method that learns a value or Q-function rather than a direct policy.
Value-based methods often represent the policy only implicitly. - implicitly (adverb)
- Indirectly, without being stated or represented directly.
The policy is represented implicitly through the Q-function. - actor-critic (noun)
- A method that combines a learned value function with a policy gradient update.
Actor-critic methods use a critic to improve the policy gradient. - hybrid (adjective)
- Combining two different approaches into one method.
Actor-critic is a hybrid of value-based and policy gradient methods. - model-based RL (noun)
- A reinforcement learning approach that learns a model of the environment's dynamics.
Model-based RL learns to predict the next state from actions. - transition model (noun)
- A learned function that predicts the next state given a current state and action.
A transition model estimates how the environment changes. - plan (verb)
- To choose a sequence of actions in advance using a model.
The agent can plan using its learned transition model. - synthetic data (noun)
- Artificially generated data rather than data collected from the real world.
A model can generate synthetic data for training. - rollout (noun)
- Running a policy step by step through an environment to collect a trajectory.
A rollout samples actions from the policy over time. - second-order (adjective)
- Using information about curvature, not just slope, in optimization.
Second-order tricks help stabilize backpropagation through a model. - numerical stability (noun)
- The property of a calculation staying accurate and well-behaved, without extreme values.
Second-order methods are used for numerical stability. - category (noun)
- A group of things sharing a common set of features.
The lecture names four main algorithm categories. - estimate (verb)
- To calculate an approximate value for something.
Policy gradient methods estimate the derivative of the objective. - derivative (noun)
- A measure of how fast a function's output changes as its input changes.
Policy gradient estimates the derivative of expected reward. - dynamic programming (noun)
- A method for solving problems by combining solutions to smaller subproblems.
A model can be used to learn a value function via dynamic programming. - unroll (verb)
- To play out a process step by step over time.
A rollout comes from unrolling the policy over time. - in principle (phrase)
- Theoretically possible, even if it's difficult in practice.
Backpropagating through a model is simple in principle.
Chapters
- 0:00 <Untitled Chapter 1>
- 0:22 Types of RL algorithms
- 1:50 Model-based RL algorithms
- 3:56 Value function based algorithms
- 4:40 Direct policy gradients
← Lecture 4, Part 3: Q-Functions and Value Functions · Lecture 4, Part 5: Comparing RL Algorithms →
