Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 21 of 99 · 25:13
Lecture 6, Part 1: Actor-Critic and Value Functions
Study guide
What this lecture covers
This part opens Berkeley CS285's lecture on actor-critic algorithms by asking how to lower the variance of the policy gradient estimator introduced in the previous lecture. It recaps the REINFORCE algorithm and its reward-to-go term, then shows that this term is a high-variance single-sample estimate of an expectation that a learned function can approximate more accurately.
By the end you should understand why the advantage function is the natural quantity to multiply into the policy gradient, why fitting the value function alone is more convenient than fitting the Q function, and how Monte Carlo and bootstrapped targets differ when training that value function.
Key ideas
- Reward-to-go (
q_hat): the sum of rewards from a time step onward in one sampled trajectory; it is an unbiased but high-variance estimate of the true expected future reward. - Advantage function:
A(s,a) = Q(s,a) - V(s), measuring how much better an action is than the policy's average behavior in that state. - Why fit V instead of Q:
Q(s,a)can be rewritten as the current reward plus the expected next-state value, so approximating the next state with the one actually observed lets the advantage be expressed using onlyV, which depends on fewer inputs and is easier to learn. - Policy evaluation: the process of estimating
V^pi, the value of the current policy at every state; the RL objective itself is the expected value ofV^piover the initial states. - Monte Carlo target: summing actual observed rewards after a state gives a target for fitting
V; it stays low bias but only reduces variance through the function approximator's generalization across similar states. - Bootstrapped target: using the observed reward plus the current value estimate of the next state gives a lower-variance but more biased target.
- Bias-variance trade-off: approximate value functions introduce some bias into the policy gradient, but this is usually worth it for the large drop in variance.
Before you watch
- Review the REINFORCE algorithm and the "orange/green/blue box" anatomy of policy gradient methods from the previous lecture, since this lecture modifies the green box.
- Be comfortable with the baseline trick for policy gradients, since the advantage function extends that idea to a state-dependent baseline.
Check your understanding
- Why does using a single sampled trajectory to estimate the reward-to-go lead to high variance?
- Derive how the advantage function follows from subtracting the value function from the Q function as a baseline.
- Why is it more convenient to fit
V(s)thanQ(s,a)orA(s,a)directly? - Compare the Monte Carlo and bootstrapped targets for fitting a value function: what does each trade off?
- How does policy evaluation relate to computing the RL objective
J(theta)?
Vocabulary
- actor-critic (noun)
- A method that combines a learned value function with a policy gradient update.
Actor-critic methods lower the variance of policy gradients. - value function (noun)
- A function estimating the expected future reward from a state under a policy.
The value function replaces the reward-to-go estimate. - single-sample estimate (noun)
- An approximation computed from just one collected data point.
Reward-to-go is a single-sample estimate with high variance. - advantage function (noun)
- A function measuring how much better an action is than the policy's average in that state.
The advantage function guides the actor-critic update. - policy evaluation (noun)
- The process of estimating how good a given policy is.
Policy evaluation computes the value function for the current policy. - Monte Carlo target (noun)
- A training target formed by summing actual observed rewards.
A Monte Carlo target has low bias but high variance. - bootstrapped target (noun)
- A training target that uses a current estimate instead of the full true outcome.
A bootstrapped target has lower variance but more bias. - bias-variance trade-off (noun)
- The balance between an estimate's systematic error and its randomness.
Approximating the value function trades bias for lower variance. - function approximator (noun)
- A model, like a neural network, used to estimate a complex function.
A function approximator generalizes value estimates across states. - generalization (noun)
- The ability to perform well on new, unseen situations.
Generalization across similar states reduces variance. - recap (verb)
- To briefly summarize something covered earlier.
The lecture recaps the REINFORCE algorithm. - approximate (verb)
- To estimate a value that is close to the true one.
A learned function can approximate the expectation more accurately. - natural quantity (phrase)
- A value that fits logically into a formula or reasoning.
The advantage function is the natural quantity to use here. - convenient (adjective)
- Easy or practical to use.
Fitting V is more convenient than fitting Q. - depend on fewer inputs (phrase)
- Requiring less information to compute a result.
V depends on fewer inputs than Q. - low bias (phrase)
- Having little systematic error on average.
A Monte Carlo target stays low bias. - observed reward (noun)
- The actual reward received during a real trajectory.
The bootstrapped target uses the observed reward plus an estimate. - worth it (phrase)
- Justified by the benefit it provides, despite a cost.
The added bias is usually worth it for the drop in variance. - current value estimate (noun)
- The value function's present prediction, before further training.
The bootstrapped target uses the current value estimate of the next state. - state-dependent (adjective)
- Changing depending on which state is being considered.
The advantage function extends the baseline to a state-dependent one.
Chapters
- 0:00 Intro
- 0:17 Recap: policy gradients
- 2:15 Improving the policy gradient
- 5:17 What about the baseline?
- 7:55 State & state-action value functions
- 11:22 Value function fitting
- 17:39 Monte Carlo evaluation with function approximation
- 20:46 Can we do better?
- 23:25 Policy evaluation examples
← Lecture 5, Part 6: The Natural Policy Gradient · Lecture 6, Part 2: Discount Factors in Actor-Critic →
