Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 24 of 99 · 15:54
Lecture 6, Part 4: Eligibility Traces and GAE
Study guide
What this lecture covers
This part returns to the bias-variance trade-off at the heart of actor-critic methods. It compares the low-variance, biased actor-critic advantage estimator against the unbiased, higher-variance Monte Carlo estimator, then shows how a state-dependent baseline (the value function) keeps the policy gradient unbiased while still reducing variance.
By the end you should be able to explain n-step returns as a way to interpolate between the two extremes, and understand how generalized advantage estimation (GAE) averages over all possible n-step returns using a single parameter.
Key ideas
- State-dependent baseline: subtracting any function of the state alone, such as the value function, keeps the policy gradient unbiased while lowering variance, unlike the biased actor-critic critic term.
- Action-dependent baseline (control variates): subtracting a Q function lowers variance further but no longer integrates to zero, requiring a correction term; this correction can sometimes be computed with low or zero variance, as in the Q-Prop method mentioned in the lecture.
- n-step returns: sum actual rewards for
nsteps, then bootstrap with the value function for the remainder, trading bias for variance depending on how largenis. - Bias-variance intuition: rewards and value estimates near the present have lower variance and are more reliable, so cutting off the Monte Carlo sum earlier and relying on the value function later tends to work better than either extreme.
- Generalized advantage estimation (GAE): a weighted average of all n-step returns, using an exponential weighting controlled by a parameter
lambda, analogous to a discount factor. - GAE recursive form: GAE can be written as a discounted, lambda-weighted sum of single-step advantage terms (
delta), giving an efficient way to compute the full weighted average. - Discount and lambda as variance control: both
gammaandlambdareduce variance by down-weighting distant, less certain rewards, at the cost of some bias.
Before you watch
- Watch Part 3 of this lecture on implementing actor-critic, since the advantage estimator discussed there is the baseline being improved here.
- Review the discount factor material from Part 2, since GAE's
lambdaplays an analogous role.
Check your understanding
- Why does subtracting a state-dependent baseline keep the policy gradient unbiased, while subtracting an action-dependent baseline does not?
- What does the n-step return estimator trade off as
nincreases? - Why is variance typically higher for rewards further in the future than for rewards close to the present?
- How does GAE combine all possible n-step returns into a single estimator?
- In what sense does the
lambdaparameter in GAE play a role similar to the discount factorgamma?
Vocabulary
- bias-variance trade-off (phrase)
- The balance between an estimate that is systematically wrong (biased) and one that is noisy (has high variance).
Choosing how far to bootstrap is a bias-variance trade-off. - actor-critic (noun)
- A reinforcement learning method that learns both a policy (actor) and a value function (critic) together.
The actor-critic method uses the critic to estimate how good actions are. - policy gradient (noun)
- A method that improves a policy by pushing its parameters in the direction that increases expected reward.
Policy gradient methods can have high variance without a good baseline. - baseline (noun)
- A reference value subtracted from a reward estimate to reduce noise without changing the expected result.
Subtracting a baseline lowers variance in the policy gradient. - state-dependent (adjective)
- Depending only on the current state, not on the action taken.
A state-dependent baseline uses only the value function of the state. - action-dependent (adjective)
- Depending on both the state and the action chosen.
An action-dependent baseline uses the Q function instead of just the value function. - unbiased (adjective)
- Correct on average, with no systematic error.
The Monte Carlo estimator is unbiased but noisy. - value function (noun)
- A function that estimates the expected total future reward starting from a state.
The critic learns the value function to reduce variance. - control variate (noun)
- A statistical trick that lowers variance in an estimate by subtracting something with a known expected value.
Using the Q function as a control variate needs a correction term. - correction term (phrase)
- An extra piece added to fix a bias introduced by an approximation.
Q-Prop adds a correction term to keep the estimate unbiased. - n-step return (noun)
- An estimate of future reward that sums real rewards for n steps, then estimates the rest with the value function.
A 5-step return mixes real rewards with a bootstrapped estimate. - bootstrap (verb)
- To estimate a future value using the network's own current prediction instead of waiting for the real outcome.
After n steps we bootstrap using the value function. - Monte Carlo estimator (noun)
- An estimate built by averaging results from actual sampled outcomes rather than a model.
The Monte Carlo estimator uses the full sum of real rewards. - generalized advantage estimation (GAE) (noun)
- A method that combines all possible n-step returns into one estimate using a weighting parameter.
GAE averages over many n-step returns to balance bias and variance. - advantage (noun)
- How much better an action is compared to the average action in that state.
The advantage tells us if this action was better than expected. - exponential weighting (phrase)
- Giving less and less importance to items further away, following a repeated multiplication pattern.
GAE uses exponential weighting controlled by lambda. - recursive (adjective)
- Defined in terms of itself, calculated step by step from previous results.
GAE has a recursive form that is efficient to compute. - discount factor (noun)
- A number between 0 and 1 that reduces the importance of future rewards.
The discount factor gamma makes distant rewards count less. - trade off (verb)
- To accept less of one good thing in order to get more of another.
Increasing n trades off bias for variance. - delta (temporal difference term) (noun)
- The single-step error between a predicted value and the value estimated from one real reward.
GAE is written as a sum of delta terms. - interpolate (verb)
- To find a value that lies between two known extremes.
N-step returns interpolate between full Monte Carlo and one-step bootstrapping.
Chapters
← Lecture 6, Part 3: Implementing Actor-Critic · Lecture 6, Part 5: Actor-Critic Summary and Examples →
