Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 17 of 99 · 14:52
Lecture 5, Part 3: Reducing Variance with Causality and Baselines
Study guide
What this lecture covers
This part addresses the high-variance problem identified earlier in the lecture by introducing two practical fixes that make policy gradients usable. It answers: how can you lower the variance of the policy gradient estimator without introducing bias?
After watching, you can explain the causality trick that replaces full trajectory reward with "reward to go," derive why subtracting a baseline from the reward leaves the gradient unbiased, and describe what the variance-minimizing optimal baseline looks like.
Key ideas
- Causality: an action at time
tcannot affect rewards at earlier time stepst' < t, since this is always true for any forward-flowing process (unlike the Markov property, which only sometimes holds). - Reward to go: exploiting causality, the reward sum multiplying
grad log piat timetcan be restricted to rewards fromtonward, denotedQ-hat(i, t); this term is a single-sample estimate of the same quantity as the Q-function from the previous lecture. - Lower variance from fewer terms: dropping past rewards from the sum reduces the total quantity being multiplied, which lowers the estimator's variance while keeping it unbiased.
- Baselines: subtracting a constant
b(such as average reward) from the reward before multiplying bygrad log pikeeps the gradient unbiased in expectation, because the extra term's expectation is provably zero — this holds for any choice ofb. - Why baselines help: without centering, an all-positive reward signal raises the probability of every trajectory, including bad ones; subtracting a baseline makes above-average trajectories more likely and below-average ones less likely, matching the intended trial-and-error behavior.
- Optimal baseline: minimizing the variance with respect to
bgivesb = E[g^2 * r] / E[g^2], a gradient-magnitude-weighted average reward, different for each parameter; in practice, the simpler average-reward baseline is used instead.
Before you watch
- Watch the earlier parts of this lecture deriving the policy gradient and demonstrating its high-variance problem.
- Know the identity
p(tau) * grad log p(tau) = grad p(tau)used in the policy gradient derivation.
Check your understanding
- How does causality justify dropping past rewards from the policy gradient sum, and why does this reduce variance?
- Why does subtracting a baseline
bfrom the reward not bias the policy gradient in expectation? - What form does the variance-minimizing optimal baseline take, and why is it rarely used in practice?
- What is "reward to go," and how does it relate to the Q-function from the previous lecture?
Vocabulary
- variance (noun)
- A measure of how spread out or inconsistent estimates are.
This lecture reduces the variance of the policy gradient. - causality (noun)
- The principle that a cause must come before its effect in time.
Causality means an action can't affect past rewards. - reward to go (noun)
- The sum of rewards received from a given time step onward.
Reward to go replaces the full trajectory reward in the gradient. - restrict (verb)
- To limit something to a smaller range or set.
The reward sum is restricted to future time steps. - baseline (noun)
- A reference value subtracted from a signal to reduce its variance.
Subtracting a baseline lowers the policy gradient's variance. - center (verb)
- To shift values so they are balanced around zero.
Centering the rewards helps the gradient behave correctly. - in expectation (phrase)
- True on average, over many repeated samples.
The baseline leaves the gradient unbiased in expectation. - provably (adverb)
- In a way that can be shown to be true with mathematical proof.
The baseline term is provably zero in expectation. - above-average (adjective)
- Better than the typical or mean value.
Above-average trajectories become more likely after subtracting a baseline. - optimal (adjective)
- The best possible choice among the available options.
The optimal baseline minimizes the gradient's variance. - gradient magnitude (noun)
- The size or strength of a gradient vector.
The optimal baseline is weighted by gradient magnitude. - unbiased (adjective)
- Not systematically wrong on average, even if noisy for a single sample.
Subtracting a baseline keeps the estimator unbiased. - single-sample estimate (noun)
- An approximation computed from just one collected data point.
Reward to go is a single-sample estimate of the Q-function. - exploit (verb)
- To make good use of a property or opportunity.
The trick exploits causality to reduce variance. - minimize (verb)
- To reduce something as much as possible.
The optimal baseline minimizes variance. - in practice (phrase)
- In real, applied use rather than only in theory.
A simpler baseline is used in practice. - practical fix (noun)
- A workable solution used to solve a real problem.
Causality and baselines are two practical fixes for high variance. - constant (noun)
- A fixed value that does not change.
The baseline can be any constant b. - signal (noun)
- A piece of information used to guide a decision or update.
An all-positive reward signal raises every trajectory's probability. - trial-and-error (noun)
- Learning by repeatedly trying actions and observing outcomes.
Baselines match the intended trial-and-error behavior.
Chapters
← Lecture 5, Part 2: Intuition and the High-Variance Problem · Lecture 5, Part 4: Off-Policy Policy Gradients with Importance Sampling →
