Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 22 of 99 · 17:29
Lecture 6, Part 2: Discount Factors in Actor-Critic
Study guide
What this lecture covers
Continuing directly from Part 1, this part assembles the pieces discussed so far into a complete batch actor-critic algorithm: sample trajectories, fit a value function, compute advantages, form the policy gradient, and take a gradient step. It then addresses a problem that arises in continuing (infinite-horizon) tasks: bootstrapped value updates can grow without bound.
You come away able to state the five-step batch actor-critic algorithm, explain why a discount factor is introduced, and distinguish two ways of adding discounting to a Monte Carlo policy gradient, including why one of them is generally preferred in practice.
Key ideas
- Batch actor-critic algorithm: generate rollouts, fit
V(s)to sampled rewards, computeA(s,a) = r + V(s') - V(s)for each transition, form the policy gradient from advantages, then update the policy. - Infinite-horizon problem: repeatedly bootstrapping value estimates in a task with no fixed end can push values toward infinity, especially when rewards are always positive.
- Discount factor (
gamma): a number between 0 and 1 (commonly 0.9 to 0.999) that down-weights future rewards, keeping value estimates finite. - Death-state interpretation: gamma can be viewed as adding a fifth "death" state to the MDP that the agent enters with probability
1 - gammaat each step, after which reward is always zero. - Two ways to discount the Monte Carlo policy gradient: multiplying only the reward-to-go by
gamma^(t'-t)(matches the critic version), versus discounting the whole trajectory from the start, which also discounts the gradient term itself bygamma^(t-1). - Preferred choice: the first option is generally used in practice, because it avoids treating later decisions as less important, which is not the behavior wanted for continuing tasks like locomotion.
- Online actor-critic: with a discount factor, actor-critic can run fully online, updating the value function and policy after every single environment step using only the current transition
(s, a, s', r).
Before you watch
- Watch Part 1 of this lecture, since it derives the advantage function and value-fitting targets used here.
- Recall the causality trick from earlier policy gradient material, since it is used to compare the two discounting options.
Check your understanding
- Why can bootstrapped value estimates diverge in an infinite-horizon task without discounting?
- Explain the "death state" interpretation of the discount factor.
- What is the mathematical difference between discounting only the reward-to-go versus discounting the entire trajectory, and why does it matter?
- Why is the fully online actor-critic algorithm able to update using only
s,a,s', andr, without needing further future states? - What trade-off does introducing a discount factor make between bias and variance?
Vocabulary
- discount factor (noun)
- A number between 0 and 1 that reduces the importance of future rewards.
The discount factor keeps value estimates finite. - continuing task (noun)
- A task with no fixed ending point.
Continuing tasks like locomotion need a discount factor. - infinite-horizon (adjective)
- Describes a task that could run indefinitely, with no fixed end.
Infinite-horizon tasks risk unbounded value growth. - unbounded (adjective)
- Having no upper or lower limit.
Without discounting, values can grow unbounded. - down-weight (verb)
- To give something less importance in a calculation.
The discount factor down-weights distant future rewards. - transition (noun)
- A single step from one state to the next after taking an action.
Online actor-critic updates after every transition. - rollout (noun)
- A sequence of states and actions generated by running a policy.
A batch actor-critic algorithm generates rollouts first. - online (adjective)
- Happening step by step during live interaction, not from a stored batch.
Online actor-critic updates after each single step. - finite (adjective)
- Having a definite, limited size, not infinite.
Discounting keeps the value estimate finite. - locomotion (noun)
- The ability to move from place to place.
Locomotion is an example of a continuing task. - assemble (verb)
- To put separate pieces together to form something complete.
This part assembles the full batch actor-critic algorithm. - advantage (noun)
- How much better an action is compared to the average for that state.
The advantage is computed for each transition. - death state (noun)
- An imagined extra state added to explain why future rewards are discounted.
The death state interpretation explains the discount factor. - down-weighted (adjective)
- Given less importance in a calculation.
Distant rewards are down-weighted by the discount factor. - gradient term (noun)
- The part of a formula representing the computed gradient.
Discounting the whole trajectory also discounts the gradient term. - step (noun)
- One single unit of interaction or update in a process.
The agent takes one step in the environment at a time. - generally preferred (phrase)
- Chosen more often because it works better in most cases.
The first discounting option is generally preferred in practice. - later decisions (phrase)
- Choices made further along in a sequence or task.
We don't want to treat later decisions as less important. - batch (noun)
- A group of examples processed together in one step.
A batch actor-critic algorithm processes many rollouts together. - bound (verb)
- To keep a value within a set limit.
The discount factor bounds the value estimate.
Chapters
- 0:00 Constructing actor critic
- 1:57 Understanding discount factors
- 6:35 Discounting policy gradients
- 14:43 Online actor critic algorithm
← Lecture 6, Part 1: Actor-Critic and Value Functions · Lecture 6, Part 3: Implementing Actor-Critic →
