Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 22 of 99 · 17:29

Lecture 6, Part 2: Discount Factors in Actor-Critic

CS 285: Lecture 6, Part 2 on YouTube

Study guide

What this lecture covers

Continuing directly from Part 1, this part assembles the pieces discussed so far into a complete batch actor-critic algorithm: sample trajectories, fit a value function, compute advantages, form the policy gradient, and take a gradient step. It then addresses a problem that arises in continuing (infinite-horizon) tasks: bootstrapped value updates can grow without bound.

You come away able to state the five-step batch actor-critic algorithm, explain why a discount factor is introduced, and distinguish two ways of adding discounting to a Monte Carlo policy gradient, including why one of them is generally preferred in practice.

Key ideas

  • Batch actor-critic algorithm: generate rollouts, fit V(s) to sampled rewards, compute A(s,a) = r + V(s') - V(s) for each transition, form the policy gradient from advantages, then update the policy.
  • Infinite-horizon problem: repeatedly bootstrapping value estimates in a task with no fixed end can push values toward infinity, especially when rewards are always positive.
  • Discount factor (gamma): a number between 0 and 1 (commonly 0.9 to 0.999) that down-weights future rewards, keeping value estimates finite.
  • Death-state interpretation: gamma can be viewed as adding a fifth "death" state to the MDP that the agent enters with probability 1 - gamma at each step, after which reward is always zero.
  • Two ways to discount the Monte Carlo policy gradient: multiplying only the reward-to-go by gamma^(t'-t) (matches the critic version), versus discounting the whole trajectory from the start, which also discounts the gradient term itself by gamma^(t-1).
  • Preferred choice: the first option is generally used in practice, because it avoids treating later decisions as less important, which is not the behavior wanted for continuing tasks like locomotion.
  • Online actor-critic: with a discount factor, actor-critic can run fully online, updating the value function and policy after every single environment step using only the current transition (s, a, s', r).

Before you watch

  • Watch Part 1 of this lecture, since it derives the advantage function and value-fitting targets used here.
  • Recall the causality trick from earlier policy gradient material, since it is used to compare the two discounting options.

Check your understanding

  1. Why can bootstrapped value estimates diverge in an infinite-horizon task without discounting?
  2. Explain the "death state" interpretation of the discount factor.
  3. What is the mathematical difference between discounting only the reward-to-go versus discounting the entire trajectory, and why does it matter?
  4. Why is the fully online actor-critic algorithm able to update using only s, a, s', and r, without needing further future states?
  5. What trade-off does introducing a discount factor make between bias and variance?

Vocabulary

discount factor (noun)
A number between 0 and 1 that reduces the importance of future rewards.
The discount factor keeps value estimates finite.
continuing task (noun)
A task with no fixed ending point.
Continuing tasks like locomotion need a discount factor.
infinite-horizon (adjective)
Describes a task that could run indefinitely, with no fixed end.
Infinite-horizon tasks risk unbounded value growth.
unbounded (adjective)
Having no upper or lower limit.
Without discounting, values can grow unbounded.
down-weight (verb)
To give something less importance in a calculation.
The discount factor down-weights distant future rewards.
transition (noun)
A single step from one state to the next after taking an action.
Online actor-critic updates after every transition.
rollout (noun)
A sequence of states and actions generated by running a policy.
A batch actor-critic algorithm generates rollouts first.
online (adjective)
Happening step by step during live interaction, not from a stored batch.
Online actor-critic updates after each single step.
finite (adjective)
Having a definite, limited size, not infinite.
Discounting keeps the value estimate finite.
locomotion (noun)
The ability to move from place to place.
Locomotion is an example of a continuing task.
assemble (verb)
To put separate pieces together to form something complete.
This part assembles the full batch actor-critic algorithm.
advantage (noun)
How much better an action is compared to the average for that state.
The advantage is computed for each transition.
death state (noun)
An imagined extra state added to explain why future rewards are discounted.
The death state interpretation explains the discount factor.
down-weighted (adjective)
Given less importance in a calculation.
Distant rewards are down-weighted by the discount factor.
gradient term (noun)
The part of a formula representing the computed gradient.
Discounting the whole trajectory also discounts the gradient term.
step (noun)
One single unit of interaction or update in a process.
The agent takes one step in the environment at a time.
generally preferred (phrase)
Chosen more often because it works better in most cases.
The first discounting option is generally preferred in practice.
later decisions (phrase)
Choices made further along in a sequence or task.
We don't want to treat later decisions as less important.
batch (noun)
A group of examples processed together in one step.
A batch actor-critic algorithm processes many rollouts together.
bound (verb)
To keep a value within a set limit.
The discount factor bounds the value estimate.

Chapters

← Lecture 6, Part 1: Actor-Critic and Value Functions · Lecture 6, Part 3: Implementing Actor-Critic →