Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 20 of 99 · 13:32

Lecture 5, Part 6: The Natural Policy Gradient

CS 285: Lecture 5, Part 6 on YouTube

Study guide

What this lecture covers

This closing part of the lecture diagnoses a numerical problem that hurts policy gradients, especially with continuous actions, and derives a fix called the natural (or covariant) policy gradient. It uses a small 1D example — a Gaussian policy with a mean parameter k and a variance parameter sigma — to show why plain gradient ascent struggles, then reframes gradient ascent as a constrained optimization problem to motivate the fix.

After watching, you can explain why some policy parameters need much smaller learning rates than others, describe gradient ascent as maximizing a linearized objective within a parameter-space trust region, and explain how replacing that region with a KL-divergence constraint on the policy distribution produces the natural gradient.

Key ideas

  • Poor conditioning example: in a small Gaussian-policy problem, the gradient with respect to sigma grows much larger than the gradient with respect to k as sigma shrinks, so the raw gradient direction doesn't point toward the optimum — the same issue as optimizing a quadratic with a badly scaled Hessian.
  • Gradient ascent as constrained optimization: a first-order gradient step can be viewed as maximizing a linear (Taylor) approximation of the objective subject to a small step-size constraint ||theta' - theta||^2 <= epsilon in parameter space.
  • Wrong space for the constraint: constraining step size in raw parameter space is awkward, because some parameters change the policy distribution a lot per unit step and others change it very little.
  • KL-divergence constraint: replacing the parameter-space constraint with a constraint on KL(pi_theta' || pi_theta) (a parameterization-independent divergence between distributions) makes step sizes consistent in policy space rather than parameter space.
  • Fisher information matrix: a second-order Taylor expansion of the KL divergence gives a quadratic form using the Fisher information matrix F = E[grad log pi * grad log pi^T], estimated from samples.
  • Natural gradient update: solving the constrained problem gives the update theta <- theta + alpha * F^-1 * grad J(theta); multiplying by F^-1 rescales the gradient so it points toward the optimum, converging faster and requiring less step-size tuning. Natural policy gradient and TRPO are named algorithms that use this idea, with TRPO deriving alpha from a chosen epsilon via conjugate gradient.

Before you watch

  • Watch the earlier parts of this lecture, especially the policy gradient derivation and implementation.
  • Basic familiarity with Taylor expansions, constrained optimization (Lagrange multipliers), and KL divergence is helpful.

Check your understanding

  1. Why does the raw policy gradient point in a poor direction when one parameter (like sigma) has a much larger gradient magnitude than another?
  2. How does viewing gradient ascent as a constrained optimization problem motivate replacing the parameter-space constraint with a KL-divergence constraint?
  3. What role does the Fisher information matrix play in the natural gradient update?
  4. What is the practical benefit of using F^-1 * grad J(theta) instead of the raw gradient?

Vocabulary

diagnose (verb)
To identify the cause of a problem.
This lecture diagnoses a numerical problem in policy gradients.
natural gradient (noun)
A modified gradient update that accounts for how parameters affect the policy's distribution.
The natural gradient rescales the update using distribution geometry.
covariant (adjective)
Changing consistently together with another quantity, regardless of how it's measured.
The natural gradient is also called the covariant gradient.
poor conditioning (noun)
A situation where a function is much steeper in some directions than others, making optimization hard.
Poor conditioning makes gradient ascent struggle.
Hessian (noun)
A matrix of second derivatives describing a function's curvature.
A badly scaled Hessian causes similar problems in optimization.
linearized (adjective)
Approximated using a straight-line version of a curved function.
Gradient ascent maximizes a linearized version of the objective.
Taylor approximation (noun)
An estimate of a function using its value and derivatives at one point.
A first-order Taylor approximation underlies plain gradient ascent.
trust region (noun)
A limited area around current parameters where an update is trusted to be safe.
The step-size constraint defines a trust region.
constraint (noun)
A limit or condition that a solution must satisfy.
The optimization is subject to a step-size constraint.
KL divergence (noun)
A measure of how different one probability distribution is from another.
The natural gradient uses a KL divergence constraint.
parameterization-independent (adjective)
Not affected by how a model's parameters happen to be defined.
KL divergence is parameterization-independent, unlike raw parameter distance.
Fisher information matrix (noun)
A matrix describing how sensitive a distribution is to changes in its parameters.
The Fisher information matrix comes from a second-order KL expansion.
quadratic form (noun)
An expression involving squared terms, often written using a matrix.
The KL expansion produces a quadratic form using the Fisher matrix.
rescale (verb)
To change the size or scale of something.
Multiplying by the inverse Fisher matrix rescales the gradient.
conjugate gradient (noun)
An iterative numerical method for solving certain optimization problems efficiently.
TRPO uses conjugate gradient to compute its update.
numerical problem (noun)
A difficulty in computation caused by how numbers behave, such as scale differences.
The lecture diagnoses a numerical problem in policy gradients.
shrink (verb)
To become smaller.
The gradient direction changes as sigma shrinks.
well-scaled (adjective)
Having consistent, appropriately sized values across dimensions.
A well-scaled gradient points toward the optimum.
second-order (adjective)
Using information about curvature, not just slope, in optimization.
The Fisher matrix comes from a second-order expansion.
estimated from samples (phrase)
Computed approximately using collected data rather than an exact formula.
The Fisher information matrix is estimated from samples.

Chapters

← Lecture 5, Part 5: Implementing Policy Gradients in Practice · Lecture 6, Part 1: Actor-Critic and Value Functions →