Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 20 of 99 · 13:32
Lecture 5, Part 6: The Natural Policy Gradient
Study guide
What this lecture covers
This closing part of the lecture diagnoses a numerical problem that hurts policy gradients, especially with continuous actions, and derives a fix called the natural (or covariant) policy gradient. It uses a small 1D example — a Gaussian policy with a mean parameter k and a variance parameter sigma — to show why plain gradient ascent struggles, then reframes gradient ascent as a constrained optimization problem to motivate the fix.
After watching, you can explain why some policy parameters need much smaller learning rates than others, describe gradient ascent as maximizing a linearized objective within a parameter-space trust region, and explain how replacing that region with a KL-divergence constraint on the policy distribution produces the natural gradient.
Key ideas
- Poor conditioning example: in a small Gaussian-policy problem, the gradient with respect to
sigmagrows much larger than the gradient with respect tokassigmashrinks, so the raw gradient direction doesn't point toward the optimum — the same issue as optimizing a quadratic with a badly scaled Hessian. - Gradient ascent as constrained optimization: a first-order gradient step can be viewed as maximizing a linear (Taylor) approximation of the objective subject to a small step-size constraint
||theta' - theta||^2 <= epsilonin parameter space. - Wrong space for the constraint: constraining step size in raw parameter space is awkward, because some parameters change the policy distribution a lot per unit step and others change it very little.
- KL-divergence constraint: replacing the parameter-space constraint with a constraint on
KL(pi_theta' || pi_theta)(a parameterization-independent divergence between distributions) makes step sizes consistent in policy space rather than parameter space. - Fisher information matrix: a second-order Taylor expansion of the KL divergence gives a quadratic form using the Fisher information matrix
F = E[grad log pi * grad log pi^T], estimated from samples. - Natural gradient update: solving the constrained problem gives the update
theta <- theta + alpha * F^-1 * grad J(theta); multiplying byF^-1rescales the gradient so it points toward the optimum, converging faster and requiring less step-size tuning. Natural policy gradient and TRPO are named algorithms that use this idea, with TRPO deriving alpha from a chosen epsilon via conjugate gradient.
Before you watch
- Watch the earlier parts of this lecture, especially the policy gradient derivation and implementation.
- Basic familiarity with Taylor expansions, constrained optimization (Lagrange multipliers), and KL divergence is helpful.
Check your understanding
- Why does the raw policy gradient point in a poor direction when one parameter (like
sigma) has a much larger gradient magnitude than another? - How does viewing gradient ascent as a constrained optimization problem motivate replacing the parameter-space constraint with a KL-divergence constraint?
- What role does the Fisher information matrix play in the natural gradient update?
- What is the practical benefit of using
F^-1 * grad J(theta)instead of the raw gradient?
Vocabulary
- diagnose (verb)
- To identify the cause of a problem.
This lecture diagnoses a numerical problem in policy gradients. - natural gradient (noun)
- A modified gradient update that accounts for how parameters affect the policy's distribution.
The natural gradient rescales the update using distribution geometry. - covariant (adjective)
- Changing consistently together with another quantity, regardless of how it's measured.
The natural gradient is also called the covariant gradient. - poor conditioning (noun)
- A situation where a function is much steeper in some directions than others, making optimization hard.
Poor conditioning makes gradient ascent struggle. - Hessian (noun)
- A matrix of second derivatives describing a function's curvature.
A badly scaled Hessian causes similar problems in optimization. - linearized (adjective)
- Approximated using a straight-line version of a curved function.
Gradient ascent maximizes a linearized version of the objective. - Taylor approximation (noun)
- An estimate of a function using its value and derivatives at one point.
A first-order Taylor approximation underlies plain gradient ascent. - trust region (noun)
- A limited area around current parameters where an update is trusted to be safe.
The step-size constraint defines a trust region. - constraint (noun)
- A limit or condition that a solution must satisfy.
The optimization is subject to a step-size constraint. - KL divergence (noun)
- A measure of how different one probability distribution is from another.
The natural gradient uses a KL divergence constraint. - parameterization-independent (adjective)
- Not affected by how a model's parameters happen to be defined.
KL divergence is parameterization-independent, unlike raw parameter distance. - Fisher information matrix (noun)
- A matrix describing how sensitive a distribution is to changes in its parameters.
The Fisher information matrix comes from a second-order KL expansion. - quadratic form (noun)
- An expression involving squared terms, often written using a matrix.
The KL expansion produces a quadratic form using the Fisher matrix. - rescale (verb)
- To change the size or scale of something.
Multiplying by the inverse Fisher matrix rescales the gradient. - conjugate gradient (noun)
- An iterative numerical method for solving certain optimization problems efficiently.
TRPO uses conjugate gradient to compute its update. - numerical problem (noun)
- A difficulty in computation caused by how numbers behave, such as scale differences.
The lecture diagnoses a numerical problem in policy gradients. - shrink (verb)
- To become smaller.
The gradient direction changes as sigma shrinks. - well-scaled (adjective)
- Having consistent, appropriately sized values across dimensions.
A well-scaled gradient points toward the optimum. - second-order (adjective)
- Using information about curvature, not just slope, in optimization.
The Fisher matrix comes from a second-order expansion. - estimated from samples (phrase)
- Computed approximately using collected data rather than an exact formula.
The Fisher information matrix is estimated from samples.
Chapters
- 0:00 <Untitled Chapter 1>
- 0:25 What else is wrong with the policy gradient?
- 4:10 Covariant/natural policy gradient
- 11:31 Advanced policy gradient topics
- 11:54 Example: policy gradient with importance sampling
- 12:25 Example: trust region policy optimization
- 12:42 Policy gradients suggested readings
← Lecture 5, Part 5: Implementing Policy Gradients in Practice · Lecture 6, Part 1: Actor-Critic and Value Functions →
