Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 38 of 99 · 5:50
Lecture 9, Part 3: Constraining Policy Gradient with KL Divergence
Study guide
What this lecture covers
Having shown that policy gradient's tractable objective is only guaranteed to improve the true RL objective when the new and old policies stay close in total variation divergence, this short part asks how to enforce that closeness in an actual algorithm. It replaces the total variation bound with the more convenient KL divergence, and shows how to enforce a KL constraint using a Lagrangian and dual gradient descent.
This directly continues from the state-distribution bound proven in Part 2 of Lecture 9, and previews the family of constrained and proximal policy gradient methods used later in the course.
Key ideas
- Why KL instead of total variation: total variation divergence is not differentiable everywhere, while KL divergence is differentiable when the two distributions share support and has tractable closed forms for many distribution families, and it upper-bounds total variation divergence, so a KL bound is a valid substitute.
- Constrained objective: maximize the importance-weighted advantage objective subject to
KL(pi_theta' || pi_theta) <= epsilon, which still improves the true RL objective for small enoughepsilon. - Lagrangian formulation: form the Lagrangian by subtracting
epsilonfrom the KL term, scale it by a multiplierlambda, and alternate maximizing over the policy parameters and taking a gradient step onlambda. - Dual gradient descent: raising
lambdawhen the constraint is violated and lowering it otherwise asymptotically finds a multiplier that enforces the constraint, a general technique the course revisits later. - Heuristic alternative: instead of solving the dual problem exactly,
lambdacan be chosen manually and the KL term used as a fixed regularizer, penalizing deviation from the old policy, with the inner maximization run incompletely for a few gradient steps. - Connection to real algorithms: this general recipe underlies methods such as guided policy search and PPO.
Before you watch
- Watch Part 2 of Lecture 9, which derives the total-variation bound that this section converts into a KL-divergence constraint.
- A basic familiarity with Lagrangian duality is helpful, though the lecture reviews the mechanics of dual gradient descent.
Check your understanding
- Why is KL divergence more convenient than total variation divergence for enforcing a policy closeness constraint?
- How does dual gradient descent adjust the Lagrange multiplier based on whether the constraint is violated?
- What is the difference between the "principled" Lagrangian approach and the heuristic fixed-penalty approach to enforcing the KL constraint?
- Which later algorithms does the lecture connect this constrained policy gradient recipe to?
Vocabulary
- enforce (verb)
- To make sure a rule or limit is actually followed.
We need to enforce closeness between old and new policies. - KL divergence (noun)
- A measure of how different one probability distribution is from another.
KL divergence replaces total variation as the closeness measure. - differentiable (adjective)
- Able to have a smooth slope calculated at every point.
KL divergence is differentiable, unlike total variation. - closed form (phrase)
- An exact formula that can be computed directly.
KL divergence has tractable closed forms for many distributions. - upper-bound (verb)
- To set a guaranteed maximum limit on a value.
KL divergence upper-bounds total variation divergence. - constrained optimization (phrase)
- Finding the best solution while obeying certain limits or rules.
The problem becomes a constrained optimization with a KL limit. - Lagrangian (noun)
- A combined objective that adds a penalty term to enforce a constraint.
We form the Lagrangian by subtracting epsilon from the KL term. - multiplier (noun)
- A number that scales how strongly a constraint is enforced.
Lambda is the multiplier controlling the KL penalty. - dual gradient descent (phrase)
- A method that alternates optimizing the main objective and adjusting the constraint's multiplier.
Dual gradient descent finds a multiplier that satisfies the constraint. - violate (verb)
- To break or fail to respect a rule.
Raise lambda when the constraint is violated. - asymptotically (adverb)
- Getting closer and closer to a result as a process continues, without ever quite reaching it abruptly.
Dual gradient descent asymptotically enforces the constraint. - heuristic (noun)
- A practical shortcut that usually works well, without a full guarantee.
A heuristic alternative fixes lambda manually. - regularizer (noun)
- An extra term added to an objective to discourage unwanted behavior.
The KL term is used as a fixed regularizer. - deviation (noun)
- How far something moves away from a reference point.
The regularizer penalizes deviation from the old policy. - incompletely (adverb)
- Not fully or entirely finished.
The inner maximization is run incompletely for a few steps. - guided policy search (phrase)
- An algorithm that guides a policy using trajectory optimization and constraints.
This recipe underlies guided policy search. - PPO (noun)
- Proximal Policy Optimization, an algorithm that limits how much a policy can change per update.
PPO uses a related closeness idea to this recipe. - convenient (adjective)
- Easy to use or work with in practice.
KL divergence is more convenient than total variation. - support (distribution) (noun)
- The set of outcomes a probability distribution assigns nonzero chance to.
KL divergence needs the two distributions to share support. - recipe (noun)
- A set sequence of steps used to build an algorithm.
This general recipe underlies several practical algorithms.
Chapters
- 0:00 Constraint derivation
- 0:25 KL divergence advantages
- 2:10 Formulating the problem
- 2:53 Dual gradient descent
← Lecture 9, Part 2: Bounding the State Distribution Mismatch · Lecture 9, Part 4: Natural Gradient and Trust Region Methods →
