Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 19 of 99 · 7:32

Lecture 5, Part 5: Implementing Policy Gradients in Practice

CS 285: Lecture 5, Part 5 on YouTube

Study guide

What this lecture covers

This part turns the policy gradient math into something you can actually code with automatic differentiation tools. It answers: how do you get PyTorch or TensorFlow to compute the policy gradient efficiently, without manually calculating a gradient vector for every sampled state-action pair?

After watching, you can write a "pseudo-loss" that tricks an autodiff package into producing the correct policy gradient, and you know practical settings — batch size, learning rate, optimizer — to expect when training with policy gradients.

Key ideas

  • Why naive computation is expensive: computing grad log pi(a|s) separately for every sampled state-action pair produces a vector as long as the number of network parameters (often millions) for each of potentially thousands of samples, which is costly in memory and compute.
  • Pseudo-loss trick: instead of the real objective, implement J-tilde, the sum of log-probabilities of sampled actions weighted by their reward-to-go (Q-hat); this quantity isn't the RL objective itself, but its gradient equals the policy gradient.
  • Reuse of maximum-likelihood code: log pi is the same cross-entropy loss (discrete actions) or squared error (Gaussian continuous actions) used in supervised learning; policy gradient implementation just multiplies per-sample likelihoods by precomputed reward-to-go values before taking the mean and calling backward.
  • Autodiff doesn't need to know Q-hat depends on theta: the reward-to-go values are treated as fixed weights, not differentiated through, which is what makes the trick work.
  • High variance changes practice: despite looking like supervised learning, policy gradient training needs much larger batch sizes (thousands to tens of thousands of samples) because gradients are noisy.
  • Optimizer and tuning: plain SGD with momentum is hard to use; Adam is a reasonable default, and expect more hyperparameter tuning than typical supervised learning, with dedicated step-size methods like natural gradient covered later.

Before you watch

  • Watch the earlier parts of this lecture on deriving the policy gradient, causality, baselines, and off-policy importance sampling.
  • Be familiar with how automatic differentiation and backpropagation work in a deep learning framework.

Check your understanding

  1. Why is computing grad log pi separately for every sampled state-action pair computationally expensive?
  2. What is the pseudo-loss J-tilde, and why does its gradient equal the true policy gradient even though it isn't the RL objective itself?
  3. How does implementing a policy gradient differ from implementing a standard maximum-likelihood loss?
  4. Why do policy gradient methods typically need much larger batch sizes than supervised learning?

Vocabulary

automatic differentiation (noun)
A technique that lets a computer calculate derivatives automatically instead of by hand.
Automatic differentiation computes the policy gradient efficiently.
pseudo-loss (noun)
A fake loss function whose gradient happens to match a different desired quantity.
The pseudo-loss trick makes autodiff produce the real policy gradient.
trick (noun)
A clever technique used to solve a problem indirectly.
The pseudo-loss is a trick to reuse existing autodiff tools.
state-action pair (noun)
A combination of a specific state and the action taken in it.
A gradient is needed for every sampled state-action pair.
cross-entropy loss (noun)
A loss used for classification that measures how well predicted probabilities match true labels.
Discrete action policies reuse cross-entropy loss code.
squared error (noun)
A loss that measures the squared difference between a prediction and the true value.
Continuous Gaussian policies use squared error as their loss.
precomputed (adjective)
Calculated in advance, before being used in a later step.
Reward-to-go values are precomputed before training.
fixed weight (noun)
A value treated as constant, not adjusted during optimization.
Reward-to-go acts as a fixed weight in the pseudo-loss.
batch size (noun)
The number of examples processed together in one training step.
Policy gradient methods need a much larger batch size.
hyperparameter (noun)
A setting chosen before training that controls how a model learns.
Tuning hyperparameters is harder for policy gradient than supervised learning.
step-size method (noun)
A technique for choosing how large each optimization update should be.
Natural gradient is a dedicated step-size method covered later.
noisy (adjective)
Containing random variation that obscures the true signal.
Policy gradients are noisy compared to supervised gradients.
framework (noun)
A set of pre-built tools used to build and train models.
PyTorch is a common deep learning framework.
efficiently (adverb)
In a way that uses time and resources well.
The pseudo-loss computes the gradient efficiently.
manually (adverb)
Done by hand, without an automatic tool.
The gradient could be computed manually, but that's expensive.
vector (noun)
An ordered list of numbers, often representing a direction or set of values.
The gradient is a vector as long as the number of parameters.
costly (adjective)
Requiring a large amount of time, memory, or computation.
Naive gradient computation is costly in memory and compute.
reasonable default (phrase)
A sensible choice to use when no better option is known.
Adam is a reasonable default optimizer for policy gradient.
momentum (noun)
An optimization technique that keeps a running average of past gradients.
Plain SGD with momentum is hard to use for policy gradient.
call backward (phrase)
To trigger the automatic computation of gradients in a deep learning framework.
We call backward after computing the pseudo-loss.

Chapters

← Lecture 5, Part 4: Off-Policy Policy Gradients with Importance Sampling · Lecture 5, Part 6: The Natural Policy Gradient →