Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 19 of 99 · 7:32
Lecture 5, Part 5: Implementing Policy Gradients in Practice
Study guide
What this lecture covers
This part turns the policy gradient math into something you can actually code with automatic differentiation tools. It answers: how do you get PyTorch or TensorFlow to compute the policy gradient efficiently, without manually calculating a gradient vector for every sampled state-action pair?
After watching, you can write a "pseudo-loss" that tricks an autodiff package into producing the correct policy gradient, and you know practical settings — batch size, learning rate, optimizer — to expect when training with policy gradients.
Key ideas
- Why naive computation is expensive: computing
grad log pi(a|s)separately for every sampled state-action pair produces a vector as long as the number of network parameters (often millions) for each of potentially thousands of samples, which is costly in memory and compute. - Pseudo-loss trick: instead of the real objective, implement
J-tilde, the sum of log-probabilities of sampled actions weighted by their reward-to-go (Q-hat); this quantity isn't the RL objective itself, but its gradient equals the policy gradient. - Reuse of maximum-likelihood code:
log piis the same cross-entropy loss (discrete actions) or squared error (Gaussian continuous actions) used in supervised learning; policy gradient implementation just multiplies per-sample likelihoods by precomputed reward-to-go values before taking the mean and calling backward. - Autodiff doesn't need to know Q-hat depends on theta: the reward-to-go values are treated as fixed weights, not differentiated through, which is what makes the trick work.
- High variance changes practice: despite looking like supervised learning, policy gradient training needs much larger batch sizes (thousands to tens of thousands of samples) because gradients are noisy.
- Optimizer and tuning: plain SGD with momentum is hard to use; Adam is a reasonable default, and expect more hyperparameter tuning than typical supervised learning, with dedicated step-size methods like natural gradient covered later.
Before you watch
- Watch the earlier parts of this lecture on deriving the policy gradient, causality, baselines, and off-policy importance sampling.
- Be familiar with how automatic differentiation and backpropagation work in a deep learning framework.
Check your understanding
- Why is computing
grad log piseparately for every sampled state-action pair computationally expensive? - What is the pseudo-loss
J-tilde, and why does its gradient equal the true policy gradient even though it isn't the RL objective itself? - How does implementing a policy gradient differ from implementing a standard maximum-likelihood loss?
- Why do policy gradient methods typically need much larger batch sizes than supervised learning?
Vocabulary
- automatic differentiation (noun)
- A technique that lets a computer calculate derivatives automatically instead of by hand.
Automatic differentiation computes the policy gradient efficiently. - pseudo-loss (noun)
- A fake loss function whose gradient happens to match a different desired quantity.
The pseudo-loss trick makes autodiff produce the real policy gradient. - trick (noun)
- A clever technique used to solve a problem indirectly.
The pseudo-loss is a trick to reuse existing autodiff tools. - state-action pair (noun)
- A combination of a specific state and the action taken in it.
A gradient is needed for every sampled state-action pair. - cross-entropy loss (noun)
- A loss used for classification that measures how well predicted probabilities match true labels.
Discrete action policies reuse cross-entropy loss code. - squared error (noun)
- A loss that measures the squared difference between a prediction and the true value.
Continuous Gaussian policies use squared error as their loss. - precomputed (adjective)
- Calculated in advance, before being used in a later step.
Reward-to-go values are precomputed before training. - fixed weight (noun)
- A value treated as constant, not adjusted during optimization.
Reward-to-go acts as a fixed weight in the pseudo-loss. - batch size (noun)
- The number of examples processed together in one training step.
Policy gradient methods need a much larger batch size. - hyperparameter (noun)
- A setting chosen before training that controls how a model learns.
Tuning hyperparameters is harder for policy gradient than supervised learning. - step-size method (noun)
- A technique for choosing how large each optimization update should be.
Natural gradient is a dedicated step-size method covered later. - noisy (adjective)
- Containing random variation that obscures the true signal.
Policy gradients are noisy compared to supervised gradients. - framework (noun)
- A set of pre-built tools used to build and train models.
PyTorch is a common deep learning framework. - efficiently (adverb)
- In a way that uses time and resources well.
The pseudo-loss computes the gradient efficiently. - manually (adverb)
- Done by hand, without an automatic tool.
The gradient could be computed manually, but that's expensive. - vector (noun)
- An ordered list of numbers, often representing a direction or set of values.
The gradient is a vector as long as the number of parameters. - costly (adjective)
- Requiring a large amount of time, memory, or computation.
Naive gradient computation is costly in memory and compute. - reasonable default (phrase)
- A sensible choice to use when no better option is known.
Adam is a reasonable default optimizer for policy gradient. - momentum (noun)
- An optimization technique that keeps a running average of past gradients.
Plain SGD with momentum is hard to use for policy gradient. - call backward (phrase)
- To trigger the automatic computation of gradients in a deep learning framework.
We call backward after computing the pseudo-loss.
Chapters
← Lecture 5, Part 4: Off-Policy Policy Gradients with Importance Sampling · Lecture 5, Part 6: The Natural Policy Gradient →
