Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 34 of 99 · 10:04
Lecture 8, Part 5: Q-Learning with Continuous Actions
Study guide
What this lecture covers
Q-learning as discussed so far assumes discrete actions, where the max over actions needed for the policy and for target values is just an exhaustive evaluation. This part asks how to perform that same max when actions are continuous, since it appears both in action selection and, more critically, in the inner-loop computation of target values.
The lecture presents three families of solutions and ends with the full pseudocode for DDPG, a continuous-action Q-learning algorithm. This extends the target-network and double Q-learning material from earlier in Lecture 8 to continuous control problems, which recur throughout the rest of the course.
Key ideas
- The core difficulty: an exact argmax over a continuous action space has no closed form, and it must be computed cheaply because it sits inside the training loop.
- Option 1, stochastic optimization: approximate the max by sampling a set of candidate actions (or using CEM/CMA-ES for a more refined search) and picking the best one; simple, easy to parallelize, less accurate in high dimensions.
- Option 2, easy-to-optimize function classes: design the Q-function so the action-dependent part is quadratic, as in NAF (normalized advantage function), giving a closed-form optimum at the cost of representational power.
- Option 3, learned maximizer: train a second network
mu_theta(s)to approximate the argmax ofQ_phi, updated by backpropagating the Q-function's gradient throughmu_thetavia the chain rule. - DDPG: combines option 3 with target networks for both
Q_phi'andmu_theta', alternating a Q-function update with an update tomu_thetathat maximizesQ_phi(s, mu_theta(s)). - Related algorithms: DDPG is closely related to the earlier NFQCA algorithm and can also be viewed as a deterministic actor-critic method; later variants include TD3 and SAC.
Walkthrough
The continuous action problem (1:56)
With discrete actions the max is a simple enumeration; with continuous actions it becomes an optimization problem that must run efficiently inside the training loop, both for the policy's action selection and, especially, for the target-value computation.
Option 1: random sampling and stochastic optimization (4:20)
The simplest approach samples a batch of candidate actions and takes the one with the highest Q-value; it's inexact but fast and trivially parallelized, and its imprecision can even reduce overestimation. More accurate options include the cross-entropy method (CEM), which iteratively refines the sampling distribution toward good regions, and CMA-ES, a more elaborate variant; both work reasonably well up to roughly 40-dimensional action spaces.
Option 2: normalized advantage functions (6:14)
Instead of approximating the max, this approach restricts the Q-function's shape so the max has a closed form: NAF has a network output a bias, a vector, and a positive-definite matrix that define a function quadratic in the action for each state. The maximizing action is then simply the vector output, and the algorithm otherwise runs unchanged, at the cost of being unable to represent Q-functions that are not quadratic in the action.
Option 3: learning an approximate maximizer (DDPG) (9:34)
A second network mu_theta(s) is trained to output the action that approximately maximizes Q_phi(s, a), found by pushing gradients through the Q-function and back through mu_theta using the chain rule. Target values then use mu_theta' in place of an explicit argmax: y = r + gamma * Q_phi'(s', mu_theta'(s')). The full algorithm mirrors standard Q-learning, adding a gradient update for theta and target-network updates (for example via Polyak averaging) for both phi' and theta'. This is essentially the DDPG algorithm, closely related to the earlier NFQCA method, with more recent successors including TD3 and SAC.
Before you watch
- Be familiar with target networks and the target-value computation from earlier in Lecture 8, since all three options here modify how that max is computed.
- Knowing the actor-critic framework helps, since DDPG can also be understood as a deterministic actor-critic method.
Check your understanding
- Why is computing the max over actions harder for continuous action spaces than discrete ones?
- What is the trade-off between random-sampling approximation and CEM/CMA-ES for approximating the max?
- How does the normalized advantage function (NAF) get a closed-form max, and what does it give up to do so?
- In DDPG, what role does the network
mu_theta(s)play, and how is it trained?
Vocabulary
- continuous action (phrase)
- An action described by a real number, rather than picked from a fixed list.
Steering a car is a continuous action. - discrete action (phrase)
- An action chosen from a fixed, countable list of options.
Moving left or right is a discrete action. - exhaustive (adjective)
- Checking every possible option, leaving nothing out.
With discrete actions, the max is an exhaustive evaluation. - closed form (phrase)
- An exact formula that can be computed directly, without searching.
There is no closed form for the max over continuous actions. - stochastic optimization (phrase)
- Finding a good solution by trying random samples.
Stochastic optimization approximates the max with sampled actions. - candidate (noun)
- A possible option being considered.
We sample several candidate actions and keep the best. - cross-entropy method (CEM) (noun)
- A method that repeatedly samples, keeps the best results, and refits the sampling distribution.
CEM gives a more refined search than random sampling. - parallelize (verb)
- To run many computations at the same time.
Random sampling is easy to parallelize. - quadratic (adjective)
- Involving squared terms, forming a curved (parabola-like) shape.
NAF makes the Q-function quadratic in the action. - closed-form optimum (phrase)
- The exact best answer computed directly, without searching.
A quadratic function has a closed-form optimum. - representational power (phrase)
- How wide a range of functions or patterns a model can capture.
NAF gives up representational power for a closed-form solution. - maximizer (noun)
- The input value that produces the largest possible output.
We train a network to approximate the maximizer of the Q-function. - backpropagate (verb)
- To send gradient information backward through a network to update its parameters.
We backpropagate the Q-function's gradient through the policy network. - chain rule (noun)
- A calculus rule for computing the derivative of a function built from other functions.
The chain rule connects the Q-function's gradient to the policy's parameters. - DDPG (noun)
- An algorithm that learns a Q-function and a deterministic policy together for continuous actions.
DDPG combines a learned maximizer with target networks. - target network (phrase)
- A slower-updating copy of a network used to compute stable targets.
DDPG uses target networks for both the Q-function and the policy. - alternate (verb)
- To take turns doing two different things.
DDPG alternates updating the Q-function and the policy. - deterministic actor-critic (phrase)
- An actor-critic method where the policy outputs one fixed action instead of a probability distribution.
DDPG can be viewed as a deterministic actor-critic method. - successor (noun)
- A later method that follows and improves on an earlier one.
TD3 and SAC are successors to DDPG. - inexact (adjective)
- Not perfectly accurate, only approximate.
Random sampling gives an inexact but fast approximation.
Chapters
← Lecture 8, Part 4: Overestimation, Double Q-Learning, N-Step Returns · Lecture 8, Part 6: Practical Tips and Q-Learning Case Studies →
