Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 80 of 99 · 9:54
Lecture 19, Control as Inference, Part 4
Study guide
What this lecture covers
This closing part of the control-as-inference lecture instantiates the soft-optimality framework as two concrete algorithms: soft Q-learning, which replaces the hard max in the Q-learning target with a soft max, and entropy-regularized policy gradient, which adds an entropy bonus to the standard policy gradient objective. It also shows that the two are closely related once the policy is written as the exponentiated soft advantage.
Key ideas
- Soft Q-learning: identical to standard Q-learning except the target value uses a soft max (log-sum-exp of next Q-values) instead of a hard max over actions; the resulting policy is the exponentiated advantage rather than the greedy action.
- Entropy-regularized policy gradient: adds the policy's entropy to the expected-reward objective; its gradient equals the standard policy gradient with
r - log piused as the reward, and the extra term from differentiatinglog pireduces to a constant that acts as a baseline and can be ignored. - Connection to soft Q-learning: substituting
log pi = Q - Vinto the entropy-regularized policy gradient produces an expression close to the soft Q-learning update, differing mainly by an off-policy correction term. - Practical benefits: entropy regularization prevents premature collapse of policy stochasticity (which would otherwise harm exploration), gives a principled way to break ties between equally good actions, improves robustness by covering multiple ways to solve a task, and makes policies easier to fine-tune for related tasks.
- Recovering hard optimality: as reward magnitude increases, or as a temperature parameter is driven toward zero, the soft-optimal policy approaches the standard deterministic optimal policy.
Before you watch
- Watch parts 1 through 3 of this lecture, which derive the soft value iteration and the entropy-augmented RL objective that these algorithms optimize.
- Review standard Q-learning and policy gradient from earlier lectures in this course, since both algorithms here are direct modifications of them.
Check your understanding
- What is the only change needed to turn standard Q-learning into soft Q-learning?
- Why does adding an entropy term to the policy gradient objective end up looking like subtracting
log pifrom the reward, with no other change to the gradient expression? - What practical problem does entropy regularization solve for on-policy policy gradient methods?
- Under what conditions does the soft-optimal policy reduce back to the standard, deterministic-optimal policy?
Vocabulary
- soft Q-learning (noun)
- A version of Q-learning that uses a soft max instead of a hard max in its target.
Soft Q-learning replaces the max operator with log-sum-exp. - entropy-regularized (adjective)
- Adjusted by adding a bonus for keeping behavior varied and less predictable.
Entropy-regularized policy gradient adds an entropy term to the reward. - baseline (noun)
- A reference value subtracted from a reward estimate to reduce variance without changing the expected result.
The constant term acts as a baseline and can be ignored. - off-policy correction (noun)
- An adjustment made when learning from data collected by a different policy.
The main difference is an off-policy correction term. - premature (adjective)
- Happening too early, before it should.
Entropy regularization prevents premature loss of randomness. - break ties (phrase)
- To choose fairly between two or more equally good options.
Entropy gives a principled way to break ties between actions. - robustness (noun)
- The ability to keep working well despite changes or disturbances.
Entropy improves robustness by covering multiple solutions. - converge to (phrase)
- To gradually approach and become the same as another value.
As temperature drops, the policy converges to the greedy one. - entropy (noun)
- A measure of how spread out or unpredictable a probability distribution is.
Entropy regularization rewards the policy for staying varied. - policy gradient (noun)
- A method that updates a policy's parameters to increase expected reward.
The entropy bonus is added to the standard policy gradient objective. - Q-learning (noun)
- A method that learns the value of taking each action in each state.
Soft Q-learning modifies the target used in standard Q-learning. - target value (noun)
- The value a learning update tries to move its estimate closer to.
Soft Q-learning changes how the target value is computed. - greedy action (noun)
- The single action currently judged to be best.
Standard Q-learning picks the greedy action rather than a soft distribution. - exponentiate (verb)
- To raise a number to a power, often e to that power.
The soft-optimal policy is the exponentiated advantage. - bonus (noun)
- An extra amount added to reward good behavior.
An entropy bonus is added to the reward objective. - stochasticity (noun)
- The quality of involving randomness.
Entropy regularization prevents premature loss of policy stochasticity. - magnitude (noun)
- The size of a quantity, regardless of its direction.
As reward magnitude increases, the policy becomes more deterministic. - temperature (noun)
- A parameter that controls how spread out or sharp a probability distribution is.
Driving the temperature toward zero recovers the deterministic policy. - principled (adjective)
- Based on a clear, justified reason rather than an arbitrary choice.
Entropy gives a principled way to break ties between actions. - instantiate (verb)
- To turn a general idea into a specific, concrete version.
This lecture instantiates soft optimality as two concrete algorithms. - augmented (adjective)
- Made larger or more complete by adding something extra.
The entropy-augmented objective adds a bonus term to the reward. - substitute (verb)
- To replace one expression with another in an equation.
Substituting log pi = Q - V links the two algorithms. - on-policy (adjective)
- Describes learning that uses data collected by the current policy itself.
Entropy regularization helps on-policy policy gradient methods explore. - collapse (noun)
- A sudden loss of variety, narrowing down to one option.
Entropy regularization prevents premature collapse of exploration. - concrete (adjective)
- Specific and real, rather than abstract.
The framework becomes two concrete algorithms in this lecture.
Chapters
- 0:00 Q-learning with soft optimality
- 2:07 Policy gradient with soft optimality
- 3:49 Policy gradient vs Q-learning
- 7:28 Benefits of soft optimality
- 9:12 Review
← Lecture 19, Control as Inference, Part 3 · Lecture 19: Control as Inference, Part 5 →
