Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 81 of 99 · 10:58

Lecture 19: Control as Inference, Part 5

CS 285: Lecture 19, Control as Inference, Part 5 on YouTube

Study guide

What this lecture covers

This closing part of the control-as-inference lecture connects the variational inference framework built earlier in the course to real algorithms and results. It asks why running the same deep RL algorithm twice can produce very different, sometimes worse, behavior, and shows how treating the policy as a distribution proportional to exponentiated Q-values fixes this by keeping multiple good strategies alive instead of collapsing onto one.

By the end, you can explain how soft Q-learning and soft actor-critic follow from the soft optimality model, why entropy-seeking pretraining speeds up fine-tuning, and where to find the original papers behind these methods.

Key ideas

  • Local optimum problem in RL: two runs of the same algorithm can commit to different, unequally good solutions because early random advantages get amplified during learning.
  • Soft Q-learning: choosing actions proportional to the exponentiated Q-value keeps probability mass on multiple promising strategies instead of committing early to one.
  • Entropy-seeking pretraining: training with a broad reward (for example, run fast in any direction) under soft Q-learning produces policies that later specialize faster than deterministic algorithms like DDPG, which must unlearn a wrong direction first.
  • Soft actor-critic (SAC): the actor-critic counterpart of soft Q-learning; its Q-update looks like ordinary actor-critic but subtracts a -log pi entropy term, and its policy update fits the variational distribution to the Q-function.
  • Message passing view: updating the Q-function corresponds to inference in the graphical model, while updating the policy corresponds to fitting the variational approximation to the posterior.
  • Robustness from entropy: because SAC learns many ways to solve a task, resulting policies (a robotic arm stacking Lego, a legged robot walking) tolerate physical perturbations better than policies trained without entropy.

Before you watch

  • Be familiar with the variational inference and soft optimality framework from the earlier parts of this control-as-inference lecture.
  • Know the basics of Q-learning and actor-critic methods, since both are extended here with an entropy term.

Check your understanding

  1. Why can two runs of the same deep RL algorithm converge to noticeably different behaviors?
  2. How does choosing actions proportional to the exponentiated Q-value change what the agent explores?
  3. Why does a policy pretrained with soft Q-learning fine-tune faster than one pretrained with DDPG?
  4. What term is added to the standard actor-critic Q-update in soft actor-critic, and what does it account for?

Vocabulary

local optimum (noun)
A solution that is better than nearby alternatives, but not necessarily the best overall.
Two training runs can settle in different local optima.
commit (verb)
To fully choose and stick to one option, giving up others.
A policy can commit early to one strategy.
amplify (verb)
To make something stronger or larger.
Small early advantages get amplified during learning.
soft actor-critic (noun)
An actor-critic algorithm that includes an entropy term to encourage exploration.
Soft actor-critic is the actor-critic version of soft Q-learning.
actor-critic (noun)
A method combining a policy (actor) and a value estimator (critic) trained together.
Soft actor-critic extends ordinary actor-critic with entropy.
message passing (noun)
Passing information between parts of a model to compute a result step by step.
Updating the Q-function is like message passing in the graphical model.
perturbation (noun)
A small disturbance or unexpected change to a system.
SAC policies tolerate physical perturbations better.
specialize (verb)
To become focused on one particular task rather than many.
Pretrained policies specialize faster with soft Q-learning.
exploration (noun)
The act of trying new actions to discover useful information.
Exploiting exponentiated Q-values changes how the agent explores.
exploit (verb)
To make full use of known good options.
A greedy policy exploits the current best action too early.
deterministic (adjective)
Always producing the same output for the same input, without randomness.
DDPG is a deterministic algorithm that must unlearn a wrong direction.
promising (adjective)
Showing signs of being good or successful.
Soft Q-learning keeps multiple promising strategies alive.
unlearn (verb)
To lose or discard a previously learned behavior.
DDPG must unlearn a wrong direction before it can specialize.
robotic arm (noun)
A mechanical arm controlled by a computer to move and manipulate objects.
A robotic arm learns to stack Lego bricks using SAC.
legged robot (noun)
A robot that moves using leg-like limbs instead of wheels.
A legged robot walks more robustly after training with entropy.
variational distribution (noun)
A simple distribution used to approximate a more complex one during training.
SAC's policy update fits the variational distribution to the Q-function.
posterior (noun)
An updated probability distribution over a hidden variable after seeing evidence.
Updating the policy fits the approximation to the posterior.
subtract (verb)
To take one quantity away from another.
SAC's Q-update subtracts a log pi entropy term.
fine-tune (verb)
To further train an already-trained model on a specific task.
Entropy-seeking pretraining speeds up later fine-tuning.
converge (verb)
To settle toward a particular, often fixed, result.
Two runs of the same algorithm can converge to different behaviors.
framework (noun)
A general structure or set of ideas used to build specific methods.
Both algorithms follow from the soft optimality framework.
tolerate (verb)
To keep working acceptably despite a disturbance.
SAC policies tolerate physical perturbations better.
viable (adjective)
Capable of working successfully in practice.
Soft Q-learning is a viable alternative to deterministic methods.
diverse (adjective)
Made up of many different kinds.
Entropy keeps a diverse set of strategies alive during learning.
counterpart (noun)
A thing that matches or corresponds to another in a different setting.
SAC is the actor-critic counterpart of soft Q-learning.

Chapters

← Lecture 19, Control as Inference, Part 4 · Lecture 20: Inverse Reinforcement Learning, Part 1 →