Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 78 of 99 · 28:38
Lecture 19, Control as Inference, Part 2
Study guide
What this lecture covers
Building on the optimality-variable graphical model from part 1, this lecture derives the three inference operations needed to make the framework useful: backward messages, the resulting near-optimal policy, and forward messages. It shows that the backward recursion is a "soft" version of value iteration, using log-sum-exp (a soft max) in place of the hard max, and that the near-optimal policy turns out to be a Boltzmann distribution over a soft advantage function.
By the end, you should be able to state the backward-message recursion, explain how it reduces to the ordinary Bellman backup in deterministic environments, describe why an action prior can be folded into the reward without loss of generality, and explain how backward and forward messages combine to give the state marginal under near-optimal behavior.
Key ideas
- Backward message recursion:
beta_t(s_t,a_t) = p(O_t|s_t,a_t) * E[beta_{t+1}(s_{t+1})], computed backward from the end of the trajectory, withbeta_Tinitialized from the last reward. - Soft value and Q functions: defining
V_t = log(beta_t(s_t))andQ_t = log(beta_t(s_t,a_t))turns the recursion intoV_t(s_t) = log(integral of exp(Q_t(s_t,a_t))), a soft max, andQ_t(s_t,a_t) = r(s_t,a_t) + log E[exp(V_{t+1}(s_{t+1}))], a soft Bellman backup. - Deterministic transitions recover the Bellman equation: when the next state is a deterministic function of the current state and action, the soft backup collapses exactly to the standard Bellman backup, but with soft max instead of hard max.
- Optimism problem under stochastic transitions: the soft backup's log-sum-exp over next states is dominated by the highest-value outcome, which makes the model act as though it can choose to get lucky rather than merely tolerate luck.
- Action prior: an assumed-uniform prior over actions given a state can, if actually non-uniform, be folded into the reward as
log p(a|s)without loss of generality, so uniformity can be assumed without losing expressiveness. - Recovered policy:
pi(a_t|s_t) = exp(Q_t(s_t,a_t) - V_t(s_t)), a Boltzmann distribution over the soft advantage; a temperature parameter can interpolate between this soft policy and the deterministic greedy policy. - Forward messages and state marginals: forward messages (probability of reaching a state given past optimality) are derived with a similar recursive, Bayes'-rule argument, and combined with backward messages give
p(s_t | O_{1:T}) ∝ beta_t(s_t) * (forward message), matching the "two cones intersecting" intuition and observed patterns in human reaching behavior.
Before you watch
- Watch part 1 of this lecture, which sets up the optimality-variable graphical model this part performs inference on.
- Review value iteration and the Bellman backup, and basic Bayes' rule manipulations, since the derivations rely heavily on both.
Check your understanding
- Why does the backward message recursion reduce exactly to the standard Bellman backup only in the deterministic-transition case?
- What is the "optimism problem" that arises when applying this soft backup to stochastic transitions, and why does it happen?
- Why can a non-uniform action prior always be absorbed into the reward function instead of being modeled separately?
- How do backward and forward messages combine to produce the state marginal
p(s_t | O_{1:T}), and what intuitive picture does the lecture give for this (the two "cones")?
Vocabulary
- recursion (noun)
- A process that repeats by using its own earlier result as input.
The backward message recursion runs from the end of the trajectory. - soft value function (noun)
- A value function computed using a smooth soft-max instead of a hard maximum.
The soft value function replaces the usual max with log-sum-exp. - log-sum-exp (noun)
- A smooth mathematical function that behaves like a soft version of the maximum.
V_t is computed with a log-sum-exp over actions. - soft max (noun)
- A smooth function that turns a set of numbers into weighted probabilities favoring the largest ones.
The recursion becomes a soft max instead of a hard max. - Bellman backup (noun)
- A single update step that improves a value estimate using neighboring values.
The soft Bellman backup adds a log-expectation term. - optimism (noun)
- An overly positive bias, expecting better outcomes than are realistic.
The optimism problem makes the model expect to get lucky. - action prior (noun)
- An assumed distribution over which actions are likely, before any learning.
A uniform action prior can be folded into the reward. - Boltzmann distribution (noun)
- A probability distribution where more likely outcomes have exponentially higher weight based on a score.
The recovered policy is a Boltzmann distribution over the advantage. - advantage function (noun)
- A measure of how much better an action is compared to the average action in that state.
The policy uses the soft advantage function. - temperature (noun)
- A parameter that controls how sharp or spread out a probability distribution is.
A temperature parameter interpolates between soft and greedy policies. - state marginal (noun)
- The probability of being in a particular state, ignoring other variables.
Combining messages gives the state marginal under near-optimal behavior. - intuition (noun)
- A natural feeling for how something works, without needing exact proof.
The two cones intersecting give an intuitive picture of the state marginal. - graphical model (noun)
- A diagram showing how a set of random variables depend on each other.
This lecture performs inference on an optimality-variable graphical model. - optimality variable (noun)
- A binary variable in the model that represents whether a time step is optimal.
Backward messages compute the probability of the optimality variable. - forward message (noun)
- A quantity computed moving forward in time that gives the probability of reaching a state given past optimality.
Forward messages are derived with a recursive Bayes'-rule argument. - backward message (noun)
- A quantity computed moving backward in time from the end of a sequence.
The backward message recursion starts from the last reward. - collapse (verb)
- To simplify down to a much simpler special case.
The soft backup collapses to the standard Bellman backup when transitions are deterministic. - without loss of generality (phrase)
- A phrase meaning an assumption can be made freely because it does not limit the result.
A uniform action prior can be assumed without loss of generality. - interpolate (verb)
- To move smoothly between two extremes, taking values in between.
The temperature parameter lets the policy interpolate between soft and greedy behavior. - deterministic (adjective)
- Always producing the same output given the same input, with no randomness.
In the deterministic case, the next state is a fixed function of the current state and action. - stochastic (adjective)
- Involving randomness, so the outcome is not always the same.
Stochastic transitions cause the optimism problem in the soft backup. - trajectory (noun)
- The full sequence of states and actions over time.
The backward message recursion runs from the end of the trajectory. - fold into (phrasal verb)
- To combine one thing into another so it becomes part of it.
The action prior can be folded into the reward function. - recover (verb)
- To arrive back at a known, familiar result from a more general one.
The soft backup recovers the ordinary Bellman backup as a special case. - greedy (adjective)
- Always picking the single best-looking option, with no randomness.
A temperature of zero gives the deterministic greedy policy. - dominate (verb)
- To have by far the largest effect, outweighing everything else.
The log-sum-exp is dominated by the highest-value outcome. - tolerate (verb)
- To accept something even though it is not ideal.
A good policy should tolerate bad luck rather than count on good luck.
Chapters
- 0:00 Inference = planning
- 0:52 Backward messages
- 6:51 A closer look at the backward pass
- 11:56 Backward pass summary
- 13:01 The action prior remember this?
- 17:58 Policy computation with value functions
- 19:09 Policy computation summary
- 20:05 Forward messages
- 25:58 Forward/backward message intersection
← Lecture 19, Control as Inference, Part 1 · Lecture 19, Control as Inference, Part 3 →
