Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 79 of 99 · 21:26
Lecture 19, Control as Inference, Part 3
Study guide
What this lecture covers
This part addresses the optimism problem identified with exact inference in the control-as-inference graphical model: because the backward pass conditions on high reward, it can bias transition probabilities toward lucky outcomes, like inferring that a lottery player is likely to win just because they hold a winning ticket. The lecture fixes this using the variational inference tools from the previous week, constraining the approximate posterior to keep the true dynamics fixed while only learning the action distribution.
By the end, you should be able to explain why exact inference distorts the dynamics, describe how restricting the variational distribution to match the true transitions and initial state removes this bias, and derive the resulting objective, which turns out to be the standard expected-reward RL objective plus an action-entropy bonus, optimized with a soft version of value iteration.
Key ideas
- Optimism problem: exact inference under the optimality-variable model computes
p(a_t | s_t, O_{1:T}), which effectively asks "given you got lucky, what did you do", not "what would you have done to be optimal", biasing the inferred behavior toward risky, high-variance outcomes. - Restricted variational distribution: define
q(s_{1:T}, a_{1:T})to use the true initial state distribution and true transition probabilities, but let onlyq(a_t|s_t)be learned, matching the graphical model with the optimality nodes removed. - Resulting lower bound: because the initial-state and transition terms cancel between
pandq, the evidence lower bound simplifies to the expected sum of rewards plus the entropy of the action distribution at each timestep. - Interpretation: maximizing this bound is exactly the standard RL objective (maximize expected reward) plus an entropy bonus that rewards stochastic, near-optimal behavior, explaining the "suboptimal monkey" behavior from part 1.
- Exponential-of-quantity solution rule: whenever an objective has the form "expected value of something minus the log probability of the distribution it's taken under," the optimal distribution is proportional to the exponential of that quantity.
- Soft value iteration backward pass: from the last timestep backward, set
Q_t = r_t + E[V_{t+1}](the ordinary, non-optimistic Bellman backup) andV_t = log integral of exp(Q_t)(a soft max), with the resulting policyq(a_t|s_t) = exp(Q_t - V_t). - Variants: a discount factor
gammacan be added (equivalent to a probability of "death" each step), and a temperaturealphacan interpolate between soft and hard max; the same recursion also has a valid infinite-horizon version.
Before you watch
- Watch parts 1 and 2 of this lecture, which introduce the optimality-variable graphical model and its exact-inference backward pass, including the optimism problem this part resolves.
- Review the variational inference lectures (evidence lower bound and the general recipe for approximating a posterior with a restricted distribution class), since this part directly reuses that machinery.
Check your understanding
- Why does exact inference in the optimality-variable model bias the inferred policy toward risky actions, using the lottery-ticket analogy?
- What specific restriction on the variational distribution
qprevents the optimism bias, and why does that restriction make the lower bound simplify so cleanly? - How does the resulting variational lower bound relate to the standard reinforcement learning objective, and what extra term does it add?
- Why does the rule "optimal distribution is the exponential of the expected quantity" apply here, and what does it say the optimal
q(a_T|s_T)looks like at the final timestep?
Vocabulary
- lottery-ticket analogy (phrase)
- A comparison used to explain a biased conclusion drawn only from a lucky outcome.
The lottery-ticket analogy shows why exact inference is too optimistic. - bias (noun)
- A systematic error that pushes results in one direction.
Exact inference introduces a bias toward risky outcomes. - restrict (verb)
- To limit something to a smaller allowed range or form.
The variational distribution is restricted to match true dynamics. - initial state distribution (noun)
- The probability distribution over which state an episode starts in.
The variational distribution uses the true initial state distribution. - cancel (verb)
- To remove or eliminate an effect by balancing it against an opposite one.
The transition terms cancel between p and q. - entropy bonus (noun)
- An extra reward added for keeping a policy's actions varied and unpredictable.
The objective adds an entropy bonus to the expected reward. - objective (noun)
- The quantity a training process tries to maximize or minimize.
The standard RL objective maximizes expected reward. - infinite-horizon (adjective)
- Continuing forever, without a fixed final timestep.
There is a valid infinite-horizon version of the recursion. - discount factor (noun)
- A number that reduces how much future rewards count compared to immediate ones.
A discount factor gamma can be added to the objective. - variational inference (noun)
- A method that approximates a hard-to-compute probability distribution with a simpler one.
The optimism problem is fixed using variational inference tools. - approximate posterior (noun)
- A simplified distribution used to stand in for a true but hard-to-compute one.
The approximate posterior q is restricted to match the true dynamics. - distort (verb)
- To change something so it no longer accurately represents the truth.
Exact inference distorts the transition probabilities toward lucky outcomes. - high-variance (adjective)
- Producing results that vary a lot between different attempts.
Exact inference biases behavior toward risky, high-variance outcomes. - recipe (noun)
- A general method or set of steps for achieving a result.
The lecture reuses the standard variational inference recipe. - simplify (verb)
- To make something easier to understand or compute.
Canceling terms lets the lower bound simplify greatly. - interpolate (verb)
- To move smoothly between two extremes, taking values in between.
A temperature parameter can interpolate between soft and hard max. - recursion (noun)
- A process that repeats by using its own earlier result as input.
The soft value iteration recursion runs backward from the last timestep. - backward pass (noun)
- Computing values by moving from the end of a sequence toward the start.
The backward pass computes Q_t and V_t at each timestep. - temperature (noun)
- A parameter that controls how sharp or spread out a probability distribution is.
A temperature alpha can interpolate between soft and hard max. - valid (adjective)
- Correct and properly justified.
There is also a valid infinite-horizon version of the recursion. - quantity (noun)
- An amount or numerical value being discussed or computed.
The optimal distribution is the exponential of that quantity. - risky (adjective)
- Involving a chance of a bad outcome.
Exact inference biases the policy toward risky actions. - effectively (adverb)
- In practice, even if not stated directly.
Exact inference effectively asks what you did given that you got lucky. - machinery (noun)
- The set of methods or tools used to solve a problem.
This part reuses the variational inference machinery from earlier. - match (verb)
- To make one thing equal to or the same as another.
The variational distribution is restricted to match the true dynamics.
Chapters
- 0:00 Intro
- 1:03 Optimism Problem
- 7:55 Variational Distribution
- 14:24 Dynamic Programming
- 18:55 Back or Pass
← Lecture 19, Control as Inference, Part 2 · Lecture 19, Control as Inference, Part 4 →
