Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 77 of 99 · 21:05
Lecture 19, Control as Inference, Part 1
Study guide
What this lecture covers
This lecture opens a new topic: reframing optimal control and reinforcement learning as probabilistic inference. It starts from an empirical observation about behavior, that humans and animals reaching a goal usually get there but rarely take the exact optimal trajectory, and asks whether a better model of "near optimal" behavior than the standard deterministic-optimal-policy framework exists.
By the end, you should understand why standard RL formulations cannot explain stochastic behavior (since fully observed MDPs admit deterministic optimal policies), how adding binary "optimality" variables to a graphical model produces trajectory probabilities that favor high reward but assign non-zero probability to any trajectory, and why this view sets up both better exploration and, in the next lecture, inverse reinforcement learning.
Key ideas
- Near-optimal, not perfectly optimal, behavior: animals and humans tend to reach a goal reliably but vary in exactly how, especially on aspects of the task that barely affect the outcome.
- Deterministic optimal policies: for any fully observed MDP with a reward linear in the state-action marginal, there exists a deterministic optimal policy, so standard RL cannot naturally explain randomized behavior.
- Optimality variable: a binary variable
O_tadded to the graphical model at each timestep, withp(O_t = true | s_t, a_t) = exp(r(s_t, a_t)), requiring rewards to be non-positive (achievable without loss of generality by shifting the reward). - Trajectory probability under optimality: conditioning on all
O_tbeing true givesp(tau | O_{1:T}) ∝ p(tau) * exp(sum of rewards along tau), so higher-reward trajectories are more likely but lower-reward ones remain possible with exponentially decreasing probability. - Explaining stochastic behavior: this model naturally explains why an agent might behave randomly among near-equally-good options while strongly avoiding clearly bad ones.
- Motivations beyond modeling animals: this framework helps explain suboptimal demonstrator behavior for inverse reinforcement learning, enables solving control problems with inference algorithms, and helps explain why stochastic policies can aid exploration and transfer.
- Backward and forward messages: inference in this chain-structured graphical model uses backward messages (probability of future optimality given the current state and action, from which the policy can be recovered) and forward messages (probability of reaching a state given past optimality), both to be derived in the next part.
Before you watch
- Review the definition of an MDP and the fact that optimal policies in fully observed MDPs can be taken to be deterministic, covered earlier in the course.
- Some familiarity with graphical models, message passing, or hidden Markov models is helpful since the inference approach parallels HMM-style forward-backward algorithms.
Check your understanding
- Why can't a standard, fully-observed MDP with a deterministic optimal policy explain the kind of variable, near-optimal behavior seen in animals?
- What role does the constraint that rewards must be negative play in defining
p(O_t = true | s_t, a_t), and why is this not a real limitation? - Under the optimality-variable model, why does a trajectory with much lower reward become exponentially less likely rather than simply impossible?
- What is the practical difference between a backward message and a forward message in this graphical model?
Vocabulary
- control as inference (noun)
- The idea of treating an optimal control problem as a probability inference problem.
Control as inference explains near-optimal, varied behavior. - graphical model (noun)
- A diagram showing how random variables depend on each other.
The lecture builds a graphical model with an optimality variable. - near-optimal (adjective)
- Close to the best possible, but not perfectly so.
Animals show near-optimal, slightly varied behavior. - deterministic (adjective)
- Always producing the exact same output for the same input, with no randomness.
A deterministic optimal policy always chooses the same action. - optimality variable (noun)
- A binary variable added to a model that represents whether a step was optimal.
The optimality variable O_t is true with probability based on reward. - binary variable (noun)
- A variable that can only take one of two values, usually true or false.
O_t is a binary variable at every timestep. - proportional to (phrase)
- Changing at the same rate as another quantity, staying in a fixed ratio.
The trajectory probability is proportional to the exponential of total reward. - exponentially (adverb)
- Increasing or decreasing very fast, based on repeated multiplication.
Low-reward trajectories become exponentially less likely. - demonstrator (noun)
- A person or agent who shows an example of how to perform a task.
The demonstrator's behavior might not be perfectly optimal. - transfer (noun)
- The reuse of knowledge or skills learned in one setting for another.
Stochastic policies can help with exploration and transfer. - backward message (noun)
- A quantity computed by working from the end of a sequence backward to the start.
The backward message gives the probability of future optimality. - forward message (noun)
- A quantity computed by working from the start of a sequence forward.
The forward message tracks the probability of reaching a state. - probabilistic inference (noun)
- The process of computing probabilities of unknown variables given known ones.
This lecture reframes control as probabilistic inference. - timestep (noun)
- One single step in a sequence over time.
The optimality variable is added at each timestep. - chain-structured (adjective)
- Arranged as a simple sequence, one item linked to the next.
The graphical model is chain-structured over timesteps. - condition on (phrasal verb)
- To treat a fact as known and calculate probabilities based on it.
We condition on all O_t being true to get the trajectory probability. - message passing (noun)
- An inference method where nodes in a graph exchange summarized information.
Message passing computes probabilities in the chain-structured model. - admit (verb)
- To allow or have as a valid possibility.
A fully observed MDP admits a deterministic optimal policy. - aspect (noun)
- One particular part or feature of something.
Behavior varies most on aspects of the task that barely affect the outcome. - reliably (adverb)
- Consistently and dependably, almost every time.
Animals reliably reach their goal even if the path varies. - achievable (adjective)
- Possible to reach or accomplish.
Non-positive rewards are achievable by simply shifting the reward. - shift (verb)
- To move a value up or down by a fixed amount.
Any reward can be shifted to make it non-positive. - non-positive (adjective)
- Zero or less than zero.
The optimality variable requires rewards to be non-positive. - without loss of generality (phrase)
- A phrase meaning an assumption can be made freely because it does not limit the result.
Non-positive rewards can be assumed without loss of generality. - favor (verb)
- To make something more likely or preferred over alternatives.
The trajectory distribution favors high-reward paths.
Chapters
- 0:00 Introduction
- 0:23 Questions
- 1:22 Human Behavior
- 4:20 Animal Behavior
- 8:43 Graphical Models
- 18:41 Inference
← Lecture 18, Variational Inference, Part 4 · Lecture 19, Control as Inference, Part 2 →
