Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 55 of 99 · 15:48
Lecture 13, Part 2: Exploration
Study guide
What this lecture covers
Continuing the exploration topic, this part presents three families of provably good strategies for the multi-armed bandit setting introduced previously: optimistic exploration (upper confidence bounds), posterior sampling (Thompson sampling), and information-gain-based methods. Each is a practical, tractable stand-in for solving the full POMDP formulation of the bandit exactly.
The lecture explains the intuition and rough guarantees behind each strategy and closes by noting the common thread: every method needs some form of uncertainty estimate and some assumption about the value of new information, since you cannot directly optimize toward reward when you don't yet know where it is.
Key ideas
- Optimistic exploration (UCB): pick the action maximizing the estimated mean reward plus a bonus proportional to uncertainty; a simple bonus of
sqrt(2 log t / N(a))achieves the asymptotically optimalO(log t)regret. - Optimism in the face of uncertainty: try an arm until you are confident it is not great, then stop; unknown arms are assumed to possibly be good.
- Thompson sampling (posterior sampling): maintain a belief distribution over the bandit's unknown parameters, sample one hypothesis, act optimally under it, then update the belief from the observed outcome.
- Model-based vs. model-free: UCB is largely model-free (it just counts pulls), while Thompson sampling and information-gain methods explicitly maintain a belief over the underlying model.
- Information gain: quantifies how much observing a variable
y(such as a reward) is expected to reduce the entropy of an unknown quantityz(such as model parameters), used to prefer actions that teach you the most. - Information-directed sampling: a decision rule that trades off expected sub-optimality (
delta_a) against information gain (g_a), roughly minimizingdelta_a^2 / g_a, so as to avoid wasted or uninformative exploration. - Common assumption across all three: since you can't directly chase reward you don't yet understand, each method substitutes a proxy value for new information -- optimism assumes unknowns might be good, Thompson sampling trusts its sampled model, and information gain values learning itself.
Before you watch
- Watch the first part of Lecture 13, which defines bandits, regret, and the exploration-exploitation trade-off used throughout this part.
Check your understanding
- Why does adding an uncertainty-scaled bonus to the estimated mean reward lead to a provably good exploration strategy?
- How does Thompson sampling differ from directly solving the bandit's POMDP formulation?
- What is information gain measuring, and why is it defined as an expectation over the unknown observation
y? - What assumption do optimism, Thompson sampling, and information gain each make about the value of unknown outcomes?
Vocabulary
- optimistic exploration (phrase)
- A strategy that assumes uncertain options might be good until proven otherwise.
Optimistic exploration tries unknown arms before giving up on them. - upper confidence bound (UCB) (phrase)
- A rule that picks the option with the best estimate plus an uncertainty bonus.
UCB adds a bonus proportional to uncertainty. - bonus (noun)
- An extra value added to encourage a certain choice.
The bonus shrinks as an arm is pulled more often. - asymptotically optimal (phrase)
- Becoming the best possible strategy as time goes to infinity.
This bonus achieves asymptotically optimal regret. - posterior sampling (phrase)
- Choosing actions by sampling one hypothesis from a belief distribution and acting on it.
Thompson sampling is a form of posterior sampling. - belief distribution (phrase)
- A probability distribution representing current uncertainty about an unknown quantity.
We maintain a belief distribution over bandit parameters. - hypothesis (noun)
- A possible explanation or guess about how things work.
We sample one hypothesis and act optimally under it. - model-based (adjective)
- Explicitly maintaining and using a model of the underlying system.
Thompson sampling is model-based, unlike UCB. - information gain (phrase)
- How much observing something reduces uncertainty about an unknown.
Information gain measures how much a reward teaches us. - entropy (noun)
- A measure of how uncertain or spread out a distribution is.
Information gain reduces the entropy of an unknown quantity. - information-directed sampling (phrase)
- A method that balances expected loss against how much is learned from an action.
Information-directed sampling avoids wasted exploration. - sub-optimality (noun)
- How much worse an option is compared to the best one.
We trade off expected sub-optimality against information gain. - proxy (noun)
- A substitute value used to stand in for something hard to measure directly.
Each method substitutes a proxy value for new information. - provably good (phrase)
- Shown by mathematical proof to perform well.
These are provably good exploration strategies. - stand-in (noun)
- Something used in place of another, harder thing.
This is a tractable stand-in for solving the full POMDP. - common thread (phrase)
- An idea shared across several different things.
The common thread is that every method needs uncertainty. - rough guarantee (phrase)
- An approximate, not exact, mathematical promise.
The lecture explains the rough guarantees behind each method. - trust (its sampled model) (verb)
- To act as if something is true, based on confidence in it.
Thompson sampling trusts its sampled model when acting. - wasted exploration (phrase)
- Effort spent trying options that don't teach anything useful.
Information gain avoids wasted exploration. - quantify (verb)
- To express something as a specific measurable amount.
Information gain quantifies how much an observation teaches us.
Chapters
- 0:00 Intro
- 0:19 How can we beat the bandit?
- 1:01 Optimistic exploration
- 4:19 Probability matching/posterior sampling
- 9:51 Information gain example
- 12:36 General themes UCB
- 14:36 Why should we care?
← Lecture 13, Part 1: Exploration · Lecture 13, Part 3: Exploration →
