Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 58 of 99 · 6:49
Lecture 13, Part 5: Exploration
Study guide
What this lecture covers
This short part adapts Thompson sampling (posterior sampling) to deep RL. Since a bandit's reward model becomes a Q-function in an MDP, the lecture describes sampling a Q-function from an ensemble, acting on it for a full episode, then updating the ensemble on the collected data, since Q-learning is off-policy and can train on data from any exploration strategy.
Key ideas
- Q-function as the sampled model: in deep RL, posterior sampling means maintaining a distribution over Q-functions and acting greedily on one sampled function per episode, rather than sampling reward parameters as in bandits.
- Bootstrap ensembles: approximate a posterior by resampling the dataset with replacement to train several independent Q-networks, then sampling one at random to act on.
- Shared-body, multi-head networks: a cheaper approximation trains one network with a shared trunk and multiple output heads, giving correlated but still somewhat varied behavior across heads.
- Why it beats random exploration: acting on one consistent sampled Q-function for a whole episode produces coherent behavior (such as reliably surfacing for air in Seaquest), unlike epsilon-greedy, which rarely chains together the many correct random actions needed for such a strategy.
- Trade-offs: bootstrapped exploration needs no reward modification and no exploration-exploitation hyperparameter, but tends to underperform well-tuned count-based or pseudo-count bonuses, and does poorly on very hard exploration problems like Montezuma's Revenge.
Before you watch
- Watch the earlier parts of Lecture 13 on Thompson sampling in bandits and on count-based and pseudo-count exploration, which this method is compared against.
Check your understanding
- Why can Q-learning train on data collected under many different sampled Q-functions without modification?
- How does acting on one sampled Q-function per episode help with games like Seaquest, where a long consistent action sequence is required?
- What is the practical trade-off between bootstrapped exploration and bonus-based methods like pseudo-counts?
Vocabulary
- Thompson sampling (noun)
- A strategy that samples one belief and acts optimally on it, then updates the belief.
Thompson sampling adapts nicely to deep RL. - posterior sampling (phrase)
- Choosing an action by drawing one hypothesis from a belief distribution.
Posterior sampling in RL means sampling a Q-function. - bootstrap ensemble (phrase)
- A set of independently trained models used together to represent uncertainty.
A bootstrap ensemble approximates a posterior over Q-functions. - resample (verb)
- To draw a new dataset from an existing one, allowing repeats.
We resample the dataset with replacement for each network. - shared trunk (phrase)
- A common set of layers used by multiple separate output branches.
A shared-body network has a shared trunk with several heads. - multi-head (adjective)
- Having several separate output branches from one shared network.
A multi-head network gives correlated but varied outputs. - correlated (adjective)
- Related to each other, tending to change together.
The heads produce correlated but somewhat varied behavior. - coherent (adjective)
- Consistent and logically connected, making sense as a whole.
Acting on one Q-function gives coherent behavior. - chain together (phrasal verb)
- To connect a series of correct steps in a row.
Random exploration rarely chains together the right actions. - hyperparameter (noun)
- A setting chosen before training that controls how a model learns.
This method needs no extra exploration hyperparameter. - underperform (verb)
- To do worse than an alternative method.
Bootstrapped exploration tends to underperform pseudo-counts. - well-tuned (adjective)
- Carefully adjusted to work as well as possible.
Well-tuned count-based bonuses often do better. - adapt (verb)
- To change an idea so it fits a new setting.
The lecture adapts Thompson sampling to deep RL. - surface for air (phrase)
- To come up to the top in order to breathe, as in the game Seaquest.
Reliably surfacing for air needs consistent behavior. - sample (a function) (verb)
- To draw one specific instance from a set of possibilities.
We sample one Q-function per episode. - consistent (adjective)
- Behaving the same way throughout, without switching randomly.
The agent needs consistent behavior across the episode. - very hard (phrase)
- Extremely difficult to solve or succeed at.
Montezuma's Revenge is a very hard exploration problem. - no reward modification (phrase)
- Not needing to change the environment's reward signal.
This method needs no reward modification. - cheaper approximation (phrase)
- A less accurate but less costly way of estimating something.
Shared-body networks are a cheaper approximation. - collected data (phrase)
- Information gathered from past interactions, used for training.
Q-learning trains on collected data from any policy.
Chapters
← Lecture 13, Part 4: Exploration · Lecture 13, Part 6: Exploration →
