Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 58 of 99 · 6:49

Lecture 13, Part 5: Exploration

CS 285: Lecture 13, Part 5 on YouTube

Study guide

What this lecture covers

This short part adapts Thompson sampling (posterior sampling) to deep RL. Since a bandit's reward model becomes a Q-function in an MDP, the lecture describes sampling a Q-function from an ensemble, acting on it for a full episode, then updating the ensemble on the collected data, since Q-learning is off-policy and can train on data from any exploration strategy.

Key ideas

  • Q-function as the sampled model: in deep RL, posterior sampling means maintaining a distribution over Q-functions and acting greedily on one sampled function per episode, rather than sampling reward parameters as in bandits.
  • Bootstrap ensembles: approximate a posterior by resampling the dataset with replacement to train several independent Q-networks, then sampling one at random to act on.
  • Shared-body, multi-head networks: a cheaper approximation trains one network with a shared trunk and multiple output heads, giving correlated but still somewhat varied behavior across heads.
  • Why it beats random exploration: acting on one consistent sampled Q-function for a whole episode produces coherent behavior (such as reliably surfacing for air in Seaquest), unlike epsilon-greedy, which rarely chains together the many correct random actions needed for such a strategy.
  • Trade-offs: bootstrapped exploration needs no reward modification and no exploration-exploitation hyperparameter, but tends to underperform well-tuned count-based or pseudo-count bonuses, and does poorly on very hard exploration problems like Montezuma's Revenge.

Before you watch

  • Watch the earlier parts of Lecture 13 on Thompson sampling in bandits and on count-based and pseudo-count exploration, which this method is compared against.

Check your understanding

  1. Why can Q-learning train on data collected under many different sampled Q-functions without modification?
  2. How does acting on one sampled Q-function per episode help with games like Seaquest, where a long consistent action sequence is required?
  3. What is the practical trade-off between bootstrapped exploration and bonus-based methods like pseudo-counts?

Vocabulary

Thompson sampling (noun)
A strategy that samples one belief and acts optimally on it, then updates the belief.
Thompson sampling adapts nicely to deep RL.
posterior sampling (phrase)
Choosing an action by drawing one hypothesis from a belief distribution.
Posterior sampling in RL means sampling a Q-function.
bootstrap ensemble (phrase)
A set of independently trained models used together to represent uncertainty.
A bootstrap ensemble approximates a posterior over Q-functions.
resample (verb)
To draw a new dataset from an existing one, allowing repeats.
We resample the dataset with replacement for each network.
shared trunk (phrase)
A common set of layers used by multiple separate output branches.
A shared-body network has a shared trunk with several heads.
multi-head (adjective)
Having several separate output branches from one shared network.
A multi-head network gives correlated but varied outputs.
correlated (adjective)
Related to each other, tending to change together.
The heads produce correlated but somewhat varied behavior.
coherent (adjective)
Consistent and logically connected, making sense as a whole.
Acting on one Q-function gives coherent behavior.
chain together (phrasal verb)
To connect a series of correct steps in a row.
Random exploration rarely chains together the right actions.
hyperparameter (noun)
A setting chosen before training that controls how a model learns.
This method needs no extra exploration hyperparameter.
underperform (verb)
To do worse than an alternative method.
Bootstrapped exploration tends to underperform pseudo-counts.
well-tuned (adjective)
Carefully adjusted to work as well as possible.
Well-tuned count-based bonuses often do better.
adapt (verb)
To change an idea so it fits a new setting.
The lecture adapts Thompson sampling to deep RL.
surface for air (phrase)
To come up to the top in order to breathe, as in the game Seaquest.
Reliably surfacing for air needs consistent behavior.
sample (a function) (verb)
To draw one specific instance from a set of possibilities.
We sample one Q-function per episode.
consistent (adjective)
Behaving the same way throughout, without switching randomly.
The agent needs consistent behavior across the episode.
very hard (phrase)
Extremely difficult to solve or succeed at.
Montezuma's Revenge is a very hard exploration problem.
no reward modification (phrase)
Not needing to change the environment's reward signal.
This method needs no reward modification.
cheaper approximation (phrase)
A less accurate but less costly way of estimating something.
Shared-body networks are a cheaper approximation.
collected data (phrase)
Information gathered from past interactions, used for training.
Q-learning trains on collected data from any policy.

Chapters

← Lecture 13, Part 4: Exploration · Lecture 13, Part 6: Exploration →