Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 59 of 99 · 12:40
Lecture 13, Part 6: Exploration
Study guide
What this lecture covers
This final part of Lecture 13 covers the third class of deep RL exploration methods: information gain. It addresses the practical question of what to gain information about (state density, or more usefully, the transition dynamics) and how to approximate the otherwise intractable information-gain computation, focusing on the variational inference approach used in the paper VIME.
The lecture closes the exploration topic with a recap of the three method families covered across the lecture and a reading list for further study.
Key ideas
- Information gain about dynamics: since reward is often sparse, information gain about the transition model
p(s'|s,a)is a more useful target than information gain about the reward itself. - Prediction gain: a crude approximation to information gain, computed as the change in log-density of a state before and after updating a density model on it, similar in spirit to pseudo-counts.
- Information gain as KL divergence: information gain about a variable
zfrom an observationyequals the KL divergence between the posteriorp(z|y)and the priorp(z). - Variational approximation (VIME): since the true posterior over dynamics-model parameters is intractable, approximate it with a variational distribution
q(theta|phi)(for example, independent Gaussians per parameter) and measure information gain as the KL divergence between the variational posterior before and after a new transition. - Bayesian neural network updates: after each transition, update the variational parameters
phitophi'; the KL divergence between the two Gaussian posteriors has a closed form and serves as the exploration bonus. - Simplifying to model error: dropping the full Bayesian treatment and just measuring how much a model's parameters or predictions change is a related, simpler family of error-based novelty bonuses.
- Recap: the lecture reiterates the three method families covered across Lecture 13 -- optimistic exploration (counts and pseudo-counts), Thompson sampling (bootstrapped Q-function ensembles), and information gain (variational approximations).
Before you watch
- Watch the earlier parts of Lecture 13, especially the pseudo-counts and bootstrapped-ensemble methods, since this part builds on and contrasts with both.
- Familiarity with KL divergence and the basic idea of variational inference is helpful, though the lecture summarizes what is needed.
Check your understanding
- Why is information gain about the dynamics generally more useful for exploration than information gain about the reward?
- How does prediction gain approximate information gain without explicitly computing pseudo-counts?
- Why does information gain equal the KL divergence between the posterior and prior over the variable of interest?
- What practical difficulty motivates using a variational approximation instead of the true posterior over dynamics-model parameters?
Vocabulary
- information gain (phrase)
- How much observing something reduces uncertainty about an unknown quantity.
Information gain about dynamics guides exploration here. - prediction gain (phrase)
- A rough measure of information gain based on how a density changes after seeing new data.
Prediction gain is a crude approximation to information gain. - KL divergence (noun)
- A measure of how different one probability distribution is from another.
Information gain equals a KL divergence between posterior and prior. - prior (noun)
- A belief held before seeing new evidence.
The KL is measured between the posterior and the prior. - posterior (noun)
- An updated belief after seeing new evidence.
The posterior changes after each new transition. - variational approximation (phrase)
- Replacing a hard-to-compute distribution with a simpler, trainable one.
VIME uses a variational approximation to the true posterior. - variational distribution (phrase)
- A chosen, simpler distribution used to approximate a true but intractable one.
The variational distribution is parameterized by phi. - closed form (phrase)
- An exact formula that can be computed directly.
The KL between two Gaussians has a closed form. - Bayesian neural network (noun)
- A neural network with a distribution over its weights instead of fixed values.
A Bayesian neural network models uncertainty over dynamics. - model error (phrase)
- The mismatch between a model's prediction and the true outcome.
A simpler family just measures model error. - recap (noun)
- A short summary of ideas already covered.
The lecture closes with a recap of three method families. - reading list (phrase)
- A set of recommended materials for further study.
The lecture ends with a reading list on exploration. - in spirit (phrase)
- Similar in overall idea, even if the details differ.
Prediction gain is similar in spirit to pseudo-counts. - practical difficulty (phrase)
- A real obstacle encountered when trying to apply an idea.
A practical difficulty motivates the variational approach. - close (a topic) (verb)
- To finish discussing a subject.
This closes the exploration topic for the course. - target (of learning) (noun)
- The thing a method is trying to learn about or predict.
The dynamics model is a more useful target than the reward. - sparse (reward) (adjective)
- Given only rarely, with long gaps of no feedback.
Reward is often sparse, so dynamics information matters more. - measure (v) (verb)
- To calculate or quantify some property.
We measure information gain as a KL divergence. - simplify (to model error) (verb)
- To make something less complex by dropping some of its detail.
We can simplify to just measuring model error. - full treatment (phrase)
- A complete, rigorous version of a method, without shortcuts.
Dropping the full Bayesian treatment gives a simpler family.
