Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 59 of 99 · 12:40

Lecture 13, Part 6: Exploration

CS 285: Lecture 13, Part 6 on YouTube

Study guide

What this lecture covers

This final part of Lecture 13 covers the third class of deep RL exploration methods: information gain. It addresses the practical question of what to gain information about (state density, or more usefully, the transition dynamics) and how to approximate the otherwise intractable information-gain computation, focusing on the variational inference approach used in the paper VIME.

The lecture closes the exploration topic with a recap of the three method families covered across the lecture and a reading list for further study.

Key ideas

  • Information gain about dynamics: since reward is often sparse, information gain about the transition model p(s'|s,a) is a more useful target than information gain about the reward itself.
  • Prediction gain: a crude approximation to information gain, computed as the change in log-density of a state before and after updating a density model on it, similar in spirit to pseudo-counts.
  • Information gain as KL divergence: information gain about a variable z from an observation y equals the KL divergence between the posterior p(z|y) and the prior p(z).
  • Variational approximation (VIME): since the true posterior over dynamics-model parameters is intractable, approximate it with a variational distribution q(theta|phi) (for example, independent Gaussians per parameter) and measure information gain as the KL divergence between the variational posterior before and after a new transition.
  • Bayesian neural network updates: after each transition, update the variational parameters phi to phi'; the KL divergence between the two Gaussian posteriors has a closed form and serves as the exploration bonus.
  • Simplifying to model error: dropping the full Bayesian treatment and just measuring how much a model's parameters or predictions change is a related, simpler family of error-based novelty bonuses.
  • Recap: the lecture reiterates the three method families covered across Lecture 13 -- optimistic exploration (counts and pseudo-counts), Thompson sampling (bootstrapped Q-function ensembles), and information gain (variational approximations).

Before you watch

  • Watch the earlier parts of Lecture 13, especially the pseudo-counts and bootstrapped-ensemble methods, since this part builds on and contrasts with both.
  • Familiarity with KL divergence and the basic idea of variational inference is helpful, though the lecture summarizes what is needed.

Check your understanding

  1. Why is information gain about the dynamics generally more useful for exploration than information gain about the reward?
  2. How does prediction gain approximate information gain without explicitly computing pseudo-counts?
  3. Why does information gain equal the KL divergence between the posterior and prior over the variable of interest?
  4. What practical difficulty motivates using a variational approximation instead of the true posterior over dynamics-model parameters?

Vocabulary

information gain (phrase)
How much observing something reduces uncertainty about an unknown quantity.
Information gain about dynamics guides exploration here.
prediction gain (phrase)
A rough measure of information gain based on how a density changes after seeing new data.
Prediction gain is a crude approximation to information gain.
KL divergence (noun)
A measure of how different one probability distribution is from another.
Information gain equals a KL divergence between posterior and prior.
prior (noun)
A belief held before seeing new evidence.
The KL is measured between the posterior and the prior.
posterior (noun)
An updated belief after seeing new evidence.
The posterior changes after each new transition.
variational approximation (phrase)
Replacing a hard-to-compute distribution with a simpler, trainable one.
VIME uses a variational approximation to the true posterior.
variational distribution (phrase)
A chosen, simpler distribution used to approximate a true but intractable one.
The variational distribution is parameterized by phi.
closed form (phrase)
An exact formula that can be computed directly.
The KL between two Gaussians has a closed form.
Bayesian neural network (noun)
A neural network with a distribution over its weights instead of fixed values.
A Bayesian neural network models uncertainty over dynamics.
model error (phrase)
The mismatch between a model's prediction and the true outcome.
A simpler family just measures model error.
recap (noun)
A short summary of ideas already covered.
The lecture closes with a recap of three method families.
reading list (phrase)
A set of recommended materials for further study.
The lecture ends with a reading list on exploration.
in spirit (phrase)
Similar in overall idea, even if the details differ.
Prediction gain is similar in spirit to pseudo-counts.
practical difficulty (phrase)
A real obstacle encountered when trying to apply an idea.
A practical difficulty motivates the variational approach.
close (a topic) (verb)
To finish discussing a subject.
This closes the exploration topic for the course.
target (of learning) (noun)
The thing a method is trying to learn about or predict.
The dynamics model is a more useful target than the reward.
sparse (reward) (adjective)
Given only rarely, with long gaps of no feedback.
Reward is often sparse, so dynamics information matters more.
measure (v) (verb)
To calculate or quantify some property.
We measure information gain as a KL divergence.
simplify (to model error) (verb)
To make something less complex by dropping some of its detail.
We can simplify to just measuring model error.
full treatment (phrase)
A complete, rigorous version of a method, without shortcuts.
Dropping the full Bayesian treatment gives a simpler family.

Chapters

← Lecture 13, Part 5: Exploration · Lecture 14, Part 1 →