Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 57 of 99 · 13:02

Lecture 13, Part 4: Exploration

CS 285: Lecture 13, Part 4 on YouTube

Study guide

What this lecture covers

This part continues the survey of optimism-based exploration methods, moving past density-model pseudo-counts to three other techniques for measuring state novelty. It covers hash-based counting, where similar states are compressed to the same code so exact counting becomes feasible again, classifier-based density estimation (EX2), which sidesteps generative modeling entirely, and heuristic prediction-error methods that use the error of a learned model as a novelty signal.

Together these give a sense of the range of practical tricks used to approximate "have I seen something like this before" without needing an exact count or a high-quality generative model.

Key ideas

  • Hashing exploration: compress each state into a short k-bit code via an encoder, then count occurrences of each code; shorter codes cause more collisions, which broadens the notion of "similar" states.
  • Learned hashing via autoencoders: training the encoder for reconstruction accuracy makes collisions more likely between states that actually look alike, rather than colliding arbitrarily.
  • EX2 classifier-based density: train a classifier to distinguish a given state from all previously seen states; an easily distinguished (novel) state gets a low estimated density, derived from the classifier's output via the Bayes-optimal classifier formula.
  • Amortized classifier: instead of training a separate classifier per state, train one network conditioned on both the exemplar state and the state being classified, updated online as new states arrive.
  • Prediction-error novelty: fit a function f_hat to approximate some target function f_star on observed states and actions; large prediction error signals states far from what has been seen, so error itself becomes the novelty bonus.
  • Random network target: a particularly simple and effective choice is to let f_star be a randomly initialized, untrained neural network, since it only needs to vary meaningfully across the state space, not be semantically meaningful.

Before you watch

  • Watch the previous part of Lecture 13, which introduces pseudo-counts from density models; this part presents these methods as alternatives to that approach.

Check your understanding

  1. How does hash-based exploration make exact counting feasible again in large state spaces?
  2. Why does training the hash encoder for reconstruction accuracy improve which states collide into the same code?
  3. How does the EX2 method turn a classifier's output into a density estimate for a state?
  4. Why can a randomly initialized, untrained network serve as a useful novelty-detection target?

Vocabulary

hash (noun)
A short code produced from an input, used to group similar items together.
We compress each state into a short hash code.
encoder (noun)
A network that compresses raw input into a shorter representation.
An encoder maps each state to a k-bit code.
collision (noun)
When two different inputs produce the same output code.
Shorter codes cause more hash collisions.
compress (verb)
To reduce something to a smaller size while keeping important information.
Hashing compresses each state to a short code.
autoencoder (noun)
A network trained to compress and then rebuild its own input.
The hash encoder can be trained as an autoencoder.
reconstruction accuracy (phrase)
How well a rebuilt output matches the original input.
Training for reconstruction accuracy improves useful collisions.
classifier-based (adjective)
Relying on a model that sorts inputs into categories.
EX2 is a classifier-based density estimation method.
distinguish (verb)
To tell two things apart based on a difference.
The classifier learns to distinguish a state from previous ones.
amortized (adjective)
Sharing one cost or computation across many uses, instead of repeating it each time.
An amortized classifier avoids training separately per state.
exemplar (noun)
A specific example used as a reference point.
The network is conditioned on an exemplar state.
prediction error (phrase)
The difference between a model's prediction and the true value.
Prediction error can serve as a novelty signal.
target function (phrase)
The function a model is trying to match or approximate.
We fit a function to approximate a target function.
random network (phrase)
A neural network with untrained, randomly chosen weights.
A random network can be a useful novelty target.
novelty (noun)
How new or unfamiliar something is compared to what's already been seen.
Prediction error signals a state's novelty.
range of tricks (phrase)
A variety of practical methods used to solve a problem.
This gives a sense of the range of tricks used in practice.
sidestep (verb)
To avoid a problem by taking a different approach.
EX2 sidesteps generative modeling entirely.
arbitrarily (adverb)
Randomly or without a clear pattern, not based on similarity.
Poor hashing causes states to collide arbitrarily.
semantically meaningful (phrase)
Carrying real, understandable meaning.
A random network doesn't need to be semantically meaningful.
particularly (adverb)
Especially, more than usual.
This is a particularly simple and effective choice.
together with (phrase)
Combined with something else.
Together with hashing, these give a sense of the range of tricks.

Chapters

← Lecture 13, Part 3: Exploration · Lecture 13, Part 5: Exploration →