Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 57 of 99 · 13:02
Lecture 13, Part 4: Exploration
Study guide
What this lecture covers
This part continues the survey of optimism-based exploration methods, moving past density-model pseudo-counts to three other techniques for measuring state novelty. It covers hash-based counting, where similar states are compressed to the same code so exact counting becomes feasible again, classifier-based density estimation (EX2), which sidesteps generative modeling entirely, and heuristic prediction-error methods that use the error of a learned model as a novelty signal.
Together these give a sense of the range of practical tricks used to approximate "have I seen something like this before" without needing an exact count or a high-quality generative model.
Key ideas
- Hashing exploration: compress each state into a short
k-bit code via an encoder, then count occurrences of each code; shorter codes cause more collisions, which broadens the notion of "similar" states. - Learned hashing via autoencoders: training the encoder for reconstruction accuracy makes collisions more likely between states that actually look alike, rather than colliding arbitrarily.
- EX2 classifier-based density: train a classifier to distinguish a given state from all previously seen states; an easily distinguished (novel) state gets a low estimated density, derived from the classifier's output via the Bayes-optimal classifier formula.
- Amortized classifier: instead of training a separate classifier per state, train one network conditioned on both the exemplar state and the state being classified, updated online as new states arrive.
- Prediction-error novelty: fit a function
f_hatto approximate some target functionf_staron observed states and actions; large prediction error signals states far from what has been seen, so error itself becomes the novelty bonus. - Random network target: a particularly simple and effective choice is to let
f_starbe a randomly initialized, untrained neural network, since it only needs to vary meaningfully across the state space, not be semantically meaningful.
Before you watch
- Watch the previous part of Lecture 13, which introduces pseudo-counts from density models; this part presents these methods as alternatives to that approach.
Check your understanding
- How does hash-based exploration make exact counting feasible again in large state spaces?
- Why does training the hash encoder for reconstruction accuracy improve which states collide into the same code?
- How does the EX2 method turn a classifier's output into a density estimate for a state?
- Why can a randomly initialized, untrained network serve as a useful novelty-detection target?
Vocabulary
- hash (noun)
- A short code produced from an input, used to group similar items together.
We compress each state into a short hash code. - encoder (noun)
- A network that compresses raw input into a shorter representation.
An encoder maps each state to a k-bit code. - collision (noun)
- When two different inputs produce the same output code.
Shorter codes cause more hash collisions. - compress (verb)
- To reduce something to a smaller size while keeping important information.
Hashing compresses each state to a short code. - autoencoder (noun)
- A network trained to compress and then rebuild its own input.
The hash encoder can be trained as an autoencoder. - reconstruction accuracy (phrase)
- How well a rebuilt output matches the original input.
Training for reconstruction accuracy improves useful collisions. - classifier-based (adjective)
- Relying on a model that sorts inputs into categories.
EX2 is a classifier-based density estimation method. - distinguish (verb)
- To tell two things apart based on a difference.
The classifier learns to distinguish a state from previous ones. - amortized (adjective)
- Sharing one cost or computation across many uses, instead of repeating it each time.
An amortized classifier avoids training separately per state. - exemplar (noun)
- A specific example used as a reference point.
The network is conditioned on an exemplar state. - prediction error (phrase)
- The difference between a model's prediction and the true value.
Prediction error can serve as a novelty signal. - target function (phrase)
- The function a model is trying to match or approximate.
We fit a function to approximate a target function. - random network (phrase)
- A neural network with untrained, randomly chosen weights.
A random network can be a useful novelty target. - novelty (noun)
- How new or unfamiliar something is compared to what's already been seen.
Prediction error signals a state's novelty. - range of tricks (phrase)
- A variety of practical methods used to solve a problem.
This gives a sense of the range of tricks used in practice. - sidestep (verb)
- To avoid a problem by taking a different approach.
EX2 sidesteps generative modeling entirely. - arbitrarily (adverb)
- Randomly or without a clear pattern, not based on similarity.
Poor hashing causes states to collide arbitrarily. - semantically meaningful (phrase)
- Carrying real, understandable meaning.
A random network doesn't need to be semantically meaningful. - particularly (adverb)
- Especially, more than usual.
This is a particularly simple and effective choice. - together with (phrase)
- Combined with something else.
Together with hashing, these give a sense of the range of tricks.
Chapters
← Lecture 13, Part 3: Exploration · Lecture 13, Part 5: Exploration →
