Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 60 of 99 · 14:17
Lecture 14, Part 1
Study guide
What this lecture covers
This lecture opens a different angle on exploration than the previous lecture's count- and bonus-based methods: what if there is no reward signal at all? The motivating idea is that an agent (or a child at play) could acquire a broad repertoire of skills through unsupervised exploration, then repurpose that knowledge quickly once a real goal is given, similar to a home robot practicing in a kitchen before being asked to do the dishes.
To build toward the algorithms covered in later parts (goal-reaching without reward, state-distribution matching, coverage-based objectives, and skill discovery), this part reviews the information-theoretic building blocks: entropy and mutual information, and introduces "empowerment" as a concrete example of how mutual information can express a meaningful RL objective.
Key ideas
- Reward-free exploration: instead of seeking reward, the agent aims to acquire diverse, reusable skills or knowledge that can later be repurposed for arbitrary goals.
- Entropy
H(p(x)): measures how broad a distribution is; a uniform distribution has maximal entropy, a distribution peaked on one value has minimal entropy. - Mutual information
I(x;y): the KL divergence between the joint distributionp(x,y)and the product of marginalsp(x)p(y); it is zero whenxandyare independent and grows as they become more dependent. - Mutual information as entropy reduction:
I(x;y) = H(y) - H(y|x), the reduction in uncertainty aboutyfrom observingx, connecting it to the information-gain idea from the previous lecture. - State marginal entropy
H(pi(s)): the entropy of the states a policy visits, used later in the lecture as a measure of how much coverage a policy achieves. - Empowerment: the mutual information between the next state and the current action,
I(s_{t+1}; a_t); maximizing it favors states with many available actions that reliably lead to many different, controllable future states.
Before you watch
- Watch the previous exploration lecture (Lecture 13), since this lecture explicitly contrasts reward-free exploration with the count- and information-gain-based bonus methods covered there.
- A basic grasp of probability distributions and KL divergence is assumed, though the lecture reviews entropy and mutual information from scratch.
Check your understanding
- Why might an agent benefit from acquiring skills before it knows what its future goals will be?
- How does mutual information relate to entropy, and why does that make it a natural measure of "informativeness"?
- Why does empowerment require both high entropy of the next state and low entropy of the next state given the action?
- What real-world intuition does the "standing in the middle of a room" example illustrate about empowerment?
Vocabulary
- reward-free (adjective)
- Having no external reward signal to guide learning.
Reward-free exploration lets the agent learn without a task reward. - repertoire (noun)
- A set of skills or behaviors an agent has available.
The agent builds up a broad repertoire of skills. - repurpose (verb)
- To use something originally built for one purpose to serve a new purpose.
Skills learned earlier can be repurposed for a new goal. - entropy (noun)
- A measure of how spread out or uncertain a distribution is.
A uniform distribution has maximal entropy. - mutual information (phrase)
- A measure of how much knowing one variable tells you about another.
Mutual information grows as two variables become more dependent. - KL divergence (noun)
- A measure of how different one probability distribution is from another.
Mutual information equals the KL divergence between joint and marginal distributions. - joint distribution (phrase)
- A probability distribution describing two variables together.
Mutual information compares the joint distribution to independent marginals. - marginal (noun)
- The probability distribution of just one variable, ignoring others.
Independence means the joint equals the product of marginals. - state marginal (phrase)
- The distribution over which states a policy tends to visit.
State marginal entropy measures how much coverage a policy has. - empowerment (noun)
- A measure of how much an agent's actions control its future states.
Empowerment favors states with many controllable outcomes. - reduction (in uncertainty) (noun)
- A decrease in how unsure we are about something.
Mutual information is the reduction in uncertainty from observing x. - reusable (adjective)
- Able to be used again for a different purpose later.
The agent aims to acquire reusable skills. - unsupervised (adjective)
- Learning without labeled examples or a given task.
The agent explores through unsupervised interaction. - at play (phrase)
- Engaged in playful, unstructured activity.
A child at play discovers new skills without a goal. - acquire (verb)
- To gain or obtain a new skill or piece of knowledge.
The agent should acquire diverse skills before its real task. - controllable (adjective)
- Able to be influenced or changed on purpose.
Empowerment favors states with many controllable futures. - coverage (noun)
- How much of the possible situations or states are included.
State marginal entropy is a measure of coverage. - informativeness (noun)
- How much useful information something provides.
Mutual information is a natural measure of informativeness. - reliably lead (phrase)
- To consistently result in a particular outcome.
Good states reliably lead to controllable future states. - building block (phrase)
- A basic component that later ideas are built from.
Entropy and mutual information are building blocks for this lecture.
Chapters
- 0:00 Introduction
- 0:50 Monday Recap
- 1:43 Exploration
- 3:35 Example
- 4:18 Outline
- 5:52 Useful Identities
- 6:45 Mutual Information
← Lecture 13, Part 6: Exploration · Lecture 14, Part 2: Learning Goal-Reaching Policies Without Rewards →
