Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 94 of 99 · 7:38
Lecture 22, Part 4: Gradient-Based Meta-Reinforcement Learning
Study guide
What this lecture covers
This segment presents gradient-based meta-learning, best known as model-agnostic meta-learning (MAML), as an alternative to the RNN-based approach from the previous part. It reframes pre-training and fine-tuning as a form of meta-learning and asks whether the fine-tuning process itself can be optimized directly.
The core idea is to make F_theta, the function that adapts to a new MDP, literally be a gradient update: starting from initial parameters theta, one policy-gradient step on a task's objective produces the adapted parameters phi_i. Meta-training then searches for an initialization theta such that a single gradient step, for any task drawn from the distribution, leads to high reward. The lecture connects this to supervised MAML, discusses why it works despite looking deceptively simple, and shows an ant robot example adapting to run in different directions after one gradient step.
Key ideas
- Pre-training and fine-tuning as meta-learning: gradient-based meta-learning treats the fine-tuning step itself as the thing being optimized, rather than just the initial features.
- F_theta as a gradient step: the adaptation function is defined as
theta + gradient of J_i(theta), i.e. one (or more) policy-gradient updates evaluated on the specific task's objective. - Second-order optimization: meta-training searches for an initial
thetasuch that applying a gradient step on any meta-training task increases that task's reward as much as possible, which is a form of second-order optimization. - Same computation graph as an RNN meta-learner: gradient descent inside
F_thetacan be treated as just another differentiable architecture, implementable with automatic differentiation, though policy-gradient second derivatives require extra care. - Favorable inductive bias: because the adaptation mechanism is a real gradient step, models trained this way can often take more gradient steps at test time than they were meta-trained with, unlike RNN-based meta-learners whose adaptation is a fixed forward pass.
- Multiple gradient steps are possible: the method is not limited to one gradient step, though the math for multiple steps and for policy gradients specifically is more involved.
Before you watch
- Watch Part 3 of this lecture, which introduces the general meta-RL framework and the RNN-based alternative that this segment contrasts with.
- Review the pre-training and fine-tuning discussion from Part 1, since this lecture directly builds on that idea.
- Familiarity with policy gradient methods is needed to follow how
F_thetais defined as a gradient update.
Check your understanding
- How does model-agnostic meta-learning (MAML) redefine
F_thetacompared to the RNN-based meta-learner from the previous part? - Why is meta-training in MAML described as a second-order optimization problem?
- What practical advantage do gradient-based meta-learners have over RNN-based ones at meta-test time?
- In the ant robot example, what happens to the meta-trained policy before any adaptation, and how does one gradient step change its behavior?
Vocabulary
- model-agnostic (adjective)
- Not depending on any specific type of model; working with many kinds.
Model-agnostic meta-learning works with any differentiable model. - initialization (noun)
- The starting values of a model's parameters before training.
MAML searches for a good initialization for fast adaptation. - gradient step (noun)
- A single update to parameters made by moving them opposite the gradient.
One gradient step adapts the policy to a new task. - second-order (adjective)
- Involving derivatives of derivatives, one level deeper than usual.
Meta-training in MAML is a second-order optimization problem. - automatic differentiation (noun)
- A technique that lets software compute exact gradients through a sequence of operations.
Automatic differentiation makes the gradient step differentiable too. - inductive bias (noun)
- The built-in assumptions a learning method uses to generalize from limited data.
The gradient-based method has a favorable inductive bias. - extrapolate (verb)
- To go beyond what was directly trained on, in a related direction.
The model can extrapolate to more gradient steps at test time. - meta-learning (noun)
- Learning how to learn quickly, so a model adapts fast to new tasks.
Gradient-based meta-learning is a form of meta-learning for RL. - meta-training (noun)
- The training process that produces a model good at adapting to new tasks.
Meta-training searches for an initialization that adapts well. - adapt (verb)
- To change so as to perform well in a new situation.
The policy adapts to a new task after one gradient step. - recurrent neural network (noun)
- A network that processes sequences by keeping an internal state over time.
The previous part used a recurrent neural network for adaptation. - differentiable (adjective)
- Able to have a gradient computed through it.
The gradient step inside F_theta is treated as a differentiable operation. - computation graph (noun)
- A diagram showing the order of operations used to compute a result.
MAML uses the same kind of computation graph as an RNN meta-learner. - forward pass (noun)
- Running inputs through a model from start to end to get an output.
An RNN meta-learner's adaptation is a fixed forward pass. - deceptively (adverb)
- In a way that seems simpler or different from how it really is.
MAML looks deceptively simple despite its second-order optimization. - distribution (noun)
- A set of possibilities together with how likely each one is.
Meta-training draws tasks from a distribution of MDPs. - alternative (noun)
- A different option that could be chosen instead.
MAML is presented as an alternative to the RNN-based approach. - favorable (adjective)
- Giving an advantage; working in one's favor.
The gradient-based method has a favorable inductive bias. - meta-test (noun)
- The stage where a meta-trained model is applied to a brand-new task.
The model can take more gradient steps at meta-test time. - evaluate (verb)
- To compute or judge the value of something.
The gradient step is evaluated on the task's own objective. - reframe (verb)
- To present an idea in a new way that changes how it is understood.
The lecture reframes fine-tuning as a form of meta-learning. - contrast (verb)
- To compare two things by pointing out their differences.
MAML is contrasted with the RNN-based meta-learner. - architecture (noun)
- The overall structure and design of a neural network.
Gradient descent can be treated as just another differentiable architecture. - derivative (noun)
- A measure of how a function's output changes as its input changes.
Policy-gradient second derivatives require extra care to compute. - segment (noun)
- One distinct part of a longer lecture or process.
This segment covers gradient-based meta-reinforcement learning.
Chapters
- 0:00 Intro
- 0:51 Model MetaLearning
- 3:00 Policy Gradients
- 4:35 Supervised MetaLearning
- 6:35 MetaLearning in Practice
- 7:08 MetaLearning References
← Lecture 22, Part 3: Meta Reinforcement Learning with RNNs · Lecture 22, Part 5: Meta-RL as Partially Observed MDPs →
