Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 92 of 99 · 9:48

Lecture 22, Part 2: What Is Meta-Learning?

CS 285: Lecture 22, Part 2: Transfer Learning & Meta-Learning on YouTube

Study guide

What this lecture covers

This short segment opens the meta-learning portion of Berkeley CS285's transfer learning unit. It frames meta-learning as a logical extension of multitask learning: instead of just solving many tasks, the goal is to use those tasks to learn how to learn new tasks more quickly.

The lecture first motivates why meta-learning matters for deep RL, given how sample-hungry model-free methods are, then works through a concrete supervised-learning example (few-shot image classification) to demystify what "learning to learn" actually means mechanically, before RL versions are introduced in later parts.

Key ideas

  • Meta-learning as learning to learn: after training on many tasks, the aim is to generalize the learning process itself, not just the solutions to those tasks.
  • Why it helps RL: a meta-learned RL method can explore more intelligently, avoid actions known to be useless, and acquire useful features faster than learning a new task from scratch.
  • Meta-training vs meta-testing: meta-training uses a set of training/test set pairs across many tasks (the source domains); meta-testing applies the learned procedure to a genuinely new task (the target domain).
  • Meta-learning as a function: standard supervised learning maps an input x to a prediction y; meta-learning maps an entire training set plus a test input to a prediction, f(D_train, x_test) = y_test.
  • RNN-based meta-learner example: a recurrent network reads in a sequence of training examples, produces a hidden state summarizing the task, and a small classifier uses that hidden state plus a test input to predict the label.
  • Two levels of optimization: generic learning finds parameters that minimize training loss; meta-learning trains a function f_theta so that the parameters it produces perform well on held-out test data across many tasks.

Before you watch

  • Watch Part 1 of this lecture first, since it defines source/target domains and transfer terminology that this segment builds on.
  • Basic familiarity with recurrent neural networks is useful for following the RNN meta-learner example.

Check your understanding

  1. How does meta-learning differ from ordinary multitask learning?
  2. In the RNN meta-learner example, what do the terms theta and phi each represent?
  3. Why might a meta-learned RL policy explore more efficiently than one trained without meta-learning?

Vocabulary

meta-learning (noun)
Training a system to learn new tasks quickly, by practicing on many tasks first.
Meta-learning teaches the model how to learn, not just one task.
learning to learn (phrase)
The idea of improving the process of learning itself, not just one skill.
Meta-learning is often described as learning to learn.
multitask learning (noun)
Training a single model to perform well on several different tasks at once.
Meta-learning builds on multitask learning.
meta-training (noun)
The training phase where a model practices across many different tasks.
Meta-training uses training and test pairs from many tasks.
meta-testing (noun)
Applying the learned learning procedure to a brand-new task.
Meta-testing checks performance on an unseen task.
few-shot (adjective)
Learning a new task from only a small number of examples.
Few-shot image classification is a common meta-learning example.
recurrent network (noun)
A neural network that processes sequences by keeping an internal memory over time.
A recurrent network reads training examples one at a time.
hidden state (noun)
The internal memory a recurrent network keeps and updates as it processes a sequence.
The hidden state summarizes what the network has seen so far.
source domain (noun)
The set of tasks or data used during training, before facing a new task.
Meta-training uses many tasks from the source domain.
target domain (noun)
The new task or setting a trained model is finally applied to.
Meta-testing checks performance on the target domain.
generalize (verb)
To perform well on new cases beyond what was directly trained on.
Meta-learning tries to generalize the learning process itself.
sample-hungry (adjective)
Needing a large amount of data or experience to learn well.
Model-free RL methods are notoriously sample-hungry.
demystify (verb)
To make something confusing easier to understand.
The supervised example helps demystify what meta-learning means.
acquire (verb)
To gain or obtain something, especially a skill or piece of information.
A meta-learned method can acquire useful features faster.
classifier (noun)
A model that assigns an input to one of several categories.
A small classifier uses the hidden state to predict the label.
held-out (adjective)
Kept separate from training so it can fairly test performance.
Meta-learning trains parameters that perform well on held-out data.
generic (adjective)
General, not specific to one particular case.
Generic learning just minimizes training loss on one task.
summarize (verb)
To give a short version that captures the main information.
The hidden state summarizes the task from the training examples.
intelligently (adverb)
In a smart, well-reasoned way.
A meta-learned agent can explore more intelligently.
segment (noun)
One distinct part of a longer lecture or unit.
This segment opens the meta-learning portion of the course.
useless (adjective)
Having no value or benefit.
A meta-learned policy can avoid actions known to be useless.
logical extension (phrase)
A natural next step that follows sensibly from an existing idea.
Meta-learning is framed as a logical extension of multitask learning.
concrete (adjective)
Specific and real, rather than abstract.
The lecture uses a concrete supervised-learning example.
procedure (noun)
A fixed series of steps for carrying out a task.
Meta-testing applies the learned procedure to a new task.
mechanically (adverb)
In terms of the actual step-by-step machinery, not just the abstract idea.
The example shows what learning to learn means mechanically.

Chapters

← Lecture 22, Part 1: Transfer Learning and Domain Adaptation · Lecture 22, Part 3: Meta Reinforcement Learning with RNNs →