Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 8 of 99 · 8:51

Lecture 2: Imitation Learning, Part 5

CS 285: Lecture 2, Imitation Learning. Part 5 on YouTube

Study guide

What this lecture covers

This closing part of the imitation learning lecture presents DAgger (Dataset Aggregation), an algorithm that fixes distributional shift by relabeling states the policy actually visits, then explains why imitation learning by itself is insufficient, motivating the shift to reward-based reinforcement learning.

After watching, you should be able to describe the DAgger loop, explain why it provably bounds error linearly rather than quadratically in trajectory length, and state why RL avoids imitation learning's dependence on costly human-labeled data.

Key ideas

  • DAgger's core idea: instead of changing the policy to fit the data, change the data to match the policy's own state distribution by running the current policy, collecting the states it visits, and asking a human to label each with the correct action.
  • The DAgger loop: train on the initial demonstrations, run the policy, label the new observations, aggregate this data with the original set, and retrain — repeating until the training distribution converges toward the policy's rollout distribution.
  • Theoretical guarantee: unlike naive behavioral cloning's quadratic-in-T error bound, DAgger provably achieves a bound linear in T, at the cost of requiring ongoing human labeling.
  • A practical limitation: labeling states offline, out of the flow of actually performing the task, can be unnatural for humans (e.g. labeling a driving image without the real-time context of reacting while driving).
  • Why imitation learning alone isn't enough: it requires plentiful human-provided data, which is expensive and limits scale; humans are also poor at specifying some kinds of low-level actions (e.g. rotor commands for aerobatics); and autonomous data collection allows machines to self-improve beyond human performance.
  • Toward reward-based objectives: the lecture reframes the same imitation-learning cost function in terms of a reward function, setting up the transition to defining more general reward functions for reinforcement learning.

Before you watch

  • Watch Lecture 2 Parts 1 through 4 first; DAgger is presented as a direct fix to the compounding-error problem developed there.

Check your understanding

  1. What does DAgger change about the data collection process compared to standard behavioral cloning?
  2. Why does DAgger achieve a linear rather than quadratic error bound?
  3. What practical difficulty can arise when asking humans to label states offline in step three of DAgger?
  4. Why might autonomous data collection be preferable to human demonstrations for some robotic tasks?

Vocabulary

DAgger (noun)
An algorithm (Dataset Aggregation) that fixes distributional shift by labeling states the trained policy actually visits.
DAgger asks a human to label new states the policy visits.
dataset aggregation (noun)
Combining new labeled data with an existing dataset over multiple rounds.
DAgger stands for dataset aggregation.
aggregate (verb)
To combine multiple pieces of data into one larger set.
New labeled data is aggregated with the original dataset.
provably (adverb)
In a way that can be shown to be true using mathematical proof.
DAgger provably achieves a linear error bound.
guarantee (noun)
A mathematically proven promise about how a method will behave.
The theoretical guarantee bounds the error linearly.
offline (adjective)
Done separately, not while actually performing the live task.
Labeling states offline can feel unnatural for humans.
real-time (adjective)
Happening immediately, as events actually occur.
Driving decisions are normally made in real time.
aerobatics (noun)
Complex, highly skilled flying maneuvers.
Humans struggle to specify exact commands for aerobatics.
autonomous (adjective)
Operating independently, without human control.
Autonomous data collection lets a robot improve on its own.
scale (verb)
To grow larger while remaining practical and effective.
Human-labeled data is expensive and hard to scale.
reward function (noun)
A function that scores how good or bad an outcome is.
The lecture reframes the cost function as a reward function.
converge (verb)
To gradually approach a stable final result.
The training distribution converges toward the rollout distribution.
distributional shift (noun)
A mismatch between the data a model was trained on and the data it later sees.
DAgger is designed to fix distributional shift directly.
compounding error (noun)
A mistake that grows worse over time because it leads to further mistakes.
DAgger reduces compounding error compared to plain behavioral cloning.
insufficient (adjective)
Not enough to fully solve a problem on its own.
Imitation learning alone is insufficient for many tasks.
plentiful (adjective)
Available in large amounts.
Imitation learning needs plentiful human-provided data.
specify (verb)
To state something in a precise, detailed way.
Humans struggle to specify exact low-level rotor commands.
low-level (adjective)
Close to the raw mechanics of a system, rather than a high-level goal.
Rotor commands are a low-level type of action.
objective (noun)
The goal that a training process tries to achieve.
The lecture moves toward a reward-based objective.
unnatural (adjective)
Not matching how something would normally happen.
Labeling states offline can feel unnatural for humans.

Chapters

← Lecture 2: Imitation Learning, Part 4 · Lecture 4: Introduction to RL Algorithms, Part 1 →