Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 5 of 99 · 23:06

Lecture 2: Imitation Learning, Part 2

CS 285: Lecture 2, Imitation Learning. Part 2 on YouTube

Study guide

What this lecture covers

This part gives the formal explanation for why behavioral cloning is not guaranteed to work, building on the informal argument from Part 1. It introduces distributional shift, defines a cost function counting mistakes, and derives a bound on the expected number of mistakes a cloned policy makes over a trajectory.

After watching, you should be able to explain distributional shift between the training and test distributions of a policy, describe the "tightrope walker" worst-case example, and state the resulting error bound and why it matters.

Key ideas

  • Distributional shift: a policy is trained under the demonstrator's observation distribution p_data(o_t) but evaluated under its own distribution p_{pi_theta}(o_t), and these differ because the policy doesn't drive exactly like the demonstrator.
  • A cost function for mistakes: define cost as 0 if the policy matches the expert's deterministic action and 1 otherwise; the goal is to minimize expected cost under the policy's own state distribution, not the training distribution.
  • The tightrope walker example: a worst-case scenario where any mistake puts the agent in a state the expert never visited, so all subsequent time steps also become mistakes; this produces an expected error on the order of epsilon * T^2, quadratic in trajectory length T.
  • General bound: even without assuming the pathological tightrope setup, the expected number of mistakes for behavioral cloning is bounded by roughly epsilon * T^2 in the worst case, derived via total variation divergence between the training and rollout state distributions.
  • Why quadratic growth is bad: a linear growth in errors with trajectory length would be expected and tolerable, but quadratic growth means long-horizon tasks accumulate error much faster, explaining why naive behavioral cloning degrades badly on longer trajectories.
  • Why practice is less pessimistic: the tightrope example is pathological because a single mistake is unrecoverable; real tasks and datasets that include recovery behavior (like the earlier left/right camera trick) let policies learn to correct small errors, keeping them closer to linear error growth.
  • A paradox: training data that includes mistakes and their recovery can make behavioral cloning more robust than "perfect" demonstration data, because it broadens the training distribution to cover near-mistake states.

Before you watch

  • Watch Lecture 2 Part 1 first; this part relies directly on its notation (states, observations, cost) and its informal compounding-error argument.
  • Comfort with basic probability notation (expectations, conditional distributions) will help with the derivation.

Check your understanding

  1. What is distributional shift, and why does behavioral cloning suffer from it while standard supervised learning does not?
  2. In the tightrope walker example, why does a single mistake lead to roughly T additional mistakes?
  3. What is the overall worst-case bound on expected mistakes for behavioral cloning, and why is quadratic growth in T a problem?
  4. Why might training data containing mistakes and recoveries actually improve a cloned policy compared to "perfect" demonstrations?

Vocabulary

formal (adjective)
Based on precise mathematical reasoning rather than intuition alone.
This lecture gives a formal analysis of behavioral cloning.
distributional shift (noun)
A mismatch between the data a model was trained on and the data it sees later.
Distributional shift causes behavioral cloning to fail.
trajectory (noun)
The sequence of states and actions a policy produces over time.
The error bound depends on the length of the trajectory.
cost function (noun)
A function that measures how bad an outcome or action is.
The cost function counts mistakes made by the policy.
rollout (noun)
Running a policy forward to see what states and actions it produces.
The policy's rollout distribution differs from the training data.
worst-case (adjective)
Describing the most extreme, least favorable possible outcome.
The tightrope walker is a worst-case example.
unrecoverable (adjective)
Impossible to fix or return from once it happens.
A single mistake on the tightrope is unrecoverable.
quadratic (adjective)
Growing in proportion to the square of a quantity.
The error bound grows quadratically with trajectory length.
linear growth (noun)
Increasing at a steady, constant rate.
Linear growth in errors would be much less severe.
total variation divergence (noun)
A mathematical measure of how different two probability distributions are.
The bound is derived using total variation divergence.
pathological (adjective)
Extreme or unusual in a way that causes serious problems.
The tightrope example is a pathological worst case.
recovery behavior (noun)
Actions that bring a system back from a mistake toward the correct path.
Datasets with recovery behavior reduce compounding error.
paradox (noun)
A surprising result that seems to contradict common sense.
It's a paradox that imperfect data can help more than perfect data.
horizon (noun)
The length of time or number of steps a task runs over.
Long-horizon tasks accumulate more error.
bound (noun)
A mathematical limit on how large or small a value can be.
The lecture derives a bound on the expected number of mistakes.
derive (verb)
To work out a result step by step from known rules.
The bound is derived using total variation divergence.
expected value (noun)
The average outcome you would expect over many repeated trials.
The expected number of mistakes is bounded by epsilon times T squared.
degrade (verb)
To become worse in quality or performance.
Naive behavioral cloning degrades badly on longer trajectories.
tolerable (adjective)
Acceptable, even if not ideal.
Linear error growth would be tolerable in practice.
assume (verb)
To take something as true without direct proof.
The bound doesn't need to assume the pathological tightrope setup.

← Lecture 2: Imitation Learning, Part 1 · Lecture 2: Imitation Learning, Part 3 →