Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 4 of 99 · 24:47

Lecture 2: Imitation Learning, Part 1

CS 285: Lecture 2, Imitation Learning. Part 1 on YouTube

Study guide

What this lecture covers

This lecture introduces the notation used throughout the course (policies, observations, actions, states) and defines behavioral cloning: training a policy with supervised learning to copy a human demonstrator. It then explains, informally, why behavioral cloning can fail and what practical tricks make it work anyway.

After watching, you should be able to distinguish observations from states, explain the Markov property, describe behavioral cloning, and explain why compounding errors are a problem in sequential decision making but not in standard supervised learning.

Key ideas

  • Policy notation: a policy Pi_theta(a|o) maps observations o to a distribution over actions a, with parameters Theta (e.g. neural network weights); it can be deterministic or stochastic.
  • States vs. observations: a state s is a complete physical description of the world sufficient to predict the future; an observation o (e.g. camera pixels) may not contain enough information to recover the state.
  • The Markov property: given the current state, the future is independent of the past — this is what defines a state, and some RL algorithms require policies conditioned on states rather than partial observations.
  • Notation conventions: the course uses s/a from Bellman-era dynamic programming; controls literature often uses x/u for the same concepts.
  • Behavioral cloning: collect (observation, action) pairs from a human demonstrator (e.g. dashboard images and steering angles) and train a policy with supervised learning to reproduce the demonstrator's actions.
  • Why it can fail: behavioral cloning violates the i.i.d. assumption of supervised learning, since actions affect future observations. Small errors push the policy into states unlike those it was trained on, where it is more likely to make bigger errors — a compounding error problem.
  • NVIDIA's fix: adding left and right dashboard cameras labeled with corrective steering commands teaches the policy how to recover from small deviations, effectively broadening the training distribution.

Walkthrough

Terminology & notation (0:00)

The lecture defines policies, observations, actions, discrete time steps, and the difference between deterministic and stochastic policies, drawing the analogy to image classifiers.

States, observations and the Markov property (4:04)

Using a cheetah-chasing-a-gazelle example, the lecture distinguishes state from observation, introduces probabilistic graphical models to describe the policy/transition relationship, and explains the Markov property as the defining feature of a state.

Imitation learning and behavioral cloning (13:05)

The lecture switches to an autonomous driving running example: images from a dashboard camera as observations, steering commands as actions. Training a neural network on this data via supervised learning is called behavioral cloning, tracing back to the 1989 ALVINN system.

Why behavioral cloning can fail, and the NVIDIA fix (16:05)

Using a state-vs-time diagram, the lecture shows how small errors compound because sequential data violates the i.i.d. assumption. It then walks through NVIDIA's behavioral cloning driving system, which mitigates the problem with left/right camera images labeled with corrective steering actions.

Before you watch

  • Watch the three parts of Lecture 1 first for the course's motivation and definition of reinforcement learning.
  • Familiarity with supervised learning (training a classifier on labeled data) helps, since the lecture repeatedly contrasts it with policy learning.

Check your understanding

  1. What is the difference between a state and an observation, and why does it matter for the Markov property?
  2. Why does behavioral cloning violate the i.i.d. assumption that standard supervised learning relies on?
  3. How does adding left and right camera images with corrective labels help NVIDIA's driving system avoid compounding errors?
  4. What does it mean for a policy to be conditioned on state versus conditioned on observation?

Vocabulary

notation (noun)
A system of symbols used to represent ideas precisely.
This lecture introduces the course's notation for policies.
policy (noun)
A rule that maps an observation or state to an action.
The policy decides what action to take at each step.
distribution (noun)
A description of how likely each possible outcome is.
The policy outputs a distribution over actions.
deterministic (adjective)
Always producing the same output for the same input, with no randomness.
A deterministic policy always picks the same action.
stochastic (adjective)
Involving randomness in its outcome.
A stochastic policy samples actions from a distribution.
state (noun)
A complete description of the world at one moment, enough to predict the future.
The state contains everything needed to predict what happens next.
observation (noun)
Partial information about the world, which may not fully describe the state.
A camera image is an observation, not the full state.
Markov property (noun)
The idea that the future depends only on the current state, not on the past.
The Markov property defines what counts as a true state.
dynamic programming (noun)
A method for solving problems by breaking them into smaller, related subproblems.
The notation comes from Bellman-era dynamic programming.
behavioral cloning (noun)
Training a policy with supervised learning to copy a human's recorded actions.
Behavioral cloning trains a driving policy from human demonstrations.
demonstrator (noun)
A person or agent whose actions are recorded to teach a policy.
The human demonstrator provides steering examples.
i.i.d. (adjective)
Short for independent and identically distributed, meaning data points don't affect each other.
Supervised learning assumes i.i.d. training data.
compounding error (noun)
A mistake that grows worse over time as it leads to further mistakes.
Compounding error is the main problem with naive behavioral cloning.
deviation (noun)
A movement away from the expected or correct path.
Corrective data helps recover from small deviations.
corrective (adjective)
Designed to fix or reverse a mistake.
Corrective steering commands teach the policy to recover.
probabilistic graphical model (noun)
A diagram showing how random variables depend on each other.
A probabilistic graphical model shows the policy-transition relationship.
analogy (noun)
A comparison between two similar things used to explain an idea.
The lecture draws an analogy to image classifiers.
dashboard camera (noun)
A camera mounted inside a vehicle facing forward.
The dashboard camera provides observations for driving.
steering command (noun)
An instruction that controls how a vehicle turns.
The policy predicts a steering command from the camera image.
mitigate (verb)
To make a problem less severe.
Extra cameras mitigate the compounding-error problem.

← Lecture 1: Introduction, Part 3 · Lecture 2: Imitation Learning, Part 2 →