Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 1 of 99 · 10:10

Lecture 1: Introduction, Part 1

CS 285: Lecture 1, Introduction. Part 1 on YouTube

Study guide

What this lecture covers

This opening part of the first CS285 lecture asks why reinforcement learning (RL) is needed at all, using robotic grasping as a motivating example. It contrasts RL with the supervised, density-estimation approach behind today's image and language generation models, then traces RL's roots back to animal behavior research and control/optimization.

After watching, you should be able to explain why supervised learning struggles with tasks like robotic grasping, describe the difference between learning to imitate data and learning to maximize a reward, and name the two historical lineages that shaped modern deep RL.

Key ideas

  • Grasping as a motivating problem: picking up objects has many special cases (rigid vs. deformable, center of mass) that are hard to hand-engineer, and even humans can't reliably label "correct" grasp locations, making standard supervised learning a poor fit.
  • Reward instead of labels: in RL, robots collect their own trial data labeled only with outcomes (success/failure), and a reward function scores these outcomes instead of providing ground-truth answers.
  • Generative models as density estimation: recent AI advances in image and text generation are framed as estimating P(X) or P(Y|X) from large datasets, learning to reproduce the distribution of human-generated content.
  • Two lineages of RL: one traces to psychology and animal behavior studies (B.F. Skinner), the other to control, optimization and evolutionary algorithms (e.g., Karl Sims's simulated creatures).
  • Deep RL as a combination: modern deep reinforcement learning merges classical RL's algorithmic ideas with large-scale optimization and neural network function approximation.
  • Emergent behavior: results like AlphaGo's "move 37" are notable because they were not copied from human play; RL can discover solutions a person would not have chosen, which matters for building systems that respond intelligently to novel situations.

Before you watch

  • No prior lecture is required; this is the first video of the course.
  • Basic familiarity with supervised learning (inputs, outputs, training on labeled data) helps you follow the contrast being drawn.

Check your understanding

  1. Why is robotic grasping difficult to frame as a standard supervised learning problem?
  2. What does a reward function provide that a labeled dataset does not?
  3. What are the two historical disciplines that influenced modern reinforcement learning?
  4. Why does the lecture consider AlphaGo's "move 37" significant?

Vocabulary

reinforcement learning (noun)
A way of training a system to make decisions by rewarding good outcomes.
Reinforcement learning lets a robot learn from trial and error.
motivating (adjective)
Giving a reason or example that makes a topic interesting.
Robotic grasping is a motivating example for this lecture.
supervised learning (noun)
Training a model using data that comes with correct answers.
Supervised learning struggles with tasks lacking clear correct labels.
density estimation (noun)
Learning the underlying pattern of how data is spread out.
Image generation is framed as density estimation.
hand-engineer (verb)
To design a solution manually, step by step, rather than learning it automatically.
Grasp rules are hard to hand-engineer for every object.
ground-truth (adjective)
Describing data known to be correct, used as the standard for comparison.
RL doesn't rely on ground-truth answers, only rewards.
reward function (noun)
A function that scores how good or bad an outcome is.
A reward function replaces labeled data in RL.
trial data (noun)
Data collected by trying actions and observing what happens.
Robots collect their own trial data during learning.
lineage (noun)
A line of historical development leading to something.
RL has two separate historical lineages.
control (noun)
The field of engineering that studies how to steer systems toward a goal.
One lineage of RL comes from control and optimization.
evolutionary algorithm (noun)
An optimization method inspired by natural selection.
Evolutionary algorithms shaped simulated creatures' behavior.
function approximation (noun)
Using a model, like a neural network, to estimate a complex function.
Deep RL combines classical algorithms with function approximation.
emergent behavior (noun)
A new pattern of behavior that arises without being directly programmed.
AlphaGo's move showed emergent behavior not copied from humans.
novel (adjective)
New and not seen before.
RL can respond intelligently to novel situations.
deformable (adjective)
Able to change shape when handled.
Deformable objects are harder to grasp than rigid ones.
rigid (adjective)
Stiff and not able to bend or change shape.
Rigid objects keep the same shape when grasped.
center of mass (noun)
The point where an object's weight is balanced.
An object's center of mass affects how it should be grasped.
outcome (noun)
The result of an action or process.
RL labels trial data only with its outcome, success or failure.
psychology (noun)
The scientific study of the mind and behavior.
One lineage of RL traces back to psychology.
simulated (adjective)
Created artificially by a computer model rather than existing in reality.
Karl Sims trained simulated creatures using evolutionary methods.

Lecture 1: Introduction, Part 2 →