Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 1 of 99 · 10:10
Lecture 1: Introduction, Part 1
Study guide
What this lecture covers
This opening part of the first CS285 lecture asks why reinforcement learning (RL) is needed at all, using robotic grasping as a motivating example. It contrasts RL with the supervised, density-estimation approach behind today's image and language generation models, then traces RL's roots back to animal behavior research and control/optimization.
After watching, you should be able to explain why supervised learning struggles with tasks like robotic grasping, describe the difference between learning to imitate data and learning to maximize a reward, and name the two historical lineages that shaped modern deep RL.
Key ideas
- Grasping as a motivating problem: picking up objects has many special cases (rigid vs. deformable, center of mass) that are hard to hand-engineer, and even humans can't reliably label "correct" grasp locations, making standard supervised learning a poor fit.
- Reward instead of labels: in RL, robots collect their own trial data labeled only with outcomes (success/failure), and a reward function scores these outcomes instead of providing ground-truth answers.
- Generative models as density estimation: recent AI advances in image and text generation are framed as estimating
P(X)orP(Y|X)from large datasets, learning to reproduce the distribution of human-generated content. - Two lineages of RL: one traces to psychology and animal behavior studies (B.F. Skinner), the other to control, optimization and evolutionary algorithms (e.g., Karl Sims's simulated creatures).
- Deep RL as a combination: modern deep reinforcement learning merges classical RL's algorithmic ideas with large-scale optimization and neural network function approximation.
- Emergent behavior: results like AlphaGo's "move 37" are notable because they were not copied from human play; RL can discover solutions a person would not have chosen, which matters for building systems that respond intelligently to novel situations.
Before you watch
- No prior lecture is required; this is the first video of the course.
- Basic familiarity with supervised learning (inputs, outputs, training on labeled data) helps you follow the contrast being drawn.
Check your understanding
- Why is robotic grasping difficult to frame as a standard supervised learning problem?
- What does a reward function provide that a labeled dataset does not?
- What are the two historical disciplines that influenced modern reinforcement learning?
- Why does the lecture consider AlphaGo's "move 37" significant?
Vocabulary
- reinforcement learning (noun)
- A way of training a system to make decisions by rewarding good outcomes.
Reinforcement learning lets a robot learn from trial and error. - motivating (adjective)
- Giving a reason or example that makes a topic interesting.
Robotic grasping is a motivating example for this lecture. - supervised learning (noun)
- Training a model using data that comes with correct answers.
Supervised learning struggles with tasks lacking clear correct labels. - density estimation (noun)
- Learning the underlying pattern of how data is spread out.
Image generation is framed as density estimation. - hand-engineer (verb)
- To design a solution manually, step by step, rather than learning it automatically.
Grasp rules are hard to hand-engineer for every object. - ground-truth (adjective)
- Describing data known to be correct, used as the standard for comparison.
RL doesn't rely on ground-truth answers, only rewards. - reward function (noun)
- A function that scores how good or bad an outcome is.
A reward function replaces labeled data in RL. - trial data (noun)
- Data collected by trying actions and observing what happens.
Robots collect their own trial data during learning. - lineage (noun)
- A line of historical development leading to something.
RL has two separate historical lineages. - control (noun)
- The field of engineering that studies how to steer systems toward a goal.
One lineage of RL comes from control and optimization. - evolutionary algorithm (noun)
- An optimization method inspired by natural selection.
Evolutionary algorithms shaped simulated creatures' behavior. - function approximation (noun)
- Using a model, like a neural network, to estimate a complex function.
Deep RL combines classical algorithms with function approximation. - emergent behavior (noun)
- A new pattern of behavior that arises without being directly programmed.
AlphaGo's move showed emergent behavior not copied from humans. - novel (adjective)
- New and not seen before.
RL can respond intelligently to novel situations. - deformable (adjective)
- Able to change shape when handled.
Deformable objects are harder to grasp than rigid ones. - rigid (adjective)
- Stiff and not able to bend or change shape.
Rigid objects keep the same shape when grasped. - center of mass (noun)
- The point where an object's weight is balanced.
An object's center of mass affects how it should be grasped. - outcome (noun)
- The result of an action or process.
RL labels trial data only with its outcome, success or failure. - psychology (noun)
- The scientific study of the mind and behavior.
One lineage of RL traces back to psychology. - simulated (adjective)
- Created artificially by a computer model rather than existing in reality.
Karl Sims trained simulated creatures using evolutionary methods.
