Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 2 of 99 · 17:54
Lecture 1: Introduction, Part 2
Study guide
What this lecture covers
This part covers the CS285 course structure (model-free RL, model-based RL, exploration, offline RL, inverse RL, meta-learning) and grading, then gets into the substance: what reinforcement learning is and how it differs from supervised learning. It closes with a tour of applications, from robotic manipulation to language model fine-tuning.
After watching, you should be able to state the definition of a policy and its role, contrast the i.i.d. assumption of supervised learning with the sequential, reward-only feedback of RL, and describe the credit assignment problem.
Key ideas
- Course structure: units move from foundations to model-free algorithms (Q-learning, policy gradient, actor-critic), model-based methods, then advanced topics (exploration, offline RL, inverse RL, meta-learning). Grading is 50% homework, 40% project, 10% quizzes.
- Reinforcement learning is two things: a mathematical formalism for learning-based decision making, and an approach for learning decision making from experience; the lecture stresses not conflating the problem with any particular solution method.
- Supervised learning's assumptions: it treats data as independent and identically distributed (i.i.d.) and assumes every input has a known correct label — assumptions that don't hold for interactive decision-making tasks.
- The RL loop: an agent observes a state
s_t, chooses an actiona_t, and the environment returns a new states_{t+1}and a reward; the goal is a policyPi_thetathat maps states to actions and maximizes cumulative reward, not just immediate reward. - Credit assignment: when an outcome is delayed, it's often unclear which earlier action caused a success or failure, unlike labeled supervised data.
- Reward framing across domains: dog training, robot locomotion, and inventory management are all reframed as (actions, observations, reward) tuples to show how broadly the RL formalism applies.
- Applications shown: robotic hammering and jumping, Atari play discovering the "tunnel through the bricks" strategy, warehouse robots sorting trash, RL-controlled traffic flow reducing jams, and RL fine-tuning of language models (RLHF) and image generation models.
Walkthrough
Course Overview (0:00)
The instructor lists the course units and grading breakdown, describes the five homeworks (imitation learning, policy gradients/Q-learning, actor-critic, model-based RL, offline RL), and explains the research-level final project, including proposal and milestone checkpoints meant to give students feedback before the final report.
Reinforcement Learning (3:10)
The lecture defines reinforcement learning as both a formalism and a class of solution methods, then walks through the RL loop: an agent picks actions, the environment returns states and rewards, and the objective is a policy that maximizes cumulative reward over time rather than reward at a single step.
Supervised Learning (7:17)
By contrast, supervised learning is defined through its i.i.d. and fully-labeled-data assumptions. The lecture explains why these assumptions break down for sequential decision problems: past actions affect future inputs, and only outcomes (not correct actions) are observed, creating the credit assignment problem.
Examples (9:18)
The lecture works through examples — training a dog, controlling a robot, managing warehouse inventory — mapping each to actions, observations, and rewards, then shows applied results: robotic hammering and quadruped jumping, Atari game play, robotic trash sorting, RL-based traffic regulation, and RL used to fine-tune language models and image generation models toward human-preferred outputs.
Before you watch
- Watch Part 1 of Lecture 1 first for the motivating grasping example and historical background.
- Basic familiarity with the idea of training a model from a labeled dataset will help the supervised-vs-RL comparison land.
Check your understanding
- Why can't reinforcement learning rely on the i.i.d. assumption that supervised learning makes?
- What is the credit assignment problem, and why does it make RL harder than standard classification?
- In the traffic example, what did the RL-controlled car optimize for instead of its own speed?
- How was reinforcement learning used to improve the image generation example with the dolphin and bicycle?
- What are the five homework topics listed for this course?
Vocabulary
- logistics (noun)
- The practical details of organizing a course, like grading and schedule.
This part covers course logistics before the main content. - model-free (adjective)
- A type of learning that doesn't build an explicit model of the environment.
Model-free RL includes Q-learning and policy gradient methods. - model-based (adjective)
- A type of learning that builds a model to predict how the environment behaves.
Model-based RL uses predictions to plan actions. - exploration (noun)
- Trying new actions to discover more about an environment.
Exploration helps an agent find better strategies. - offline RL (noun)
- Learning a policy from a fixed dataset without further interaction.
Offline RL doesn't require live trial and error. - inverse RL (noun)
- Learning a reward function by observing someone's behavior.
Inverse RL infers rewards from expert demonstrations. - meta-learning (noun)
- Learning how to learn new tasks more quickly.
Meta-learning helps an agent adapt to new tasks fast. - formalism (noun)
- A precise, structured way of describing a problem mathematically.
RL is both a formalism and a set of solution methods. - conflate (verb)
- To wrongly treat two different things as if they were the same.
Don't conflate the RL problem with any one solution method. - i.i.d. (adjective)
- Short for independent and identically distributed, meaning data points don't affect each other.
Supervised learning assumes i.i.d. data. - policy (noun)
- A rule that maps a situation to an action for an agent to take.
The goal of RL is to find a good policy. - cumulative reward (noun)
- The total reward added up over many steps, not just one.
The policy is trained to maximize cumulative reward. - credit assignment (noun)
- The problem of figuring out which past action caused a later outcome.
Credit assignment is hard when rewards are delayed. - locomotion (noun)
- The ability to move from one place to another.
Robot locomotion is a common RL application. - inventory management (noun)
- Deciding how much stock to keep and order over time.
Inventory management can be framed as an RL problem. - fine-tuning (noun)
- Further training an already-trained model on a more specific task.
RLHF is used for fine-tuning language models. - grading (noun)
- The way scores are assigned in a course.
Grading is split between homework, project, and quizzes. - proposal (noun)
- A written plan describing what a project intends to do.
Students submit a proposal before starting the final project. - milestone (noun)
- A checkpoint marking progress partway through a longer project.
The course includes milestone checkpoints for the final project. - regulate (verb)
- To control something so it stays within a desired range.
RL is used to regulate traffic flow.
Chapters
← Lecture 1: Introduction, Part 1 · Lecture 1: Introduction, Part 3 →
