Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 2 of 99 · 17:54

Lecture 1: Introduction, Part 2

CS 285: Lecture 1, Introduction. Part 2 on YouTube

Study guide

What this lecture covers

This part covers the CS285 course structure (model-free RL, model-based RL, exploration, offline RL, inverse RL, meta-learning) and grading, then gets into the substance: what reinforcement learning is and how it differs from supervised learning. It closes with a tour of applications, from robotic manipulation to language model fine-tuning.

After watching, you should be able to state the definition of a policy and its role, contrast the i.i.d. assumption of supervised learning with the sequential, reward-only feedback of RL, and describe the credit assignment problem.

Key ideas

  • Course structure: units move from foundations to model-free algorithms (Q-learning, policy gradient, actor-critic), model-based methods, then advanced topics (exploration, offline RL, inverse RL, meta-learning). Grading is 50% homework, 40% project, 10% quizzes.
  • Reinforcement learning is two things: a mathematical formalism for learning-based decision making, and an approach for learning decision making from experience; the lecture stresses not conflating the problem with any particular solution method.
  • Supervised learning's assumptions: it treats data as independent and identically distributed (i.i.d.) and assumes every input has a known correct label — assumptions that don't hold for interactive decision-making tasks.
  • The RL loop: an agent observes a state s_t, chooses an action a_t, and the environment returns a new state s_{t+1} and a reward; the goal is a policy Pi_theta that maps states to actions and maximizes cumulative reward, not just immediate reward.
  • Credit assignment: when an outcome is delayed, it's often unclear which earlier action caused a success or failure, unlike labeled supervised data.
  • Reward framing across domains: dog training, robot locomotion, and inventory management are all reframed as (actions, observations, reward) tuples to show how broadly the RL formalism applies.
  • Applications shown: robotic hammering and jumping, Atari play discovering the "tunnel through the bricks" strategy, warehouse robots sorting trash, RL-controlled traffic flow reducing jams, and RL fine-tuning of language models (RLHF) and image generation models.

Walkthrough

Course Overview (0:00)

The instructor lists the course units and grading breakdown, describes the five homeworks (imitation learning, policy gradients/Q-learning, actor-critic, model-based RL, offline RL), and explains the research-level final project, including proposal and milestone checkpoints meant to give students feedback before the final report.

Reinforcement Learning (3:10)

The lecture defines reinforcement learning as both a formalism and a class of solution methods, then walks through the RL loop: an agent picks actions, the environment returns states and rewards, and the objective is a policy that maximizes cumulative reward over time rather than reward at a single step.

Supervised Learning (7:17)

By contrast, supervised learning is defined through its i.i.d. and fully-labeled-data assumptions. The lecture explains why these assumptions break down for sequential decision problems: past actions affect future inputs, and only outcomes (not correct actions) are observed, creating the credit assignment problem.

Examples (9:18)

The lecture works through examples — training a dog, controlling a robot, managing warehouse inventory — mapping each to actions, observations, and rewards, then shows applied results: robotic hammering and quadruped jumping, Atari game play, robotic trash sorting, RL-based traffic regulation, and RL used to fine-tune language models and image generation models toward human-preferred outputs.

Before you watch

  • Watch Part 1 of Lecture 1 first for the motivating grasping example and historical background.
  • Basic familiarity with the idea of training a model from a labeled dataset will help the supervised-vs-RL comparison land.

Check your understanding

  1. Why can't reinforcement learning rely on the i.i.d. assumption that supervised learning makes?
  2. What is the credit assignment problem, and why does it make RL harder than standard classification?
  3. In the traffic example, what did the RL-controlled car optimize for instead of its own speed?
  4. How was reinforcement learning used to improve the image generation example with the dolphin and bicycle?
  5. What are the five homework topics listed for this course?

Vocabulary

logistics (noun)
The practical details of organizing a course, like grading and schedule.
This part covers course logistics before the main content.
model-free (adjective)
A type of learning that doesn't build an explicit model of the environment.
Model-free RL includes Q-learning and policy gradient methods.
model-based (adjective)
A type of learning that builds a model to predict how the environment behaves.
Model-based RL uses predictions to plan actions.
exploration (noun)
Trying new actions to discover more about an environment.
Exploration helps an agent find better strategies.
offline RL (noun)
Learning a policy from a fixed dataset without further interaction.
Offline RL doesn't require live trial and error.
inverse RL (noun)
Learning a reward function by observing someone's behavior.
Inverse RL infers rewards from expert demonstrations.
meta-learning (noun)
Learning how to learn new tasks more quickly.
Meta-learning helps an agent adapt to new tasks fast.
formalism (noun)
A precise, structured way of describing a problem mathematically.
RL is both a formalism and a set of solution methods.
conflate (verb)
To wrongly treat two different things as if they were the same.
Don't conflate the RL problem with any one solution method.
i.i.d. (adjective)
Short for independent and identically distributed, meaning data points don't affect each other.
Supervised learning assumes i.i.d. data.
policy (noun)
A rule that maps a situation to an action for an agent to take.
The goal of RL is to find a good policy.
cumulative reward (noun)
The total reward added up over many steps, not just one.
The policy is trained to maximize cumulative reward.
credit assignment (noun)
The problem of figuring out which past action caused a later outcome.
Credit assignment is hard when rewards are delayed.
locomotion (noun)
The ability to move from one place to another.
Robot locomotion is a common RL application.
inventory management (noun)
Deciding how much stock to keep and order over time.
Inventory management can be framed as an RL problem.
fine-tuning (noun)
Further training an already-trained model on a more specific task.
RLHF is used for fine-tuning language models.
grading (noun)
The way scores are assigned in a course.
Grading is split between homework, project, and quizzes.
proposal (noun)
A written plan describing what a project intends to do.
Students submit a proposal before starting the final project.
milestone (noun)
A checkpoint marking progress partway through a longer project.
The course includes milestone checkpoints for the final project.
regulate (verb)
To control something so it stays within a desired range.
RL is used to regulate traffic flow.

Chapters

← Lecture 1: Introduction, Part 1 · Lecture 1: Introduction, Part 3 →