Seyed Masoud Hosseini · Overview · Study log · Ideas · Transcript · RSS feed
Deep Reinforcement Learning
Berkeley CS285, Sergey Levine · AI · Spring 2028 · Planned
Lectures
- Lecture 1: Introduction, Part 1 (10:10)
- Lecture 1: Introduction, Part 2 (17:54)
- Lecture 1: Introduction, Part 3 (29:01)
- Lecture 2: Imitation Learning, Part 1 (24:47)
- Lecture 2: Imitation Learning, Part 2 (23:06)
- Lecture 2: Imitation Learning, Part 3 (32:23)
- Lecture 2: Imitation Learning, Part 4 (8:31)
- Lecture 2: Imitation Learning, Part 5 (8:51)
- Lecture 4: Introduction to RL Algorithms, Part 1 (26:28)
- Lecture 4: Introduction to RL Algorithms, Part 2 (7:23)
- Lecture 4, Part 3: Q-Functions and Value Functions (9:11)
- Lecture 4, Part 4: Types of RL Algorithms (5:50)
- Lecture 4, Part 5: Comparing RL Algorithms (9:14)
- Lecture 4, Part 6: Examples of Deep RL Algorithms (3:04)
- Lecture 5, Part 1: Deriving the Policy Gradient (14:01)
- Lecture 5, Part 2: Intuition and the High-Variance Problem (13:17)
- Lecture 5, Part 3: Reducing Variance with Causality and Baselines (14:52)
- Lecture 5, Part 4: Off-Policy Policy Gradients with Importance Sampling (15:40)
- Lecture 5, Part 5: Implementing Policy Gradients in Practice (7:32)
- Lecture 5, Part 6: The Natural Policy Gradient (13:32)
- Lecture 6, Part 1: Actor-Critic and Value Functions (25:13)
- Lecture 6, Part 2: Discount Factors in Actor-Critic (17:29)
- Lecture 6, Part 3: Implementing Actor-Critic (18:31)
- Lecture 6, Part 4: Eligibility Traces and GAE (15:54)
- Lecture 6, Part 5: Actor-Critic Summary and Examples (3:36)
- Lecture 7, Part 1: From Actor-Critic to Policy Iteration (16:16)
- Lecture 7, Part 2: Fitted Value and Fitted Q-Iteration (15:08)
- Lecture 7, Part 3: Q-Learning and Exploration (11:57)
- Lecture 7, Part 4: Why Fitted Q-Iteration Doesn't Converge (17:01)
- Lecture 8, Part 1: Replay Buffers and the Correlation Problem (13:37)
- Lecture 8, Part 2: Target Networks for Q-Learning (11:56)
- Lecture 8, Part 3: A Unified View of Q-Learning (8:35)
- Lecture 8, Part 4: Overestimation, Double Q-Learning, N-Step Returns (23:41)
- Lecture 8, Part 5: Q-Learning with Continuous Actions (10:04)
- Lecture 8, Part 6: Practical Tips and Q-Learning Case Studies (11:00)
- Lecture 9, Part 1: Why Does Policy Gradient Work? (21:23)
- Lecture 9, Part 2: Bounding the State Distribution Mismatch (18:48)
- Lecture 9, Part 3: Constraining Policy Gradient with KL Divergence (5:50)
- Lecture 9, Part 4: Natural Gradient and Trust Region Methods (21:08)
- Lecture 10, Part 1: Introduction to Model-Based Planning (19:19)
- Lecture 10, Part 2: Stochastic Optimization for Planning (23:11)
- Lecture 10, Part 3: Trajectory Optimization with the LQR (23:49)
- Lecture 10, Part 4: Extending LQR to Stochastic and Nonlinear Systems (13:23)
- Lecture 10, Part 5: A Case Study in Optimal Control (6:21)
- Lecture 11, Part 1: Model-Based RL and Distributional Shift (17:46)
- Lecture 11, Part 2: Why Model-Based RL Underperforms, and Uncertainty (9:35)
- Lecture 11, Part 3: Estimating Epistemic Uncertainty with Neural Networks (17:08)
- Lecture 11, Part 4: Planning with Uncertainty-Aware Models (6:53)
- Lecture 11, Part 5: Model-Based RL with Image Observations (17:22)
- Lecture 12, Part 1: Model-Based RL with Policies (15:05)
- Lecture 12, Part 2: Model-Based RL with Policies (15:04)
- Lecture 12, Part 3: Model-Based RL with Policies (12:53)
- Lecture 12, Part 4: Model-Based RL with Policies (29:20)
- Lecture 13, Part 1: Exploration (19:51)
- Lecture 13, Part 2: Exploration (15:48)
- Lecture 13, Part 3: Exploration (14:31)
- Lecture 13, Part 4: Exploration (13:02)
- Lecture 13, Part 5: Exploration (6:49)
- Lecture 13, Part 6: Exploration (12:40)
- Lecture 14, Part 1 (14:17)
- Lecture 14, Part 2: Learning Goal-Reaching Policies Without Rewards (15:27)
- Lecture 14, Part 3: State Marginal Matching and Intrinsic Motivation (13:52)
- Lecture 14, Part 4: Learning Diverse Skills (8:02)
- Lecture 15, Part 1: What Is Offline Reinforcement Learning? (38:01)
- Lecture 15, Part 2: Offline RL by Importance Sampling (25:33)
- Lecture 15, Part 3: Classic Offline RL with Linear Value Functions (21:20)
- Lecture 16, Part 1: Policy Constraints and Implicit Q-Learning (31:58)
- Lecture 16, Part 2: Conservative Q-Learning (CQL) (7:33)
- Lecture 16, Part 3: Model-Based Offline RL (18:16)
- Lecture 16, Part 4: Offline RL in Practice, Applications, and Open Problems (12:19)
- Lecture 17, Part 1: RL Theory (50:09)
- Lecture 17, Part 2: RL Theory (21:58)
- Lecture 18, Variational Inference, Part 1 (20:12)
- Lecture 18, Variational Inference, Part 2 (19:36)
- Lecture 18, Variational Inference, Part 3 (17:52)
- Lecture 18, Variational Inference, Part 4 (25:29)
- Lecture 19, Control as Inference, Part 1 (21:05)
- Lecture 19, Control as Inference, Part 2 (28:38)
- Lecture 19, Control as Inference, Part 3 (21:26)
- Lecture 19, Control as Inference, Part 4 (9:54)
- Lecture 19: Control as Inference, Part 5 (10:58)
- Lecture 20: Inverse Reinforcement Learning, Part 1 (23:50)
- Lecture 20: Inverse Reinforcement Learning, Part 2 (12:28)
- Lecture 20: Inverse Reinforcement Learning, Part 3 (8:21)
- Lecture 20: Inverse Reinforcement Learning, Part 4 (14:21)
- Guest Lecture: Eric Mitchell on RLHF, Algorithms and Applications (54:28)
- Guest Lecture: Andrea Zanette on Statistical Foundations of RL (1:00:14)
- Lecture 21: RL with Sequence Models & Language Models, Part 1 (29:54)
- Lecture 21: RL with Sequence Models & Language Models, Part 2 (23:39)
- Lecture 21: RL with Sequence Models & Language Models, Part 3 (16:58)
- Lecture 22, Part 1: Transfer Learning and Domain Adaptation (44:17)
- Lecture 22, Part 2: What Is Meta-Learning? (9:48)
- Lecture 22, Part 3: Meta Reinforcement Learning with RNNs (11:39)
- Lecture 22, Part 4: Gradient-Based Meta-Reinforcement Learning (7:38)
- Lecture 22, Part 5: Meta-RL as Partially Observed MDPs (13:26)
- Lecture 23, Part 1: Challenges and Open Problems in Deep RL (28:24)
- Lecture 23, Part 2: Three Perspectives on What RL Is (38:00)
- Guest Lecture: Aviral Kumar on Offline RL for Pre-training (56:17)
- Guest Lecture: Dorsa Sadigh on Interactive Learning (1:01:41)
Notes
No notes yet.
References
No references yet.
Study log
No log entries for this course yet.