Seyed Masoud Hosseini · Overview · Study log · Ideas · Transcript · RSS feed

Deep Reinforcement Learning

Berkeley CS285, Sergey Levine · AI · Spring 2028 · Planned

Lectures

  1. Lecture 1: Introduction, Part 1 (10:10)
  2. Lecture 1: Introduction, Part 2 (17:54)
  3. Lecture 1: Introduction, Part 3 (29:01)
  4. Lecture 2: Imitation Learning, Part 1 (24:47)
  5. Lecture 2: Imitation Learning, Part 2 (23:06)
  6. Lecture 2: Imitation Learning, Part 3 (32:23)
  7. Lecture 2: Imitation Learning, Part 4 (8:31)
  8. Lecture 2: Imitation Learning, Part 5 (8:51)
  9. Lecture 4: Introduction to RL Algorithms, Part 1 (26:28)
  10. Lecture 4: Introduction to RL Algorithms, Part 2 (7:23)
  11. Lecture 4, Part 3: Q-Functions and Value Functions (9:11)
  12. Lecture 4, Part 4: Types of RL Algorithms (5:50)
  13. Lecture 4, Part 5: Comparing RL Algorithms (9:14)
  14. Lecture 4, Part 6: Examples of Deep RL Algorithms (3:04)
  15. Lecture 5, Part 1: Deriving the Policy Gradient (14:01)
  16. Lecture 5, Part 2: Intuition and the High-Variance Problem (13:17)
  17. Lecture 5, Part 3: Reducing Variance with Causality and Baselines (14:52)
  18. Lecture 5, Part 4: Off-Policy Policy Gradients with Importance Sampling (15:40)
  19. Lecture 5, Part 5: Implementing Policy Gradients in Practice (7:32)
  20. Lecture 5, Part 6: The Natural Policy Gradient (13:32)
  21. Lecture 6, Part 1: Actor-Critic and Value Functions (25:13)
  22. Lecture 6, Part 2: Discount Factors in Actor-Critic (17:29)
  23. Lecture 6, Part 3: Implementing Actor-Critic (18:31)
  24. Lecture 6, Part 4: Eligibility Traces and GAE (15:54)
  25. Lecture 6, Part 5: Actor-Critic Summary and Examples (3:36)
  26. Lecture 7, Part 1: From Actor-Critic to Policy Iteration (16:16)
  27. Lecture 7, Part 2: Fitted Value and Fitted Q-Iteration (15:08)
  28. Lecture 7, Part 3: Q-Learning and Exploration (11:57)
  29. Lecture 7, Part 4: Why Fitted Q-Iteration Doesn't Converge (17:01)
  30. Lecture 8, Part 1: Replay Buffers and the Correlation Problem (13:37)
  31. Lecture 8, Part 2: Target Networks for Q-Learning (11:56)
  32. Lecture 8, Part 3: A Unified View of Q-Learning (8:35)
  33. Lecture 8, Part 4: Overestimation, Double Q-Learning, N-Step Returns (23:41)
  34. Lecture 8, Part 5: Q-Learning with Continuous Actions (10:04)
  35. Lecture 8, Part 6: Practical Tips and Q-Learning Case Studies (11:00)
  36. Lecture 9, Part 1: Why Does Policy Gradient Work? (21:23)
  37. Lecture 9, Part 2: Bounding the State Distribution Mismatch (18:48)
  38. Lecture 9, Part 3: Constraining Policy Gradient with KL Divergence (5:50)
  39. Lecture 9, Part 4: Natural Gradient and Trust Region Methods (21:08)
  40. Lecture 10, Part 1: Introduction to Model-Based Planning (19:19)
  41. Lecture 10, Part 2: Stochastic Optimization for Planning (23:11)
  42. Lecture 10, Part 3: Trajectory Optimization with the LQR (23:49)
  43. Lecture 10, Part 4: Extending LQR to Stochastic and Nonlinear Systems (13:23)
  44. Lecture 10, Part 5: A Case Study in Optimal Control (6:21)
  45. Lecture 11, Part 1: Model-Based RL and Distributional Shift (17:46)
  46. Lecture 11, Part 2: Why Model-Based RL Underperforms, and Uncertainty (9:35)
  47. Lecture 11, Part 3: Estimating Epistemic Uncertainty with Neural Networks (17:08)
  48. Lecture 11, Part 4: Planning with Uncertainty-Aware Models (6:53)
  49. Lecture 11, Part 5: Model-Based RL with Image Observations (17:22)
  50. Lecture 12, Part 1: Model-Based RL with Policies (15:05)
  51. Lecture 12, Part 2: Model-Based RL with Policies (15:04)
  52. Lecture 12, Part 3: Model-Based RL with Policies (12:53)
  53. Lecture 12, Part 4: Model-Based RL with Policies (29:20)
  54. Lecture 13, Part 1: Exploration (19:51)
  55. Lecture 13, Part 2: Exploration (15:48)
  56. Lecture 13, Part 3: Exploration (14:31)
  57. Lecture 13, Part 4: Exploration (13:02)
  58. Lecture 13, Part 5: Exploration (6:49)
  59. Lecture 13, Part 6: Exploration (12:40)
  60. Lecture 14, Part 1 (14:17)
  61. Lecture 14, Part 2: Learning Goal-Reaching Policies Without Rewards (15:27)
  62. Lecture 14, Part 3: State Marginal Matching and Intrinsic Motivation (13:52)
  63. Lecture 14, Part 4: Learning Diverse Skills (8:02)
  64. Lecture 15, Part 1: What Is Offline Reinforcement Learning? (38:01)
  65. Lecture 15, Part 2: Offline RL by Importance Sampling (25:33)
  66. Lecture 15, Part 3: Classic Offline RL with Linear Value Functions (21:20)
  67. Lecture 16, Part 1: Policy Constraints and Implicit Q-Learning (31:58)
  68. Lecture 16, Part 2: Conservative Q-Learning (CQL) (7:33)
  69. Lecture 16, Part 3: Model-Based Offline RL (18:16)
  70. Lecture 16, Part 4: Offline RL in Practice, Applications, and Open Problems (12:19)
  71. Lecture 17, Part 1: RL Theory (50:09)
  72. Lecture 17, Part 2: RL Theory (21:58)
  73. Lecture 18, Variational Inference, Part 1 (20:12)
  74. Lecture 18, Variational Inference, Part 2 (19:36)
  75. Lecture 18, Variational Inference, Part 3 (17:52)
  76. Lecture 18, Variational Inference, Part 4 (25:29)
  77. Lecture 19, Control as Inference, Part 1 (21:05)
  78. Lecture 19, Control as Inference, Part 2 (28:38)
  79. Lecture 19, Control as Inference, Part 3 (21:26)
  80. Lecture 19, Control as Inference, Part 4 (9:54)
  81. Lecture 19: Control as Inference, Part 5 (10:58)
  82. Lecture 20: Inverse Reinforcement Learning, Part 1 (23:50)
  83. Lecture 20: Inverse Reinforcement Learning, Part 2 (12:28)
  84. Lecture 20: Inverse Reinforcement Learning, Part 3 (8:21)
  85. Lecture 20: Inverse Reinforcement Learning, Part 4 (14:21)
  86. Guest Lecture: Eric Mitchell on RLHF, Algorithms and Applications (54:28)
  87. Guest Lecture: Andrea Zanette on Statistical Foundations of RL (1:00:14)
  88. Lecture 21: RL with Sequence Models & Language Models, Part 1 (29:54)
  89. Lecture 21: RL with Sequence Models & Language Models, Part 2 (23:39)
  90. Lecture 21: RL with Sequence Models & Language Models, Part 3 (16:58)
  91. Lecture 22, Part 1: Transfer Learning and Domain Adaptation (44:17)
  92. Lecture 22, Part 2: What Is Meta-Learning? (9:48)
  93. Lecture 22, Part 3: Meta Reinforcement Learning with RNNs (11:39)
  94. Lecture 22, Part 4: Gradient-Based Meta-Reinforcement Learning (7:38)
  95. Lecture 22, Part 5: Meta-RL as Partially Observed MDPs (13:26)
  96. Lecture 23, Part 1: Challenges and Open Problems in Deep RL (28:24)
  97. Lecture 23, Part 2: Three Perspectives on What RL Is (38:00)
  98. Guest Lecture: Aviral Kumar on Offline RL for Pre-training (56:17)
  99. Guest Lecture: Dorsa Sadigh on Interactive Learning (1:01:41)

Notes

No notes yet.

References

No references yet.

Study log

No log entries for this course yet.