Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 48 of 99 · 6:53
Lecture 11, Part 4: Planning with Uncertainty-Aware Models
Study guide
What this lecture covers
This part closes out the uncertainty-aware modeling discussion from Part 3 by showing exactly how to plan using an ensemble of uncertainty-aware dynamics models, then reviews experimental evidence that this matters a great deal in practice. It's the practical payoff of the epistemic uncertainty machinery just introduced.
After watching, you'll be able to describe the sampling procedure for evaluating a candidate action sequence's expected reward across an ensemble of models, and you'll know the rough magnitude of improvement uncertainty-aware planning produced in real reported experiments.
Key ideas
- Averaging reward across models: instead of optimizing reward under one dynamics model, optimize the average reward across
Nsampled models (or the full ensemble), each contributing its own predicted trajectory. - Sampling procedure: for a candidate action sequence, sample a model from the posterior (e.g. pick one ensemble member), roll out states through it, compute the reward, and repeat to average over models.
- Compatible with standard planners: this scheme plugs into random shooting, CEM, or (with extra machinery like the reparameterization trick, covered later) LQR-style methods.
- Large empirical gains: on HalfCheetah, adding bootstrap-ensemble epistemic uncertainty to model-based RL version 1.5 raised reward from around 500 to over 6,000 in a similar amount of training time.
- Real-world robotic result: an in-hand manipulation robot using an ensemble-based model-based RL method learned to reliably rotate objects in its palm within a few hours of real-world interaction.
Walkthrough
How to plan with uncertainty (0:12)
The lecture extends the standard planning objective (sum of rewards over a horizon under one model) to an ensemble setting: optimize the action sequence that maximizes reward averaged over multiple sampled models. The described procedure samples a model from the posterior (trivial with a bootstrap ensemble — just pick one member), rolls out predicted states and rewards under it, and repeats across models and time steps to accumulate the average reward for a candidate action sequence, which is then optimized with a standard method like random shooting or CEM.
Example: model-based RL with ensembles (4:02)
Referencing the paper "Deep Reinforcement Learning in a Handful of Trials," the lecture reports that adding bootstrap-ensemble epistemic uncertainty raised HalfCheetah reward from roughly 500 (version 1.5 without uncertainty) to over 6,000, in comparable training time — a dramatic demonstration that uncertainty awareness matters most in low-data regimes.
More recent example: PDDM (4:41)
A real-world robotic hand manipulation experiment is described, using an ensemble-based model-based RL method (PDDM) to learn in-hand object rotation directly from interaction, reaching a reliable full 180-degree turn after a few hours of real-world training.
Further readings (5:40)
The lecture points to several further readings: PILCO (Deisenroth), an older foundational paper using Gaussian processes that established the importance of epistemic uncertainty in model-based RL; the "handful of trials" ensemble paper; and papers combining models with value functions and integrating epistemic uncertainty at multiple points in the pipeline.
Before you watch
- Watch Lecture 11, Part 3 first for the definition of epistemic uncertainty and how bootstrap ensembles approximate the model posterior.
- Recall the model-based RL version 1.5 (MPC) procedure from Part 1, since this part directly extends its planning step.
Check your understanding
- How does planning with an ensemble of models change the objective being optimized, compared to planning with a single model?
- Describe the steps of the sampling procedure used to estimate the expected reward of a candidate action sequence under model uncertainty.
- By roughly how much did epistemic uncertainty estimation improve HalfCheetah performance in the reported experiment, and what does that suggest about when uncertainty matters most?
- What real-world task did the PDDM-style robotic hand experiment demonstrate, and how long did learning take?
- Which planning methods can this ensemble-averaging approach be combined with, according to the lecture?
Vocabulary
- ensemble (noun)
- A group of several models trained together and combined for a task.
We plan through an ensemble of dynamics models. - posterior (noun)
- The updated belief distribution over a value after seeing data.
We sample a model from the posterior over models. - average reward (phrase)
- The mean reward across several trials or models.
We optimize the average reward across sampled models. - reparameterization trick (phrase)
- A technique for making random sampling differentiable so gradients can flow through it.
LQR-style methods need the reparameterization trick with uncertainty. - empirical gain (phrase)
- An improvement measured in real experiments.
The empirical gains from uncertainty were dramatic. - in-hand manipulation (phrase)
- The task of a robotic hand moving or repositioning an object it holds.
The robot learned in-hand manipulation of an object. - robotic (adjective)
- Related to machines that can sense and act in the physical world.
This is a real-world robotic result. - rotate (verb)
- To turn something around a fixed point or axis.
The robot learned to rotate objects in its palm. - handful of trials (phrase)
- A very small number of attempts or episodes.
The paper is called 'Deep RL in a Handful of Trials.' - Gaussian process (phrase)
- A probabilistic model that predicts a distribution over possible functions.
PILCO used a Gaussian process for its dynamics model. - value function (noun)
- A function that estimates expected future reward from a state.
Some methods combine models with value functions. - reliable (adjective)
- Consistently able to be trusted to work correctly.
The robot learned a reliable 180-degree turn. - comparable (adjective)
- Roughly equal or similar in amount.
The gain appeared in comparable training time. - candidate action sequence (phrase)
- A possible full plan of actions being considered.
We evaluate each candidate action sequence across models. - sampling procedure (phrase)
- A defined set of steps for randomly drawing examples.
The sampling procedure picks one ensemble member at a time. - accumulate (verb)
- To gradually gather or add up over time.
We accumulate the average reward across models and steps. - practical payoff (phrase)
- A real, useful benefit gained from applying an idea.
This is the practical payoff of the uncertainty machinery. - plug into (phrasal verb)
- To fit and work together with an existing system.
This scheme plugs into random shooting or CEM. - magnitude (noun)
- The size or scale of a quantity.
We look at the rough magnitude of improvement. - close out (phrasal verb)
- To finish or conclude a topic.
This part closes out the uncertainty-aware modeling discussion.
Chapters
- 0:00 <Untitled Chapter 1>
- 0:12 How to plan with uncertainty
- 4:02 Example: model-based RL with ensembles
- 4:41 More recent example: PDDM
- 5:40 Further readings
← Lecture 11, Part 3: Estimating Epistemic Uncertainty with Neural Networks · Lecture 11, Part 5: Model-Based RL with Image Observations →
