Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 72 of 99 · 21:58

Lecture 17, Part 2: RL Theory

CS 285: Lecture 17, Part 2: RL Theory on YouTube

Study guide

What this lecture covers

This part continues the CS285 theory lecture, moving from the model-based analysis of part 1 to an idealized version of fitted Q-iteration, a model-free method. Because real fitted Q-iteration is not guaranteed to converge, the lecture studies a simplified version amenable to analysis, and asks how far the resulting q_hat_k ends up from q_star after many iterations.

You'll see the error decomposed into two separate sources: sampling error, from using an empirical Bellman operator t_hat built from averaged rewards and transitions instead of the true operator t, and approximation error, from fitting the Q-function inexactly at each iteration. The lecture then combines both to produce a bound on the asymptotic error of approximate fitted Q-iteration.

Key ideas

  • Idealized fitted Q-iteration: models the effect of averaging many sample losses for the same state-action pair as a backup under an empirical transition model p_hat and empirical reward r_hat.
  • Sampling error: the gap between the empirical Bellman operator t_hat and the true operator t, bounded using the same concentration inequalities from part 1 and scaling as 1/sqrt(n).
  • Approximation error: the gap between the fitted q_hat_{k+1} and the exact backup t q_hat_k, assumed bounded by some epsilon_k in the infinity norm.
  • Contraction property: the Bellman backup shrinks the infinity-norm distance between two Q-functions by a factor of gamma, which is what lets the per-iteration errors be summed into a geometric series.
  • Forgetting the initialization: as the number of iterations grows, the effect of the starting q_hat_0 vanishes because gamma^k goes to zero.
  • Combined error bound: the asymptotic error is proportional to 1/(1-gamma) times the largest per-iteration error, and because sampling error itself contains a 1/(1-gamma) factor, the overall bound is quadratic in the effective horizon 1/(1-gamma).

Before you watch

  • Watch part 1 of this lecture first: it introduces the concentration inequalities and the 1/(1-gamma) horizon argument reused here.
  • Review the Bellman contraction property and the definition of fitted Q-iteration from the earlier value-based methods lectures.

Check your understanding

  1. What are the two distinct sources of error in idealized fitted Q-iteration, and which one comes from having a finite number of samples?
  2. How does the contraction property of the Bellman backup let you turn a per-iteration error bound into a bound on the limiting error as k approaches infinity?
  3. Why does the final error bound end up quadratic in 1/(1-gamma) rather than linear?
  4. Why does the lecture call the infinity-norm assumption on fitting error a strong assumption, and what alternative does it mention?

Vocabulary

fitted Q-iteration (noun)
A method that repeatedly fits a Q-function to targets computed from the previous Q-function.
The lecture analyzes an idealized version of fitted Q-iteration.
idealized (adjective)
Simplified to make analysis easier, ignoring some real-world complications.
The idealized version is easier to analyze than real fitted Q-iteration.
converge (verb)
To gradually approach and settle near a final value.
Real fitted Q-iteration is not guaranteed to converge.
decompose (verb)
To break something into separate, simpler parts.
The total error is decomposed into two separate sources.
empirical (adjective)
Based on observed data or samples rather than exact theory.
The empirical Bellman operator is built from sampled data.
operator (noun)
A mathematical rule that transforms one function into another.
The Bellman operator maps one Q-function to an updated one.
sampling error (noun)
The gap between a true value and an estimate made from limited samples.
Sampling error shrinks as more data is collected.
approximation error (noun)
The gap caused by fitting a function imperfectly rather than exactly.
Approximation error comes from imperfect fitting at each step.
contraction (noun)
A property where an operation always shrinks the distance between two things.
The Bellman backup has a contraction property with factor gamma.
geometric series (noun)
A sum of terms where each term is a fixed multiple of the one before it.
The errors are summed as a geometric series.
asymptotic (adjective)
Describing behavior as some quantity grows very large, such as many iterations.
The asymptotic error bound describes behavior after many iterations.
bootstrapped (adjective)
Built using its own earlier estimates rather than fully independent data.
The bootstrapped target introduces its own error term.
concentration inequality (noun)
A mathematical bound showing that a sample average is close to its true value with high probability.
Sampling error is bounded using concentration inequalities.
infinity norm (noun)
A way of measuring the size of the largest difference between two functions.
Approximation error is measured in the infinity norm.
horizon (noun)
The effective length of time or number of steps a problem looks ahead over.
The bound is quadratic in the effective horizon 1/(1-gamma).
quadratic (adjective)
Growing in proportion to the square of a quantity.
The final error bound is quadratic in the horizon term.
bound (noun)
A mathematical limit that a quantity cannot exceed.
The lecture derives a bound on the asymptotic error.
model-free (adjective)
Describing a method that does not build an explicit model of the environment's dynamics.
Fitted Q-iteration is a model-free method.
backup (noun)
An update that computes a new value estimate from other value estimates.
The Bellman backup shrinks distance between Q-functions by a factor of gamma.
initialization (noun)
The starting values a process begins from before it runs.
The effect of the starting q_hat_0 initialization vanishes over time.
vanish (verb)
To shrink to nothing.
The effect of the initial guess vanishes as gamma^k goes to zero.
proportional (adjective)
Changing at a constant rate relative to another quantity.
The asymptotic error is proportional to the largest per-iteration error.
strong assumption (phrase)
A condition that is unlikely to hold exactly in practice but is assumed for the analysis.
Bounding the fitting error in the infinity norm is a strong assumption.
alternative (noun)
A different option that could be used instead.
The lecture mentions an alternative to the infinity-norm assumption.
gap (noun)
The difference between two values.
Sampling error is the gap between the empirical and true operators.

Chapters

← Lecture 17, Part 1: RL Theory · Lecture 18, Variational Inference, Part 1 →