Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 72 of 99 · 21:58
Lecture 17, Part 2: RL Theory
Study guide
What this lecture covers
This part continues the CS285 theory lecture, moving from the model-based analysis of part 1 to an idealized version of fitted Q-iteration, a model-free method. Because real fitted Q-iteration is not guaranteed to converge, the lecture studies a simplified version amenable to analysis, and asks how far the resulting q_hat_k ends up from q_star after many iterations.
You'll see the error decomposed into two separate sources: sampling error, from using an empirical Bellman operator t_hat built from averaged rewards and transitions instead of the true operator t, and approximation error, from fitting the Q-function inexactly at each iteration. The lecture then combines both to produce a bound on the asymptotic error of approximate fitted Q-iteration.
Key ideas
- Idealized fitted Q-iteration: models the effect of averaging many sample losses for the same state-action pair as a backup under an empirical transition model
p_hatand empirical rewardr_hat. - Sampling error: the gap between the empirical Bellman operator
t_hatand the true operatort, bounded using the same concentration inequalities from part 1 and scaling as1/sqrt(n). - Approximation error: the gap between the fitted
q_hat_{k+1}and the exact backupt q_hat_k, assumed bounded by someepsilon_kin the infinity norm. - Contraction property: the Bellman backup shrinks the infinity-norm distance between two Q-functions by a factor of
gamma, which is what lets the per-iteration errors be summed into a geometric series. - Forgetting the initialization: as the number of iterations grows, the effect of the starting
q_hat_0vanishes becausegamma^kgoes to zero. - Combined error bound: the asymptotic error is proportional to
1/(1-gamma)times the largest per-iteration error, and because sampling error itself contains a1/(1-gamma)factor, the overall bound is quadratic in the effective horizon1/(1-gamma).
Before you watch
- Watch part 1 of this lecture first: it introduces the concentration inequalities and the
1/(1-gamma)horizon argument reused here. - Review the Bellman contraction property and the definition of fitted Q-iteration from the earlier value-based methods lectures.
Check your understanding
- What are the two distinct sources of error in idealized fitted Q-iteration, and which one comes from having a finite number of samples?
- How does the contraction property of the Bellman backup let you turn a per-iteration error bound into a bound on the limiting error as
kapproaches infinity? - Why does the final error bound end up quadratic in
1/(1-gamma)rather than linear? - Why does the lecture call the infinity-norm assumption on fitting error a strong assumption, and what alternative does it mention?
Vocabulary
- fitted Q-iteration (noun)
- A method that repeatedly fits a Q-function to targets computed from the previous Q-function.
The lecture analyzes an idealized version of fitted Q-iteration. - idealized (adjective)
- Simplified to make analysis easier, ignoring some real-world complications.
The idealized version is easier to analyze than real fitted Q-iteration. - converge (verb)
- To gradually approach and settle near a final value.
Real fitted Q-iteration is not guaranteed to converge. - decompose (verb)
- To break something into separate, simpler parts.
The total error is decomposed into two separate sources. - empirical (adjective)
- Based on observed data or samples rather than exact theory.
The empirical Bellman operator is built from sampled data. - operator (noun)
- A mathematical rule that transforms one function into another.
The Bellman operator maps one Q-function to an updated one. - sampling error (noun)
- The gap between a true value and an estimate made from limited samples.
Sampling error shrinks as more data is collected. - approximation error (noun)
- The gap caused by fitting a function imperfectly rather than exactly.
Approximation error comes from imperfect fitting at each step. - contraction (noun)
- A property where an operation always shrinks the distance between two things.
The Bellman backup has a contraction property with factor gamma. - geometric series (noun)
- A sum of terms where each term is a fixed multiple of the one before it.
The errors are summed as a geometric series. - asymptotic (adjective)
- Describing behavior as some quantity grows very large, such as many iterations.
The asymptotic error bound describes behavior after many iterations. - bootstrapped (adjective)
- Built using its own earlier estimates rather than fully independent data.
The bootstrapped target introduces its own error term. - concentration inequality (noun)
- A mathematical bound showing that a sample average is close to its true value with high probability.
Sampling error is bounded using concentration inequalities. - infinity norm (noun)
- A way of measuring the size of the largest difference between two functions.
Approximation error is measured in the infinity norm. - horizon (noun)
- The effective length of time or number of steps a problem looks ahead over.
The bound is quadratic in the effective horizon 1/(1-gamma). - quadratic (adjective)
- Growing in proportion to the square of a quantity.
The final error bound is quadratic in the horizon term. - bound (noun)
- A mathematical limit that a quantity cannot exceed.
The lecture derives a bound on the asymptotic error. - model-free (adjective)
- Describing a method that does not build an explicit model of the environment's dynamics.
Fitted Q-iteration is a model-free method. - backup (noun)
- An update that computes a new value estimate from other value estimates.
The Bellman backup shrinks distance between Q-functions by a factor of gamma. - initialization (noun)
- The starting values a process begins from before it runs.
The effect of the starting q_hat_0 initialization vanishes over time. - vanish (verb)
- To shrink to nothing.
The effect of the initial guess vanishes as gamma^k goes to zero. - proportional (adjective)
- Changing at a constant rate relative to another quantity.
The asymptotic error is proportional to the largest per-iteration error. - strong assumption (phrase)
- A condition that is unlikely to hold exactly in practice but is assumed for the analysis.
Bounding the fitting error in the infinity norm is a strong assumption. - alternative (noun)
- A different option that could be used instead.
The lecture mentions an alternative to the infinity-norm assumption. - gap (noun)
- The difference between two values.
Sampling error is the gap between the empirical and true operators.
Chapters
← Lecture 17, Part 1: RL Theory · Lecture 18, Variational Inference, Part 1 →
