Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Machine Learning · Lecture 20 of 21 · 1:21:06

Lecture 19: State-Action Rewards, Finite-Horizon MDPs and LQR

Lecture 19 - Reward Model & Linear Dynamical System | Stanford CS229: Machine Learning (Autumn 2018) on YouTube

Study guide

What this lecture covers

This lecture extends the MDP framework in two small but useful ways before introducing one powerful special case that can be solved exactly. It first shows how to let rewards depend on the action taken, not just the state, and how to model tasks that end after a fixed number of steps (a finite horizon) rather than running forever with a discount factor. Both generalizations make certain robotics and factory-automation problems easier to describe, and the finite-horizon case introduces the idea of a policy that changes with time.

The second half of the lecture introduces the linear quadratic regulator (LQR), a setting where the state transitions are linear in the state and action, and the reward is a negative quadratic cost. Under these two assumptions, the lecture shows that the optimal value function is exactly quadratic and the optimal policy is exactly linear in the state, computable by a backward dynamic-programming recursion with no function approximation at all. It also covers two ways to obtain the linear dynamics matrices, from data or by linearizing a nonlinear physics model, and ends with the observation that the optimal LQR policy does not depend on the amount of noise in the system.

Key ideas

  • State-action rewards: letting R depend on both state and action (R(s, a)) lets a designer penalize costly or aggressive actions, such as sudden control-stick movements on a helicopter, and shifts the max over actions outside the reward term in Bellman's equation.
  • Finite-horizon MDP: replaces the discount factor with a fixed horizon T, after which the process ends; the total payoff is a finite, undiscounted sum of rewards over T steps.
  • Non-stationary policy: in a finite-horizon MDP the optimal action can depend on how much time is left, so the optimal policy is written Pi*_t(s) rather than a single time-independent Pi*(s).
  • Finite-horizon dynamic programming: V*_t(s) is computed backward from the final time step T, where the optimal action is whatever maximizes the immediate reward, down to t = 0, using V*_t as a function of V*_(t+1).
  • LQR's two assumptions: state transitions are linear-Gaussian, s_(t+1) = A*s_t + B*a_t + w_t, and the reward is a negative quadratic cost, R(s,a) = -(s^T U s + a^T V a) with U and V positive semidefinite.
  • Exact quadratic value function: under these assumptions, V*_t(s) turns out to be exactly a quadratic function of s at every time step, computed by a backward recursion with no approximation.
  • Linear optimal policy: the optimal action at each time step is an exact linear function of the current state, Pi*_t(s) = L_t * s, not merely an approximation fit to a linear function.
  • Noise independence: the optimal LQR policy does not depend on the covariance of the noise term w_t, even though the value function does; this is a property specific to LQR, not reinforcement learning in general.

Walkthrough

State-action rewards (4:23)

The lecture generalizes the reward function from R(s) to R(s, a), so different actions in the same state can carry different costs, for example penalizing motion versus staying still in the grid-world robot, or penalizing aggressive control-stick movements on a helicopter. Bellman's equation adjusts accordingly, with the max over actions now applying to the immediate reward plus the discounted future value together, since the immediate reward itself depends on the chosen action. Value iteration still works unchanged in this formulation.

Finite-horizon MDPs and non-stationary policies (11:04)

Rather than discounting rewards over an infinite horizon, a finite-horizon MDP runs for a fixed number of steps T and then stops, with the total payoff simply the undiscounted sum of rewards collected. The lecture shows that in this setting the optimal action can depend on how much time remains, using a grid-world example where a robot near a small reward and a large reward should chase the large one if there is enough time left but settle for the small one if the clock is nearly out. This makes the optimal policy non-stationary, written Pi*_t(s), in contrast to the stationary policies from earlier lectures.

Non-stationary dynamics and rewards in practice (17:29)

The lecture extends the finite-horizon idea further by allowing the state transition probabilities and reward function themselves to change over time, giving examples such as a commercial airplane's dynamics changing as it burns off a third of its weight in fuel, driving conditions changing with rush-hour traffic or rain, and labor costs in a 24-hour factory varying by time of day. These are modeled by indexing P and R by time, t.

Dynamic programming for the finite-horizon value function (22:02)

The lecture derives a backward recursion for V*_t(s): at the final time step T, the optimal value is just the best immediate reward available, since there is no future left; at earlier steps, V*_t(s) is the max over actions of the immediate reward plus the expected value of V*_(t+1) at the resulting next state. Working backward from V*_T down to V*_0 gives the optimal value function, and hence the optimal non-stationary policy, at every time step and state.

LQR: linear dynamics and quadratic cost (33:22)

The linear quadratic regulator framework applies when the state transitions are linear in the previous state and action plus Gaussian noise, s_(t+1) = A*s_t + B*a_t + w_t, and the reward is a negative quadratic form, R(s,a) = -(s^T U s + a^T V a), with U and V positive semidefinite so the reward is always non-positive. A simple example sets U and V to identity matrices to penalize a helicopter for deviating from a hovering state near zero or for making large control actions.

Getting the matrices A and B: learning versus linearizing (40:35)

The lecture describes two ways to obtain the linear dynamics matrices. The first, as in the previous lecture, fits A and B by regression on recorded trajectories from flying a real system such as a helicopter at low speeds. The second linearizes a nonlinear physics model around a typical state and action (often the zero, or hovering, state) by taking a first-order Taylor approximation, which is a good fit as long as the system stays close to that operating point; the lecture works through the one-dimensional case before generalizing to state and action together.

The LQR dynamic-programming solution and the noise-independence fact (57:07)

Applying the same backward dynamic-programming idea as before, the lecture shows that if V*_(t+1) is a quadratic function of the state (parameterized by a matrix Phi_(t+1) and scalar Psi_(t+1)), then V*_t is also exactly quadratic with the same form, and the optimal action a_t turns out to be an exact linear function of the state, Pi*_t(s) = L_t * s, derived by maximizing the quadratic expression and setting its derivative to zero. Running this recursion from t = T down to t = 0 yields the exact optimal policy with no function approximation, only the initial assumptions of linear dynamics and quadratic cost. The lecture closes with a specific, LQR-only observation: L_t depends on Phi_(t+1) but not on Psi_(t+1), and the noise covariance Sigma_w only ever affects Psi, so the optimal policy does not depend on the noise level at all, even though the value function does.

Before you watch

  • Review Bellman's equation, value iteration, and fitted value iteration from the previous lectures in this course, since this lecture builds directly on them.
  • Be comfortable with basic matrix algebra, including positive semidefinite matrices and quadratic forms.
  • Recall the discussion from the previous lecture of building simulators from data or from physics, which this lecture reuses for obtaining the matrices A and B.

Check your understanding

  1. How does adding state-action rewards change the placement of the max in Bellman's equation, and why?
  2. Why is the optimal policy in a finite-horizon MDP generally non-stationary, while the optimal policy in the earlier discounted, infinite-horizon MDPs was stationary?
  3. What two assumptions does the LQR framework require, and what do they buy you in terms of solving for the optimal policy exactly?
  4. What are the two methods described for obtaining the matrices A and B in a linear dynamical system model?
  5. Why does the optimal LQR policy not depend on the covariance of the noise term, and does this property generalize to other reinforcement learning settings?

Vocabulary

generalize (verb)
To make an idea apply more broadly than before.
The lecture generalizes the reward function to depend on both state and action.
penalize (verb)
To give a lower score or reward as punishment for something undesirable.
We penalize aggressive control-stick movements on a helicopter.
aggressive (adjective)
Forceful or sudden, more than needed.
Aggressive actions cost more reward than smooth ones.
finite horizon (noun)
A fixed, limited number of time steps after which a process ends.
A finite-horizon MDP runs for exactly T steps and then stops.
undiscounted (adjective)
Not reduced in value for happening later in time.
The finite-horizon payoff is an undiscounted sum of rewards.
non-stationary (adjective)
Changing over time rather than staying fixed.
The optimal policy in a finite-horizon MDP is non-stationary.
stationary (adjective)
Staying the same and not depending on time.
Earlier lectures used a stationary, time-independent policy.
dynamic programming (noun)
A method that solves a problem by breaking it into smaller steps solved in order, reusing earlier results.
Dynamic programming computes V*_t backward from the final step.
backward recursion (noun)
A calculation done step by step starting from the end and working toward the beginning.
The value function is computed by a backward recursion from T to 0.
linear quadratic regulator (LQR) (noun)
A control method for systems with linear dynamics and a cost that is a squared (quadratic) function.
LQR gives an exact solution for the optimal policy.
quadratic (adjective)
Involving a squared term, like x squared, in a formula.
The reward is a negative quadratic cost in state and action.
positive semidefinite (adjective)
A property of a matrix meaning the quadratic form it defines is never negative.
U and V must be positive semidefinite so the reward stays non-positive.
Gaussian noise (noun)
Random variation that follows a bell-shaped (normal) probability distribution.
The state transition includes Gaussian noise w_t.
covariance (noun)
A measure of how much two random quantities vary together, or here, the spread of noise.
The optimal LQR policy does not depend on the noise covariance.
linearize (verb)
To approximate a curved or complex function with a straight-line (linear) one near a point.
You can linearize the helicopter's physics model around a hovering state.
Taylor approximation (noun)
A way to approximate a function near a point using its value and slope there.
A first-order Taylor approximation linearizes the nonlinear dynamics.
operating point (noun)
The typical state and action a system runs around, used as a reference for approximation.
The model is linearized around the hovering operating point.
hover (verb)
To stay in one place in the air without moving forward or falling.
The helicopter needs to hover steadily near zero velocity.
recursion (noun)
A process that repeats by applying the same rule to its own previous result.
The LQR solution comes from a backward recursion over time steps.
derive (verb)
To work out a result step by step from known facts or rules.
The lecture derives the exact form of the optimal policy.
matrix (noun)
A rectangular grid of numbers used in linear algebra.
A and B are matrices describing the linear dynamics.
identity matrix (noun)
A square matrix with ones on the diagonal and zeros elsewhere, representing no change.
Setting U and V to identity matrices penalizes any deviation equally.
regression (noun)
A statistical method for finding the relationship between inputs and an output.
The matrices A and B can be fit by regression on flight data.
parameterize (verb)
To describe something using a set of adjustable numbers.
The value function is parameterized by a matrix and a scalar.
scalar (noun)
A single number, as opposed to a vector or matrix.
Psi is a scalar term in the quadratic value function.

Chapters

From the YouTube description

For more information about Stanford’s Artificial Intelligence professional and graduate programs, visit: https://stanford.io/ai

Andrew Ng
Adjunct Professor of Computer Science
https://www.andrewng.org/

To follow along with the course schedule and syllabus, visit:
http://cs229.stanford.edu/syllabus-autumn2018.html

← Lecture 18: Continuous-State MDPs and Fitted Value Iteration · Lecture 20: RL Debugging, Diagnostics, and Policy Search →