Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Machine Learning · Lecture 18 of 21 · 1:19:14

Lecture 17: MDPs, Value Iteration and Policy Iteration

Lecture 17 - MDPs & Value/Policy Iteration | Stanford CS229: Machine Learning Andrew Ng (Autumn2018) on YouTube

Study guide

What this lecture covers

This lecture answers the question left open in the previous session: given a Markov decision process, how do you actually compute the optimal policy? It builds up the machinery step by step, defining the value function for a fixed policy, deriving Bellman's equation for it, defining the optimal value function, and then presenting two algorithms, value iteration and policy iteration, that compute it.

The lecture then moves from the idealized setting, where the state transition probabilities are known, to the practical one, where they must be estimated from a robot's experience. It closes with a well-known difficulty in reinforcement learning: the tension between exploiting what the agent already believes is the best policy and exploring to discover better options, illustrated with epsilon-greedy exploration and a real-world example from online advertising. After watching, you should be able to derive Bellman's equation, implement value iteration or policy iteration for a small discrete-state MDP, and explain why an exploration strategy is needed on top of a greedy policy.

Key ideas

  • Value function V_pi: the expected total discounted payoff of starting in a state and following policy pi from then on.
  • Bellman's equation: V_pi(s) = R(s) + gamma * sum_s' P(s, pi(s), s') V_pi(s'), expressing a state's value as its immediate reward plus the discounted expected value of the next state.
  • Solving for V_pi directly: for a fixed policy and a finite number of states, Bellman's equation is a linear system with one equation and one unknown per state, solvable exactly with a linear algebra solver.
  • Optimal value function V*: V*(s) = max_pi V_pi(s), and it satisfies its own Bellman equation with a max over actions instead of a fixed policy.
  • Value iteration: repeatedly applying the optimal-value Bellman equation to an initial guess (often all zeros) converges geometrically fast to V*, after which the optimal action at each state is the argmax over actions of the equation's right-hand side.
  • Policy iteration: alternates between solving the linear system for the current policy's value function and updating the policy to be greedy with respect to that value function, converging in a finite number of steps to the exact optimal policy.
  • Estimating transition probabilities: when P(s, a) is unknown, it is estimated from observed frequencies of ending up in each next state after taking each action, with a simple heuristic (uniform probability) for state-action pairs never observed.
  • Exploration versus exploitation: a purely greedy policy can get stuck exploiting a known reward and never discover a better one; epsilon-greedy exploration takes a random action with small probability epsilon to keep exploring the state space.

Walkthrough

Recap: MDPs and the roadmap for finding an optimal policy (2:49)

The lecture recaps the five-tuple MDP definition from before and the 11-state grid-world example, then frames the day's goal: since the number of possible policies grows exponentially with the number of states, brute-force search is infeasible, so the lecture will define the value function V_pi, the optimal value function V*, and the optimal policy Pi* in order, then derive an efficient way to compute them.

The value function and Bellman's equation (11:38)

V_pi(s) is defined as the expected total discounted payoff of starting in state s and executing policy pi. The lecture derives Bellman's equation by splitting this expectation into the immediate reward received on waking up in s plus the discounted expected value of whatever state comes next, arriving at V_pi(s) = R(s) + gamma * E[V_pi(s')]. For a small MDP this turns into a linear system: with 11 states there are 11 equations in 11 unknowns, solvable directly with a linear algebra solver instead of iterating.

Defining the optimal value function (22:44)

The lecture defines V*(s) as the maximum of V_pi(s) over every possible policy, and derives a second Bellman equation for it that replaces the fixed policy with a max over actions: the best expected payoff from state s is the immediate reward plus the discounted value of the best next state reachable by any action. This equation immediately gives a recipe for the optimal action at each state, Pi*(s) = argmax_a of that same expression, once V* is known.

Value iteration (36:17)

Value iteration initializes an estimate of V to zero for every state and repeatedly applies the optimal-value Bellman equation as an update rule, either synchronously (updating all states at once from the previous estimate) or asynchronously (updating states one at a time using freshly updated values). Both variants converge, and with a discount factor like 0.99 the error shrinks geometrically each iteration, so a few hundred iterations bring the estimate very close to V*. The lecture works a concrete example, computing the expected future reward of moving west versus north from a given state using the converged V*, and shows west wins.

Policy iteration (53:04)

Policy iteration takes the opposite approach: it initializes a policy randomly, solves the linear system for that policy's exact value function, then updates the policy greedily with respect to that value function, and repeats. Unlike value iteration, which approaches V* asymptotically without ever reaching it exactly, policy iteration converges to the exact optimal policy in a finite number of iterations. It is efficient for small state spaces where solving the linear system is cheap, but value iteration scales better to large state spaces because it avoids repeatedly inverting a large system.

Estimating unknown transitions and exploration versus exploitation (59:00)

For a real robot, the state transition probabilities are usually unknown and must be estimated from observed frequencies of outcomes after taking each action in each state, with unseen state-action pairs defaulting to a uniform distribution over next states; the lecture notes that MDP solvers, unlike naive Bayes, are not very sensitive to these estimates being imperfect. The resulting loop, act, estimate transitions, solve for a new policy, repeat, converges to a good policy but has a blind spot: a purely greedy policy can settle on a known reward and never discover a better one elsewhere. This exploration-versus-exploitation trade-off, illustrated with online advertising as well as robotics, motivates epsilon-greedy exploration, where the agent follows its current policy most of the time but takes a random action with probability epsilon to keep discovering the state space; the lecture notes this always converges to the optimal policy for a discrete-state MDP, though possibly slowly.

Before you watch

  • Review the MDP five-tuple (states, actions, state transition probabilities, discount factor, reward function) and the grid-world example from the previous lecture in this course.
  • Be comfortable with expectations of random variables and with solving small linear systems of equations.
  • Recall how gradient descent converges asymptotically, since the lecture draws a direct comparison to how value iteration approaches V*.

Check your understanding

  1. How does Bellman's equation let you set up and solve for V_pi as a linear system rather than iterating?
  2. What is the difference between the Bellman equation for V_pi and the one for V*, and how does the second one give you the optimal action directly?
  3. Why does policy iteration reach the exact optimal policy in finitely many steps while value iteration only approaches V* in the limit?
  4. When state transition probabilities are unknown, how are they estimated from a robot's experience, and what happens for state-action pairs never observed?
  5. Why can a purely greedy policy fail to find the best available reward, and how does epsilon-greedy exploration address this?

Vocabulary

roadmap (noun)
A plan showing the steps you will take to reach a goal.
The lecture sets out a roadmap for finding the optimal policy.
brute-force (adjective)
Solving a problem by trying every possible option instead of using a clever method.
Brute-force search over policies is infeasible because there are too many of them.
infeasible (adjective)
Not possible to do in practice.
Checking every policy by hand is infeasible for large state spaces.
Markov decision process (MDP) (noun)
A mathematical model of an agent choosing actions in states to get rewards over time.
The grid world is described as an 11-state MDP.
policy (noun)
A rule that tells an agent which action to take in each state.
We want to find the policy that gives the most reward.
value function (noun)
A function that gives the expected total reward from a state if you follow a certain policy.
V_pi(s) is the value function for policy pi.
expected (adjective)
Describing an average result you would get if something were repeated many times.
The expected payoff accounts for randomness in the outcome.
payoff (noun)
The total reward or benefit received.
We want to maximize the total discounted payoff.
discount factor (noun)
A number less than 1 that makes future rewards worth less than immediate ones.
With a discount factor of 0.99, far-off rewards matter less.
Bellman's equation (noun)
An equation that expresses a state's value as its immediate reward plus the discounted value of the next state.
Bellman's equation lets us solve for V_pi directly.
linear system (noun)
A set of equations where every unknown appears only multiplied by a constant, no powers or products.
With 11 states, Bellman's equation becomes an 11-equation linear system.
solver (noun)
A program or method that finds the answer to a set of equations.
A linear algebra solver finds V_pi exactly.
optimal (adjective)
The best possible, giving the highest result.
V* is the optimal value function.
argmax (noun)
The input that produces the largest output of a function.
The best action is the argmax over actions of the Bellman expression.
value iteration (noun)
A method that repeats an update rule on value estimates until they get close to the true optimal values.
Value iteration starts from zero and slowly approaches V*.
converge (verb)
To gradually get closer and closer to a fixed result.
The value estimates converge to V* after enough iterations.
geometrically (adverb)
In a way where the error shrinks by a constant fraction each step, very quickly.
The error in value iteration shrinks geometrically with each update.
synchronously (adverb)
Doing all updates at the same time, using only old values.
Synchronously, every state is updated at once from the previous estimate.
asynchronously (adverb)
Doing updates one at a time, sometimes using values that were already updated.
Asynchronously, later states can use freshly updated earlier values.
policy iteration (noun)
A method that alternates between computing a policy's exact value and improving the policy based on that value.
Policy iteration converges to the exact optimal policy in finitely many steps.
greedy (adjective)
Always choosing the action that looks best right now, without considering exploration.
The updated policy is greedy with respect to the current value function.
finite (adjective)
Having a limited, countable number, not endless.
Policy iteration converges in a finite number of steps.
asymptotically (adverb)
Getting closer and closer to a value without ever exactly reaching it.
Value iteration approaches V* asymptotically.
state transition probability (noun)
The chance of moving to a particular next state after taking an action.
State transition probabilities describe how the robot's state changes.
frequency (noun)
How often something happens, counted from observations.
We estimate transitions from the observed frequency of each outcome.
heuristic (noun)
A simple, practical rule that works reasonably well even if not exact.
A uniform-probability heuristic handles state-action pairs never observed.
uniform distribution (noun)
A probability spread equally across all possible outcomes.
Unseen transitions default to a uniform distribution over next states.
naive Bayes (noun)
A simple classification method that assumes features are independent given the class.
MDP solvers are less sensitive to estimation errors than naive Bayes is.
sensitive (adjective)
Reacting a lot to small changes or errors.
The algorithm is not very sensitive to small errors in the transition estimates.
exploration (noun)
Trying new or uncertain actions to learn more about the environment.
Exploration helps the agent discover rewards it hasn't found yet.
exploitation (noun)
Using what you already know to get the best result right now.
Exploitation means always picking the currently best-known action.
trade-off (noun)
A balance between two things where gaining one means losing some of the other.
The exploration-exploitation trade-off affects how fast the agent finds the best policy.
epsilon-greedy (adjective)
Describing a strategy that usually picks the best-known action but sometimes picks a random one.
Epsilon-greedy exploration takes a random action with small probability epsilon.
blind spot (noun)
A weakness or area a method fails to notice or handle well.
A purely greedy policy has a blind spot: it may never explore better options.
get stuck (phrasal verb)
To become unable to make progress or change.
A greedy policy can get stuck exploiting a known but weaker reward.

Chapters

From the YouTube description

For more information about Stanford’s Artificial Intelligence professional and graduate programs, visit: https://stanford.io/ai

Andrew Ng
Adjunct Professor of Computer Science
https://www.andrewng.org/

To follow along with the course schedule and syllabus, visit:
http://cs229.stanford.edu/syllabus-autumn2018.html

← Lecture 16: Independent Component Analysis & Reinforcement Learning · Lecture 18: Continuous-State MDPs and Fitted Value Iteration →