Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Deep Reinforcement Learning · Lecture 26 of 99 · 16:16
Lecture 7, Part 1: From Actor-Critic to Policy Iteration
Study guide
What this lecture covers
Opening Berkeley CS285's lecture on value function methods, this part asks whether the explicit policy network used in actor-critic is even necessary. It shows that taking the action that maximizes the advantage function at every state produces a policy at least as good as the one that advantage was computed for, no matter how bad the original policy was.
You come away understanding the idea behind policy iteration, an algorithm that alternates between evaluating a value function and setting a new implicit policy via an argmax, and how this reduces to a dynamic-programming procedure in the special case of a small, tabular MDP with known transition probabilities.
Key ideas
- Implicit policy via argmax: taking
argmax_a A^pi(s,a)gives an action at least as good as one sampled frompi, for any policypi, which means acting greedily on the advantage improves the policy without needing an explicit policy network. - Policy iteration: alternates between evaluating the advantage (or value function) of the current policy and constructing a new deterministic policy that assigns probability 1 to the argmax action.
- Tabular setting: when the state and action spaces are small and discrete and the transition probabilities are known, the value function can be stored explicitly as a table rather than approximated with a neural network.
- Policy evaluation by dynamic programming: repeatedly applying the Bellman backup
V(s) = r(s, pi(s)) + gamma * E[V(s')]across the whole table converges to the true value functionV^pi. - Value iteration: a simplification that skips explicit policy representation by building a table of Q values (
Q(s,a) = r(s,a) + gamma * E[V(s')]) and then settingV(s)to the max over actions in each row, which is equivalent to the argmax-based update.
Before you watch
- Watch the actor-critic lecture (Lecture 6) in full, since this lecture opens by directly extending its batch actor-critic algorithm.
- Be comfortable with the definitions of the advantage, value, and Q functions from that lecture.
Check your understanding
- Why does
argmax_a A^pi(s,a)always produce an action at least as good as sampling frompi, regardless of how goodpiis? - What are the two steps of policy iteration, and what does each one compute?
- Why does the tabular setting allow the value function to be stored exactly rather than approximated?
- How does value iteration simplify policy iteration by working with the Q function instead of alternating value and policy updates?
Vocabulary
- policy iteration (noun)
- An algorithm that repeatedly evaluates a policy and then improves it by choosing the best action at each state.
Policy iteration alternates between evaluation and improvement. - advantage function (noun)
- A function that measures how much better an action is than the average action at that state.
Taking the argmax of the advantage function improves the policy. - argmax (noun)
- The input that gives the largest output of a function.
The best action is the argmax of the advantage function. - implicit policy (phrase)
- A policy defined indirectly through a value function, rather than as its own separate network.
Acting greedily on the advantage gives an implicit policy. - explicit (adjective)
- Directly stated or represented, not derived from something else.
Actor-critic uses an explicit policy network. - deterministic (adjective)
- Always producing the same result given the same input, with no randomness.
Policy iteration produces a deterministic policy after each update. - tabular (adjective)
- Stored as a table with one entry per state, rather than approximated by a model.
In the tabular setting, the value function is stored exactly. - transition probability (phrase)
- The chance of moving from one state to another after taking an action.
Dynamic programming needs known transition probabilities. - dynamic programming (noun)
- A method that solves a problem by breaking it into smaller repeated sub-steps and reusing results.
Policy evaluation is done here by dynamic programming. - Bellman backup (noun)
- An update rule that recalculates a state's value from the reward and the value of the next state.
Repeating the Bellman backup converges to the true value function. - converge (verb)
- To gradually approach and settle at a final, stable answer.
This procedure converges to the true value function. - value iteration (noun)
- An algorithm that repeatedly updates a table of Q-values and takes the best value at each state.
Value iteration skips explicitly storing a separate policy. - Q function (noun)
- A function that estimates the expected future reward of taking a specific action in a specific state.
Value iteration builds a table of Q function values. - greedy (adjective)
- Always choosing the option that looks best right now.
The new policy acts greedily with respect to the value function. - sample (verb)
- To draw one example randomly, following some probability distribution.
In actor-critic, actions are sampled from the policy. - regardless of (phrase)
- Without being affected by something; no matter what.
The improvement holds regardless of how bad the original policy is. - batch actor-critic (phrase)
- An actor-critic method that collects a full batch of data before updating.
This lecture extends the batch actor-critic algorithm. - extend (verb)
- To build further on an existing idea or method.
This lecture extends actor-critic toward policy iteration. - small and discrete (phrase)
- Having a limited, countable number of possible values.
A small and discrete state space allows exact tables. - equivalent (adjective)
- Having the same effect or meaning, even if written differently.
Value iteration is equivalent to the argmax-based update.
Chapters
- 0:00 Recap: actor-critic
- 1:34 Can we omit policy gradient completely?
- 5:00 Policy iteration High level idea
- 10:09 Policy iteration with dynamic programming
- 12:35 Even simpler dynamic programming
← Lecture 6, Part 5: Actor-Critic Summary and Examples · Lecture 7, Part 2: Fitted Value and Fitted Q-Iteration →
