Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Deep Reinforcement Learning · Lecture 74 of 99 · 19:36

Lecture 18, Variational Inference, Part 2

CS 285: Lecture 18, Variational Inference, Part 2 on YouTube

Study guide

What this lecture covers

This part derives the core tool of variational inference: a tractable lower bound on log p(x_i), built by approximating the true posterior p(z|x_i) with a simple distribution q_i(z). It shows, via Jensen's inequality, how this bound arises, and then, via KL divergence, why maximizing the bound with respect to q_i tightens it and justifies the expected log likelihood objective introduced in part 1.

By the end, you should be able to derive the evidence lower bound (ELBO), explain the role entropy plays in keeping q_i from collapsing onto a single point, and describe the resulting alternating optimization: maximize the ELBO with respect to q_i to tighten the bound, then with respect to the model parameters to raise the likelihood.

Key ideas

  • Jensen's inequality: for a concave function like the logarithm, log(E[y]) >= E[log(y)], which lets the log of an intractable expectation be lower-bounded by a tractable expectation of a log.
  • Evidence lower bound (ELBO): log p(x_i) >= E_{q_i(z)}[log p(x_i|z) + log p(z)] + H(q_i), tractable because it only needs samples from q_i(z) and, for distributions like Gaussians, a closed-form entropy.
  • Entropy's role: the entropy term keeps q_i(z) spread out rather than collapsing onto the single most likely point, so q_i tends to cover the true posterior rather than just its peak.
  • KL divergence: measures how different two distributions are; KL(q_i(z) || p(z|x_i)) is zero exactly when q_i equals the true posterior.
  • Exact identity: log p(x_i) = KL(q_i(z) || p(z|x_i)) + ELBO, which shows the ELBO is always a valid lower bound (since KL divergence is non-negative) and that minimizing the KL divergence is equivalent to maximizing the ELBO with respect to q_i.
  • Alternating optimization: maximize the ELBO with respect to q_i to tighten the bound, then with respect to the model parameters theta to improve the likelihood, and repeat.
  • Scalability problem: fitting a separate q_i (its own mean and variance) for every data point becomes impractical with millions of data points, motivating the move to amortized inference.

Before you watch

  • Watch part 1 of this lecture: it introduces latent variable models and the expected log likelihood objective that this part justifies.
  • Some familiarity with entropy and basic information theory helps, though the lecture recaps both entropy and KL divergence from first principles.

Check your understanding

  1. How does Jensen's inequality turn the intractable log p(x_i) into a tractable lower bound?
  2. Why does maximizing the ELBO with respect to q_i alone not just collapse q_i onto the single highest-probability value of z?
  3. What does the identity log p(x_i) = KL(q_i(z) || p(z|x_i)) + ELBO tell you about when the ELBO equals the true log likelihood?
  4. Why does fitting a separate Gaussian q_i per data point become impractical at scale, and what alternative does the lecture point toward?

Vocabulary

lower bound (noun)
A value that a quantity is guaranteed not to fall below.
The evidence lower bound is always below the true log likelihood.
Jensen's inequality (noun)
A math rule saying the log of an average is at least the average of the logs.
Jensen's inequality justifies the lower bound.
concave (adjective)
Curving downward, like the top of a hill, on a graph.
The logarithm is a concave function.
evidence lower bound (noun)
A tractable value that is always less than or equal to the true log probability of the data.
We maximize the evidence lower bound instead of the true likelihood.
entropy (noun)
A measure of how spread out or uncertain a probability distribution is.
The entropy term keeps the distribution spread out.
collapse (verb)
To shrink down to a single point or a very narrow range.
Without entropy, the distribution could collapse onto one value.
KL divergence (noun)
A measure of how different two probability distributions are from each other.
KL divergence is zero only when the two distributions match exactly.
closed-form (adjective)
Expressible as an exact formula, without needing further approximation.
Entropy has a closed-form expression for Gaussians.
alternating optimization (noun)
A method that improves two things in turn, one at a time.
Training uses alternating optimization between q and theta.
scalability (noun)
The ability of a method to keep working well as size or amount grows.
Fitting a separate distribution per point creates a scalability problem.
amortized (adjective)
Sharing a one-time cost across many uses instead of repeating it each time.
Amortized inference avoids fitting a new distribution per data point.
identity (noun)
A mathematical equation that is always exactly true.
The identity relates the log likelihood to the KL divergence and the bound.
derive (verb)
To work out a result step by step from known facts.
This part derives the evidence lower bound using Jensen's inequality.
arise (verb)
To come into existence as a result of something else.
The lower bound arises directly from Jensen's inequality.
tighten (verb)
To make a bound closer to the true value it is approximating.
Maximizing the ELBO with respect to q_i tightens the bound.
non-negative (adjective)
Equal to or greater than zero.
KL divergence is always non-negative.
peak (noun)
The single highest point of a function or distribution.
Without entropy, q_i could collapse onto the posterior's peak.
information theory (noun)
The mathematical study of measuring and transmitting information.
Entropy and KL divergence come from information theory.
first principles (phrase)
The most basic facts, used to build up an explanation from the ground up.
The lecture recaps entropy from first principles.
impractical (adjective)
Not sensible or workable to do in real conditions.
Fitting a separate q_i per data point is impractical at scale.
motivate (verb)
To give a reason that leads to a particular choice.
The scalability problem motivates amortized inference.
cover (verb)
To include or account for a wide range of possibilities.
A good q_i should cover the true posterior, not just its peak.
spread out (phrase)
Distributed over a wide range rather than concentrated in one place.
The entropy term keeps q_i spread out.
justify (verb)
To give a good reason showing something is correct.
The identity justifies using the expected log likelihood objective.
intuitively (adverb)
In a way that feels natural to understand, without needing a full proof.
Entropy intuitively keeps the approximation from collapsing.

Chapters

← Lecture 18, Variational Inference, Part 1 · Lecture 18, Variational Inference, Part 3 →