Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Probability · Lecture 73 of 76 · 27:50

An Inference Example

An Inference Example on YouTube

Study guide

What this lecture covers

This review problem starts from a joint PDF of two random variables X and Y that is constant (2/3) over an irregular region and works through several estimation techniques applied to it. The lecture shows how to read a conditional expectation directly off a picture of the joint density when the density is piecewise uniform, then formalizes the answer algebraically.

Building on that estimator, the lecture computes its conditional and unconditional mean squared error, marginalizes the joint PDF to find the distribution of X, derives the linear LMS estimator for comparison, and explains why a MAP estimator breaks down in this particular setting. It is a consolidation lecture, pulling together tools of conditional expectation, variance, and estimation from earlier in the course.

Key ideas

  • Reading conditional expectation from a picture: when a joint PDF is piecewise uniform, slicing the region at a fixed X value gives a uniform conditional distribution of Y, whose expectation is just the midpoint of the slice.
  • Piecewise LMS estimator: here the LMS estimator of Y given X is X/2 for X between 0 and 1, and X - 1/2 for X between 1 and 2, with a kink at X = 1.
  • Conditional variance equals conditional MSE for the LMS estimator: because the LMS estimator is the conditional expectation, the conditional mean squared error reduces to the conditional variance of Y given X.
  • Marginalizing a joint PDF: integrating the joint density over Y for each region of X gives the marginal PDF of X, needed to compute unconditional expectations.
  • Law of iterated expectations for MSE: the overall mean squared error equals the expectation, over X, of the conditional mean squared error.
  • Linear LMS estimator: computed from E[Y], E[X], Var(X), and Cov(X,Y), it approximates the true (kinked) LMS estimator with a single straight line, at the cost of allowing impossible negative estimates of Y.
  • MAP estimator can be ill-defined: when the conditional distribution of Y given X is uniform, every value of Y has equal posterior density, so there is no unique maximum and no sensible MAP estimate.

Walkthrough

Reading the joint PDF and finding the LMS estimator visually (0:00)

The lecture presents a joint PDF of X and Y that is a constant 2/3 over a bounded region, and shows that because it is uniform, the conditional expectation of Y given X can be found visually as the midpoint of each vertical slice, giving a piecewise linear function with a kink at X = 1.

Computing the conditional mean squared error (4:08)

The lecture defines conditional MSE and shows it equals the conditional variance of Y given X, since the LMS estimator is the conditional mean. Using the uniform-distribution variance formula (width squared over 12), it computes this variance for both regions of X.

Marginalizing to find the distribution of X and the overall MSE (8:13)

To get the unconditional mean squared error, the lecture marginalizes the joint PDF over Y to find f_X(x), piecewise 2/3 x and 2/3, then applies the law of iterated expectations to combine the conditional MSE with this marginal density, arriving at a final MSE of 5/72.

Deriving the linear LMS estimator (13:27)

The lecture computes E[X], Var(X), E[Y] (via iterated expectations using the earlier LMS result), and Cov(X,Y) (via a double integral over the joint PDF), then assembles the linear LMS formula to get an explicit straight-line estimator.

Comparing estimators and their mean squared errors (21:43)

Plotting the linear LMS estimator against the LMS estimator shows they are close but the linear version lacks the kink and can dip below zero for small X, an impossible value for Y. The lecture explains that the LMS estimator must have the smallest MSE by construction, so the linear LMS estimator's MSE can only be equal or larger.

Why the MAP estimator fails here (25:50)

Because each conditional slice of Y given X is uniform, every value of Y in the slice has the same conditional density, so there is no single maximizing value and the MAP rule gives no unique estimate.

Before you watch

  • Review conditional expectation and conditional variance for jointly distributed random variables.
  • Be familiar with the formulas for LMS and linear LMS estimators and how covariance and variance are computed from a joint PDF.
  • Recall how to marginalize a joint PDF to find a single variable's distribution.

Check your understanding

  1. Why can the conditional expectation of Y given X be read directly off a picture when the joint PDF is piecewise uniform?
  2. Why does the conditional MSE of the LMS estimator equal the conditional variance of Y given X?
  3. How is the law of iterated expectations used to go from conditional MSE to overall MSE?
  4. Why does forcing the estimator to be linear allow it to produce impossible values, and why is that acceptable in some situations?
  5. Why is the MAP estimator undefined in this problem, and what does that reveal about when MAP is a useful estimation method?

Vocabulary

joint PDF (noun)
A function giving the relative likelihood of two random variables taking specific values together.
The joint PDF of X and Y is constant over an irregular region.
piecewise uniform (adjective)
Constant within separate sections, but possibly different between sections.
The joint density is piecewise uniform over the region.
slice (noun)
A thin cross-section of a region taken at one fixed value.
We take a vertical slice at a fixed X to find the conditional distribution of Y.
kink (noun)
A sudden change in direction or slope in a graph or function.
The LMS estimator has a kink at X = 1.
conditional variance (noun)
The variance of a random variable, calculated using extra known information.
The conditional variance of Y given X gives the conditional MSE.
marginalize (verb)
To find the distribution of one variable by summing or integrating out the other variable from a joint distribution.
We marginalize the joint PDF over Y to find the distribution of X.
marginal PDF (noun)
The probability density of a single variable, found by integrating a joint density over the other variable.
The marginal PDF of X is piecewise linear.
double integral (noun)
An integral computed over two variables at once, often used to find area or expected values over a region.
We compute the covariance using a double integral over the joint PDF.
region (noun)
An area in a graph or plane defined by certain boundaries.
The joint PDF is nonzero only over a bounded region.
dip below zero (phrase)
To fall to a value less than zero.
The linear estimator can dip below zero for small X.
ill-defined (adjective)
Not clearly or uniquely specified.
The MAP estimator is ill-defined when the posterior is flat.
maximizing value (phrase)
The input that produces the largest possible output of a function.
There is no unique maximizing value of Y in a flat posterior.
consolidation (lecture) (noun)
A review session that brings together several earlier topics into one worked example.
This is a consolidation lecture pulling together estimation tools.
irregular (region) (adjective)
Not having a simple, regular shape.
The joint PDF is defined over an irregular region.
bounded (adjective)
Having limits that keep values within a fixed range.
The joint PDF is nonzero only over a bounded area.
algebraically (adverb)
Using symbols and equations rather than pictures or intuition.
The estimator is formalized algebraically after being read off a picture.
reduce to (phrasal verb)
To simplify down to a more basic form.
The conditional MSE reduces to the conditional variance.
width (noun)
The size of an interval from one end to the other.
The variance formula uses the width of the interval squared over 12.
by construction (phrase)
True automatically because of how something was defined or built.
The LMS estimator has the lowest MSE by construction.
visually (adverb)
By looking at a picture or graph, rather than by calculation.
The estimator can be found visually from the joint PDF.

Chapters

From the YouTube description

MIT 6.041SC Probabilistic Systems Analysis and Applied Probability, Fall 2013
View the complete course: http://ocw.mit.edu/6-041SCF13
Instructor: Jimmy Li

License: Creative Commons BY-NC-SA
More information at http://ocw.mit.edu/terms
More courses at http://ocw.mit.edu

← Inferring a Parameter of the Uniform Distribution, Part 2 · Lecture 23: Classical Statistical Inference I →