Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Probability · Lecture 47 of 76 · 10:10

Using the Conditional Expectation and Variance

Using the Conditional Expectation and Variance on YouTube

Study guide

What this lecture covers

This is a worked recitation problem applying the law of total variance to a joint density. X and Y are uniformly distributed over a parallelogram, and the problem finds the variance of X + Y. It demonstrates a choice that matters in practice: which variable to condition on, and why picking the one that keeps conditional slices a constant width makes the calculation far easier.

After watching, you should be able to read a joint density defined on a region, choose a sensible conditioning variable based on the region's geometry, derive conditional distributions from geometric slices, and combine the resulting pieces using the law of total variance.

Key ideas

  • Why condition on X and not Y: slicing the parallelogram at a fixed x always gives an interval of the same width, so the conditional distribution of Y given X is uniform with a constant-width shift; slicing at a fixed y gives varying widths, which would be harder to work with.
  • Conditional distribution from geometry: for x in [0, 1], Y given X = x is uniform between x and x + 1, since conditioning on a uniform joint density preserves uniformity within the conditional region.
  • Conditional expectation as a linear shift: E[X + Y | X] = x + E[Y|X] = x + (midpoint of the interval) = 2x + 1/2.
  • Conditional variance ignores the constant shift: Var(X + Y | X) = Var(Y|X) = 1/12, using the uniform-variance formula (b - a)^2 / 12.
  • Law of total variance combines both pieces: Var(X + Y) = Var(2X + 1/2) + E[1/12] = 4*Var(X) + 1/12.
  • Marginalizing the joint density: integrating the joint density over y at each x shows the marginal density of X is uniform on [0, 1], giving Var(X) = 1/12 and a final answer of 5/12.

Before you watch

  • Watch the main lecture on iterated expectations for the law of total variance.
  • Be comfortable computing marginal densities from a joint density, and know the variance formula for a uniform random variable.

Check your understanding

  1. Why does conditioning on X, rather than Y, simplify this problem given the shape of the parallelogram?
  2. Why does the constant term inside a conditional expectation not affect the conditional variance?
  3. Derive the marginal density of X by marginalizing the joint density over y.
  4. Walk through both terms of the law of total variance calculation for Var(X + Y).

Vocabulary

parallelogram (noun)
A four-sided shape with two pairs of parallel sides.
X and Y are uniform over a parallelogram.
slice (noun)
A thin cut or cross-section taken through a shape.
A slice at fixed x always has the same width.
constant width (phrase)
A width that stays the same, no matter where it is measured.
The parallelogram has slices of constant width when sliced at fixed x.
linear shift (phrase)
A movement of a value by adding or subtracting a fixed related amount.
The conditional expectation is a linear shift depending on x.
marginalize (verb)
To find one variable's distribution by summing or integrating out another variable.
We marginalize the joint density over y to get the density of X.
geometry (noun)
The shape and arrangement of a figure or region.
The geometry of the parallelogram decides which variable to condition on.
joint density (phrase)
A function describing the probability pattern of two variables together.
X and Y follow a joint density that is uniform over the parallelogram.
uniform distribution (phrase)
A distribution where every value in a range is equally likely.
Y given X follows a uniform distribution over an interval of width 1.
conditional expectation (phrase)
The average value of one variable, calculated once another variable's value is known.
The conditional expectation of X + Y given X is a linear shift.
conditional variance (phrase)
The spread of one variable's values, calculated once another variable's value is known.
The conditional variance ignores the constant shift in the expectation.
law of total variance (phrase)
A rule that splits an overall variance into a part from averages and a part from spread within each case.
The law of total variance combines both pieces to give the final answer.
midpoint (noun)
The point exactly halfway between two values.
The conditional expectation uses the midpoint of the interval.
integrate (verb)
To sum up continuously over a range to find a total or average.
We integrate the joint density over y to find the marginal density of X.
interval (noun)
A range of values between two numbers.
Y given X is uniform over an interval from x to x+1.
in practice (phrase)
In real, actual use, as opposed to just in theory.
The choice of conditioning variable matters a lot in practice.
sensible (adjective)
Reasonable and practical, showing good judgment.
Choosing a sensible conditioning variable makes the problem much easier.
derive (verb)
To work out a result step by step from known facts or rules.
We derive the conditional distribution from the region's geometry.
combine (verb)
To bring separate parts together into one result.
The final step combines the two pieces of the total variance formula.
width (noun)
The distance across something, from one side to the other.
Slicing at fixed x always gives an interval of the same width.
straightforward (adjective)
Simple and easy to follow, without unnecessary complication.
Conditioning on X makes the calculation straightforward.
calculate (verb)
To work out a numerical answer using math.
We calculate the variance of X + Y using the law of total variance.
choice (noun)
A decision between two or more options.
The choice of which variable to condition on affects how hard the problem is.
matter (verb)
To be important or to make a real difference.
Which variable you condition on matters a great deal here.
constant (adjective)
Staying the same and not changing.
A constant shift inside an expectation does not affect the variance.
formula (noun)
A fixed rule expressed with symbols for calculating something.
The uniform-variance formula gives (b-a)^2/12.

Chapters

From the YouTube description

MIT 6.041SC Probabilistic Systems Analysis and Applied Probability, Fall 2013
View the complete course: http://ocw.mit.edu/6-041SCF13
Instructor: Katie Szeto

License: Creative Commons BY-NC-SA
More information at http://ocw.mit.edu/terms
More courses at http://ocw.mit.edu

← Widgets and Crates · A Random Number of Coin Flips →