Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Probability · Lecture 20 of 76 · 50:52

Lecture 6: Discrete Random Variables II

6. Discrete Random Variables II on YouTube

Study guide

What this lecture covers

This lecture continues the study of discrete random variables, reviewing expected value and the expected value rule before introducing variance, standard deviation, conditional PMFs and conditional expectation, the memorylessness property of the geometric random variable, the total expectation theorem, and joint PMFs of two random variables.

After watching, you should be able to compute and interpret variance and standard deviation, explain why expectation does not commute with nonlinear functions, compute conditional PMFs and conditional expectations, use the memorylessness property and total expectation theorem to derive the expected value of a geometric random variable without heavy algebra, and read joint, marginal, and conditional PMFs from a table.

Key ideas

  • Variance and standard deviation: variance Var(X) = E[(X - E[X])^2] measures the average squared distance from the mean; standard deviation is its square root, restoring the original units of X.
  • Nonlinearity breaks averaging shortcuts: E[g(X)] generally does not equal g(E[X]) for nonlinear g, and E[XY] generally does not equal E[X]E[Y], illustrated with a speed/time example where average speed times average time does not equal average distance.
  • Conditional PMF and conditional expectation: conditioning on an event rescales the probabilities of the remaining outcomes to sum to 1 while preserving their relative proportions; conditional expectation is an ordinary expectation computed with these conditional probabilities.
  • Memorylessness of the geometric distribution: given that the first few coin tosses were tails, the number of further tosses needed to get a head has the same distribution as starting fresh, because coin tosses are independent.
  • Total expectation theorem: E[X] = sum over scenarios of P(scenario) * E[X | scenario], letting you compute an expectation by dividing into cases.
  • Joint, marginal, and conditional PMFs: the joint PMF p_{X,Y}(x,y) gives the probability of a pair of outcomes; summing over one variable gives the marginal PMF of the other; conditioning on a fixed value of one variable and rescaling gives the conditional PMF of the other.

Walkthrough

Review of PMFs and expectation (1:00)

The lecture recaps random variables, PMFs, and expected value from the previous lecture, along with the expected value rule for computing E[g(X)] directly from the PMF of X, and linearity of expectation for linear functions.

Variance and standard deviation (6:05)

Noting that the average signed distance from the mean is always zero, the lecture motivates variance as the average squared distance from the mean, derives it as an application of the expected value rule, and introduces standard deviation as its square root to restore the original units, illustrated with a two-speed travel example (1 mph vs. 200 mph, each with probability 1/2).

Why you can't reason on the average with nonlinear functions (14:12)

Continuing the travel example, the lecture computes E[T] (expected time) and shows that E[T] * E[V] does not equal E[TV] (which is exactly the fixed distance, 200), demonstrating that expectation does not commute with nonlinear functions like products or reciprocals.

Conditional PMFs and conditional expectation (17:13)

Conditional PMFs are introduced as ordinary PMFs applied to a new, conditioned model: the relative probabilities of remaining outcomes stay proportional but are rescaled to sum to 1. Conditional expectation is then just an ordinary expectation computed with these conditional probabilities, and every formula for expectations has a direct conditional counterpart.

Memorylessness of the geometric random variable (24:24)

Using a thought experiment comparing two coin-flippers, one who already flipped two tails and one starting fresh, the lecture argues intuitively and then formally (via PMF shifting) that the number of remaining flips until the first head has the same geometric distribution regardless of past failures — the memorylessness property.

Total expectation theorem and the expected value of a geometric random variable (34:44)

The total expectation theorem is stated as the PMF analogue of the law of total probability. Applying it by conditioning on whether the first toss is heads or tails, and using memorylessness to relate the conditional expectation for the "tails" case back to E[X] itself, the lecture derives E[X] = 1/p for a geometric random variable without evaluating an infinite sum directly.

Joint, marginal, and conditional PMFs (41:57)

To capture the relationship between two random variables from the same experiment (such as height and weight), the lecture introduces the joint PMF p_{X,Y}(x,y), represented as a table, and shows how to recover a marginal PMF by summing over the other variable, and a conditional PMF by fixing one variable's value and rescaling the corresponding slice of the table.

Before you watch

  • Review PMFs, expected value, and the expected value rule from the previous lecture in this course, since this lecture builds directly on them.
  • Be comfortable with conditional probability and the law of total probability from earlier in the course.

Check your understanding

  1. Why does variance use squared distance from the mean rather than signed distance?
  2. In the travel example, why does E[TV] differ from E[T] * E[V]?
  3. How does the memorylessness property simplify the derivation of the geometric random variable's expected value?
  4. What does the total expectation theorem let you do that computing an expectation directly might not?
  5. How do you recover a marginal PMF from a joint PMF table, and how does that differ from computing a conditional PMF?

Vocabulary

standard deviation (noun)
The square root of variance, giving a spread measure in the same units as the original variable.
Standard deviation restores the original units after taking variance.
commute (verb)
To give the same result regardless of the order operations are done in.
Expectation and nonlinear functions do not commute in general.
conditional expectation (noun)
The expected value of a random variable computed using conditional probabilities.
Conditional expectation uses the rescaled probabilities after conditioning.
memorylessness (noun)
A property where past outcomes don't affect the probability of future ones.
Memorylessness means the geometric distribution forgets past failed tosses.
total expectation theorem (noun)
A rule for finding an expectation by averaging over conditional expectations in different scenarios.
The total expectation theorem breaks the calculation into simple cases.
joint PMF (noun)
A function giving the probability of a specific pair of values for two random variables together.
The joint PMF shows the probability of each height-and-weight pair.
marginal PMF (noun)
The probability distribution of one random variable alone, found by summing over the other.
Summing the joint PMF's rows gives the marginal PMF of height.
shortcut (averaging) (noun)
A quick reasoning trick that avoids full calculation, but only works for linear cases.
There is no shortcut for averaging over a nonlinear function.
signed distance (noun)
A distance that keeps its positive or negative direction, rather than only its size.
The average signed distance from the mean is always zero.
units (physical) (noun)
The measurement scale a quantity is expressed in, such as meters or volts.
Standard deviation restores the original units of the variable.
rescale (verb)
To adjust values so they sum or fit correctly after a change.
Conditioning rescales probabilities to sum to 1 again.
counterpart (noun)
A matching version of a concept applied in a different setting.
Every expectation formula has a conditional counterpart.
shift (a PMF) (verb)
To move a distribution's values by a fixed amount.
Memorylessness is shown formally via PMF shifting.
table (data) (noun)
A grid used to organize values by row and column.
The joint PMF is represented as a table.
recover (marginal) (verb)
To calculate one distribution from a larger joint one.
Summing over y recovers the marginal PMF of x.
speed and time example (noun)
A classic illustration showing why averages don't multiply the way you might expect.
The speed and time example shows E of TV differs from E of T times E of V.
fixed distance (noun)
A distance value that stays the same no matter how speed and time vary.
The trip's fixed distance stays 200 no matter the speed chosen.
restate (verb)
To say something again, often in a clearer or more formal way.
The total expectation theorem restates the law of total probability for PMFs.
direct counterpart (noun)
A matching version of a concept in a related but different setting.
Every expectation formula has a direct conditional counterpart.
continue (a topic) (verb)
To keep developing an idea introduced earlier.
This lecture continues the study of discrete random variables.

Chapters

From the YouTube description

MIT 6.041 Probabilistic Systems Analysis and Applied Probability, Fall 2010
View the complete course: http://ocw.mit.edu/6-041F10
Instructor: John Tsitsiklis

License: Creative Commons BY-NC-SA
More information at http://ocw.mit.edu/terms
More courses at http://ocw.mit.edu

← PMF of a Function of a Random Variable · Flipping a Coin a Random Number of Times →