Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Probability · Lecture 68 of 76 · 11:24

Using the Central Limit Theorem

Using the Central Limit Theorem on YouTube

Study guide

What this lecture covers

This problem applies the central limit theorem to a factory production scenario: daily gadget output is normal, i.i.d., and the questions involve totals over many days. It works through three parts: computing a tail probability for a fixed number of days, finding the largest number of days satisfying a probability constraint, and computing a probability involving a random first-crossing day.

It builds on the central limit theorem and continuity correction from the main lecture. After watching, you should be able to standardize a sum of i.i.d. normal variables, apply continuity correction, solve backward for a sample size given a target probability, and rewrite an event about a random index as an equivalent event about a fixed sum.

Key ideas

  • Setup: daily gadget output X_i is normal with mean 5 and variance 9, i.i.d. across days; sums of these are analyzed using the CLT.
  • Standardizing a sum: for n days, the sum has mean 5n and variance 9n; subtracting the mean and dividing by the standard deviation converts any probability question into one about a standard normal Z.
  • Continuity correction: since the sum is a discrete-like quantity, using a half-unit adjustment (e.g., 439.5 instead of 440) gives a more accurate normal approximation.
  • Solving backward for n: given a target probability bound (such as wanting a tail probability under 0.05), you can look up the corresponding normal quantile and solve the resulting inequality for the largest valid n.
  • Random index trick: the event "the first day the total exceeds 1000 is day 220 or later" is equivalent to "the sum over the first 219 days is still at most 1000," turning a question about a random stopping day into an ordinary CLT calculation on a fixed sum.

Before you watch

  • Review the central limit theorem and continuity correction from the main lecture on the CLT.
  • Be comfortable standardizing sums of independent normal random variables and reading normal distribution tables.

Check your understanding

  1. Why does using 439.5 instead of 440 (or 439) improve the normal approximation to a discrete sum?
  2. How would you solve for the largest n satisfying a probability bound like P(sum >= 200 + 5n) <= 0.05?
  3. Why is the event "the first day total exceeds 1000 occurs on day 220 or later" equivalent to "the sum through day 219 is at most 1000"?
  4. How does the mean and variance of the sum change as the number of days n increases, and how does that affect the standardized variable?

Vocabulary

central limit theorem (noun)
A rule saying that the sum of many independent random values is close to a normal distribution.
We use the central limit theorem to approximate the total gadget output.
i.i.d. (independent and identically distributed) (phrase)
Describes random variables that are independent of each other and all follow the same distribution.
Daily output is i.i.d. across days.
normal distribution (noun)
A common bell-shaped probability distribution described by a mean and a variance.
Each day's output is normal with mean 5 and variance 9.
mean (noun)
The average value of a random variable.
The mean of the daily output is 5 units.
variance (noun)
A number that measures how spread out a random variable's values are.
The variance of daily output is 9.
standard deviation (noun)
The square root of the variance; it measures typical spread in the same units as the data.
We divide by the standard deviation to standardize the sum.
standardize (verb)
To convert a random variable into a standard form by subtracting its mean and dividing by its standard deviation.
We standardize the sum to turn it into a standard normal variable Z.
standard normal (noun)
The normal distribution with mean 0 and variance 1, usually called Z.
After standardizing, we look up the value in the standard normal table.
tail probability (noun)
The probability that a random variable is far above or below its typical value.
We compute the tail probability that the sum exceeds a large number.
approximation (noun)
A result that is close to the exact answer but not perfectly exact.
The CLT gives a good approximation even though the sum isn't exactly normal.
continuity correction (noun)
A small adjustment used when approximating a whole-number quantity with a smooth normal curve.
We use 439.5 instead of 440 as a continuity correction.
half-unit adjustment (phrase)
Shifting a value by 0.5 to better match a discrete quantity with a continuous approximation.
The half-unit adjustment makes the normal approximation more accurate.
quantile (noun)
The value below which a given probability of the distribution falls.
We look up the quantile that corresponds to a 5% tail probability.
inequality (noun)
A mathematical statement that one quantity is greater than, less than, or not equal to another.
We solve the inequality to find the largest valid n.
solve backward (phrase)
To start from a known result and work out what input value would produce it.
We solve backward for n given a target probability of 0.05.
bound (noun)
A limit that a quantity must not go above or below.
We want the probability to stay under a bound of 0.05.
stopping day (noun)
The specific day when a running total first reaches or crosses a target value.
The stopping day is the first day the total output exceeds 1000.
random index (noun)
A variable, like a day number, whose exact value is itself uncertain until observed.
The random index trick rewrites a question about the stopping day using a fixed sum.
equivalent (adjective)
Having the same meaning or the same effect, even if written differently.
The two events turn out to be equivalent.
exceed (verb)
To go above a certain amount.
We want the day when total production first exceeds 1000 gadgets.
worked calculation (phrase)
A full step-by-step solution to a numerical problem.
The lecture shows three worked calculations using the CLT.
trick (noun)
A clever method that makes a hard problem easier to solve.
The random-index trick turns a tricky question into a standard CLT calculation.

Chapters

From the YouTube description

MIT 6.041SC Probabilistic Systems Analysis and Applied Probability, Fall 2013
View the complete course: http://ocw.mit.edu/6-041SCF13
Instructor: Jagdish Ramakrishnan

License: Creative Commons BY-NC-SA
More information at http://ocw.mit.edu/terms
More courses at http://ocw.mit.edu

← Probability Bounds · Lecture 21: Bayesian Statistical Inference I →