Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Probability · Lecture 68 of 76 · 11:24
Using the Central Limit Theorem
Study guide
What this lecture covers
This problem applies the central limit theorem to a factory production scenario: daily gadget output is normal, i.i.d., and the questions involve totals over many days. It works through three parts: computing a tail probability for a fixed number of days, finding the largest number of days satisfying a probability constraint, and computing a probability involving a random first-crossing day.
It builds on the central limit theorem and continuity correction from the main lecture. After watching, you should be able to standardize a sum of i.i.d. normal variables, apply continuity correction, solve backward for a sample size given a target probability, and rewrite an event about a random index as an equivalent event about a fixed sum.
Key ideas
- Setup: daily gadget output
X_iis normal with mean 5 and variance 9, i.i.d. across days; sums of these are analyzed using the CLT. - Standardizing a sum: for
ndays, the sum has mean5nand variance9n; subtracting the mean and dividing by the standard deviation converts any probability question into one about a standard normalZ. - Continuity correction: since the sum is a discrete-like quantity, using a half-unit adjustment (e.g., 439.5 instead of 440) gives a more accurate normal approximation.
- Solving backward for n: given a target probability bound (such as wanting a tail probability under 0.05), you can look up the corresponding normal quantile and solve the resulting inequality for the largest valid
n. - Random index trick: the event "the first day the total exceeds 1000 is day 220 or later" is equivalent to "the sum over the first 219 days is still at most 1000," turning a question about a random stopping day into an ordinary CLT calculation on a fixed sum.
Before you watch
- Review the central limit theorem and continuity correction from the main lecture on the CLT.
- Be comfortable standardizing sums of independent normal random variables and reading normal distribution tables.
Check your understanding
- Why does using 439.5 instead of 440 (or 439) improve the normal approximation to a discrete sum?
- How would you solve for the largest
nsatisfying a probability bound likeP(sum >= 200 + 5n) <= 0.05? - Why is the event "the first day total exceeds 1000 occurs on day 220 or later" equivalent to "the sum through day 219 is at most 1000"?
- How does the mean and variance of the sum change as the number of days
nincreases, and how does that affect the standardized variable?
Vocabulary
- central limit theorem (noun)
- A rule saying that the sum of many independent random values is close to a normal distribution.
We use the central limit theorem to approximate the total gadget output. - i.i.d. (independent and identically distributed) (phrase)
- Describes random variables that are independent of each other and all follow the same distribution.
Daily output is i.i.d. across days. - normal distribution (noun)
- A common bell-shaped probability distribution described by a mean and a variance.
Each day's output is normal with mean 5 and variance 9. - mean (noun)
- The average value of a random variable.
The mean of the daily output is 5 units. - variance (noun)
- A number that measures how spread out a random variable's values are.
The variance of daily output is 9. - standard deviation (noun)
- The square root of the variance; it measures typical spread in the same units as the data.
We divide by the standard deviation to standardize the sum. - standardize (verb)
- To convert a random variable into a standard form by subtracting its mean and dividing by its standard deviation.
We standardize the sum to turn it into a standard normal variable Z. - standard normal (noun)
- The normal distribution with mean 0 and variance 1, usually called Z.
After standardizing, we look up the value in the standard normal table. - tail probability (noun)
- The probability that a random variable is far above or below its typical value.
We compute the tail probability that the sum exceeds a large number. - approximation (noun)
- A result that is close to the exact answer but not perfectly exact.
The CLT gives a good approximation even though the sum isn't exactly normal. - continuity correction (noun)
- A small adjustment used when approximating a whole-number quantity with a smooth normal curve.
We use 439.5 instead of 440 as a continuity correction. - half-unit adjustment (phrase)
- Shifting a value by 0.5 to better match a discrete quantity with a continuous approximation.
The half-unit adjustment makes the normal approximation more accurate. - quantile (noun)
- The value below which a given probability of the distribution falls.
We look up the quantile that corresponds to a 5% tail probability. - inequality (noun)
- A mathematical statement that one quantity is greater than, less than, or not equal to another.
We solve the inequality to find the largest valid n. - solve backward (phrase)
- To start from a known result and work out what input value would produce it.
We solve backward for n given a target probability of 0.05. - bound (noun)
- A limit that a quantity must not go above or below.
We want the probability to stay under a bound of 0.05. - stopping day (noun)
- The specific day when a running total first reaches or crosses a target value.
The stopping day is the first day the total output exceeds 1000. - random index (noun)
- A variable, like a day number, whose exact value is itself uncertain until observed.
The random index trick rewrites a question about the stopping day using a fixed sum. - equivalent (adjective)
- Having the same meaning or the same effect, even if written differently.
The two events turn out to be equivalent. - exceed (verb)
- To go above a certain amount.
We want the day when total production first exceeds 1000 gadgets. - worked calculation (phrase)
- A full step-by-step solution to a numerical problem.
The lecture shows three worked calculations using the CLT. - trick (noun)
- A clever method that makes a hard problem easier to solve.
The random-index trick turns a tricky question into a standard CLT calculation.
Chapters
From the YouTube description
MIT 6.041SC Probabilistic Systems Analysis and Applied Probability, Fall 2013
View the complete course: http://ocw.mit.edu/6-041SCF13
Instructor: Jagdish Ramakrishnan
License: Creative Commons BY-NC-SA
More information at http://ocw.mit.edu/terms
More courses at http://ocw.mit.edu
← Probability Bounds · Lecture 21: Bayesian Statistical Inference I →
