Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Probability · Lecture 69 of 76 · 48:49
Lecture 21: Bayesian Statistical Inference I
Study guide
What this lecture covers
This lecture opens the course's final unit on statistical inference: how you go from real-world data to a model of what's happening, and how to use that model to make estimates or decisions. It surveys where inference shows up (polling, medical trials, recommendation systems, signal processing, curve fitting) and frames the field conceptually as an application of probability tools already covered in the course, while noting that, unlike ordinary probability problems, inference problems don't have a single "correct" method.
The lecture develops the Bayesian approach specifically: treating an unknown quantity as a random variable with a prior distribution, then using Bayes' rule to compute its posterior distribution given data. It introduces two ways to report a single point estimate from a posterior — the maximum a posteriori (MAP) estimate and the conditional expectation — and proves that the conditional expectation minimizes mean squared error. After watching, you should be able to distinguish hypothesis testing from estimation problems, distinguish Bayesian from classical (non-Bayesian) approaches to an unknown quantity, compute MAP and conditional-expectation estimates from a posterior distribution, and explain why conditional expectation is optimal under squared error.
Key ideas
- Inference vs. probability: probability problems have a single correct answer given the model; inference problems involve choosing among multiple reasonable methods to fit a model to observed data, with no universally "correct" choice.
- Hypothesis testing vs. estimation: hypothesis testing problems involve an unknown quantity from a small discrete set of options, and aim to maximize the probability of a correct decision; estimation problems involve a continuous unknown quantity, and aim to minimize the size of the error.
- Bayesian vs. classical inference: the classical approach treats the unknown quantity
thetaas a fixed but unknown constant; the Bayesian approach treatsthetaas a random variable with a subjective prior distribution reflecting initial beliefs, updated via Bayes' rule after observing data. - Posterior distribution: applying Bayes' rule with a prior and a model of the data-generating process (
P(X | theta)) yields the posterior distribution or density ofthetagiven the observed data. - MAP estimate: reporting the value of
thetawith the highest posterior probability (discrete case) or highest posterior density (continuous case) minimizes the probability of an incorrect decision in the discrete case; in the continuous case its optimality rationale is weaker since any single continuous value has probability zero. - Conditional expectation estimate: reporting
E[theta | X]is a different, commonly used point estimate, especially natural for continuous unknowns. - Conditional expectation minimizes mean squared error: for the "no data" warm-up case, minimizing
E[(theta - c)^2]overcgivesc = E[theta], with resulting error equal toVar(theta); conditioning on data, the same argument showsE[theta | X]minimizes mean squared error in the conditional universe, and this optimality holds on average over all data too. - Practical difficulties: real problems require choosing a plausible prior (a judgment call) and often involve multi-dimensional integrals over vectors of unknowns and data, which can be computationally difficult even though the underlying formula is simple.
Walkthrough
From real phenomena to models (0:00)
The lecture frames statistical inference as the bridge between real-world data and probabilistic models, giving examples across polling, medical trials, the Netflix recommendation competition, financial data, and signal processing, and noting the field's long history in astronomical curve-fitting.
Why inference differs from probability, and the risk of misuse (5:11)
Unlike probability problems with a unique correct answer, inference problems admit multiple reasonable methods. The lecture warns that statistics is frequently misapplied when people plug data into a method without understanding its assumptions and guarantees.
System identification and communication as the same problem (9:16)
Using a signal-plus-noise example (X = A*S + noise), the lecture shows that estimating an unknown channel gain A (system identification) and decoding an unknown transmitted signal S (communication) are mathematically the same linear inference problem, just with different unknowns.
Classifying inference problems (12:19)
The lecture distinguishes hypothesis testing (discrete unknown, minimize probability of error) from estimation (continuous unknown, minimize size of error), using the airplane-radar and polling examples respectively.
Bayesian vs. classical philosophies (14:20)
Two philosophical approaches to an unknown theta are contrasted: classical statistics treats it as a fixed unknown number; Bayesian statistics models it as a random variable with a prior distribution reflecting subjective belief, even when the quantity isn't inherently random.
Bayes' rule and the posterior (18:24)
The lecture shows how Bayes' rule combines a prior P(theta) with a model of the data given theta to compute the posterior P(theta | X), for discrete and continuous cases, and applies this to estimating the coefficients of a parabolic trajectory from noisy position measurements, noting that unknowns and data are typically vectors rather than single variables.
MAP estimates and reporting a single answer (25:31)
Using a coin-bias example, the lecture contrasts the classical estimator (sample proportion of heads) with the Bayesian approach (choosing a prior, then computing the posterior). It introduces the MAP estimate as the posterior mode, which minimizes the probability of an incorrect decision in the discrete case, and notes the weaker justification for this choice with continuous densities.
Conditional expectation and the least-squares warm-up (34:43)
As a warm-up without data, the lecture minimizes E[(theta - c)^2] over constants c, showing the optimal choice is c = E[theta], with mean squared error equal to Var(theta).
Proving conditional expectation is optimal given data (38:50)
Applying the same argument within the "conditional universe" created by observing X, the lecture shows the optimal point estimate given data is E[theta | X]. It then proves this beats any other estimator g(X) in mean squared error for every value of X, and therefore on average, establishing conditional expectation as the optimal estimator under squared error.
Practical complications (45:58)
The lecture closes by noting that while the theory is simple, applying it involves choosing a defensible prior and computing potentially difficult multi-dimensional integrals, motivating a simpler alternative to be introduced in the next lecture.
Before you watch
- Review Bayes' rule for both discrete and continuous random variables, and conditional expectation.
- Recall the weak law of large numbers and the polling example, referenced here as a classical estimator.
Check your understanding
- How does the Bayesian approach to an unknown quantity differ philosophically from the classical approach, and how does this affect the tools each uses?
- What is the maximum a posteriori (MAP) estimate, and why is its optimality rationale different in discrete versus continuous settings?
- How is it proven that
E[theta | X]minimizes mean squared error, and why does this proof extend from a fixed value ofXto an average over all values ofX? - Why can implementing Bayesian estimation in practice be difficult even though the guiding formula is simple?
Vocabulary
- statistical inference (noun)
- The process of using observed data to learn about an unknown quantity or model.
Statistical inference bridges real-world data and probability models. - bridge (noun)
- Something that connects two different things or ideas.
Inference acts as a bridge between raw data and a working model. - polling (noun)
- Asking a sample of people questions to estimate opinions of a whole population.
Polling is one common application of statistical inference. - signal processing (noun)
- The field of analyzing and interpreting signals, such as sound or radio waves.
Signal processing uses inference to recover a message from noisy data. - curve fitting (noun)
- Finding a mathematical curve that best matches a set of observed data points.
Astronomers used curve fitting long before modern statistics existed. - misapply (verb)
- To use a method in the wrong way or in a situation it doesn't fit.
Statistics is often misapplied when people ignore a method's assumptions. - assumption (noun)
- Something taken to be true without direct proof, used as a starting point.
Every statistical method relies on certain assumptions about the data. - system identification (noun)
- Estimating the unknown properties of a system from its observed input and output.
System identification estimates the unknown channel gain A. - channel gain (noun)
- A number describing how much a signal is strengthened or weakened as it passes through a channel.
We want to estimate the channel gain A from noisy measurements. - decode (verb)
- To figure out the original message from a signal that has been changed or hidden.
The receiver must decode the transmitted signal S. - hypothesis testing (noun)
- A method for choosing between a small number of possible explanations based on data.
Hypothesis testing decides whether the radar detected a plane or just noise. - estimation (noun)
- The process of calculating an approximate value for an unknown continuous quantity.
Estimation problems involve unknowns that can take any value on a scale. - Bayesian (adjective)
- Relating to an approach that treats an unknown quantity as random and updates beliefs using data.
The Bayesian approach models theta with a prior distribution. - classical (statistics) (adjective)
- Relating to an approach that treats an unknown quantity as a fixed, non-random number.
The classical approach treats theta as a fixed unknown constant. - prior distribution (noun)
- The initial belief about an unknown quantity before seeing any data.
We choose a prior distribution to reflect our beliefs about theta. - subjective (adjective)
- Based on personal judgment rather than a single objective fact.
The prior reflects a subjective belief about the coin's bias. - posterior distribution (noun)
- The updated belief about an unknown quantity after using Bayes' rule with observed data.
Bayes' rule turns the prior into the posterior distribution. - Bayes' rule (noun)
- A formula for updating the probability of something after new evidence is observed.
We apply Bayes' rule to combine the prior with the data. - MAP (maximum a posteriori) estimate (noun)
- The single value of an unknown quantity with the highest posterior probability or density.
The MAP estimate is the mode of the posterior distribution. - mode (noun)
- The value that occurs with the highest probability or density.
The MAP estimate picks the mode of the posterior. - conditional expectation (noun)
- The expected value of a random variable, calculated using extra known information.
The conditional expectation E[theta | X] is a common point estimate. - mean squared error (noun)
- The average of the squared differences between estimates and true values.
Conditional expectation minimizes the mean squared error. - optimal (adjective)
- Being the best possible choice according to some measure.
The conditional expectation is the optimal estimator under squared error. - coin bias (noun)
- The unknown probability that a coin lands heads.
We use a coin-bias example to compare classical and Bayesian estimates. - point estimate (noun)
- A single number chosen to represent an unknown quantity, instead of a whole distribution.
MAP and conditional expectation are two kinds of point estimate. - trajectory (noun)
- The path an object follows over time, often described by a mathematical curve.
We estimate the coefficients of a parabolic trajectory from noisy data. - multi-dimensional (adjective)
- Involving more than one variable or direction at once.
Real inference problems often need multi-dimensional integrals. - integral (noun)
- A mathematical operation that adds up infinitely many small pieces, often to find area or total probability.
Computing the posterior can require a difficult integral. - warm-up (problem) (noun)
- A simple version of a problem used to build understanding before the full version.
The no-data case is a warm-up before adding observations. - philosophical (adjective)
- Relating to basic beliefs or ways of thinking about a subject.
Bayesian and classical statistics reflect different philosophical views.
Chapters
- 0:00 <Untitled Chapter 1>
- 2:30 Netflix Competition
- 5:22 Relation between the Field of Inference and the Field of Probability
- 8:57 Generalities
- 12:29 Classification of Inference Problems
- 14:43 Model the Quantity That Is Unknown
- 19:14 Bayes Rule
- 25:23 Example of an Estimation Problem with Discrete Data
- 31:02 Maximum a Posteriori Probability Estimate
- 35:46 Point Estimate
- 40:28 Conclusion
- 47:13 Issue Is that this Is a Formula That's Extremely Nice and Compact and Simple that You Can Write with Minimal Ink but behind It There Could Be Hidden a Huge Amount of Calculation So Doing any Sort of Calculations That Involve Multiple Random Variables Really Involves Calculating Multi-Dimensional Integrals and Multi-Dimensional Integrals Are Hard To Compute So Implementing Actually this Calculating Machine Here May Not Be Easy Might Be Complicated Computationally It's Also Complicated in Terms of Not Being Able To Derive Intuition about It So Perhaps You Might Want To Have a Simpler Version a Simpler Alternative to this Formula That's Easier To Work with and Easier To Calculate
From the YouTube description
MIT 6.041 Probabilistic Systems Analysis and Applied Probability, Fall 2010
View the complete course: http://ocw.mit.edu/6-041F10
Instructor: John Tsitsiklis
License: Creative Commons BY-NC-SA
More information at http://ocw.mit.edu/terms
More courses at http://ocw.mit.edu
← Using the Central Limit Theorem · Lecture 22: Bayesian Statistical Inference II →
