Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Probability · Lecture 71 of 76 · 24:51
Inferring a Parameter of the Uniform Distribution, Part 1
Study guide
What this lecture covers
This is a worked review problem that reuses the Bayesian inference tools from Chapter 8 of the course. Romeo and Juliet, introduced earlier when Juliet's arrival delay was modeled as uniform between 0 and 1 hour, return with a twist: the upper bound of that uniform distribution, called theta, is now itself unknown and treated as a random variable, uniform between 0 and 1. The lecture derives the posterior distribution of theta after observing one delay, then extends the result to n independent observations.
By the end you should be able to set up a prior and a conditional (likelihood) distribution for a simple two-layer random model, apply Bayes' rule to get a posterior, and read off how each new data point narrows the range of plausible values for an unknown parameter. This sets up Part 2, which uses the posterior to build point estimators.
Key ideas
- Two layers of randomness: theta (the true upper bound of the delay distribution) is itself uniform between 0 and 1, and the observed delay
xis uniform between 0 and theta, conditional on theta. - Prior: the initial belief about theta before seeing data, here a flat uniform density on
[0,1]. - Likelihood: the conditional PDF of
xgiven theta,1/thetaforxbetween 0 and theta. - Bayes' rule for the posterior:
f(theta|x) = f(theta) * f(x|theta) / f(x), where the denominator is a normalizing constant that does not depend on theta. - Posterior shape: for one observation, the posterior is proportional to
1/thetaon the range[x, 1], so it rules out any theta smaller than the observed delay. - Conditional independence across samples: given theta, multiple observed delays are treated as independent, so their joint likelihood is the product of individual
1/thetaterms. - Sufficient statistic: with n observations, only the maximum observed delay (
x_bar) matters; the posterior becomes proportional to1/theta^non[x_bar, 1]. - More data sharpens the posterior: each additional observation both narrows the valid range for theta and makes the posterior more peaked near the sufficient statistic.
Walkthrough
Setting up the problem (1:00)
The lecture restates the classic Romeo-and-Juliet delay problem, now adding uncertainty about the uniform distribution's own upper bound theta. It defines theta as uniform on [0,1] and the delay x, conditional on theta, as uniform on [0, theta].
Writing down the prior and conditional distributions (2:01)
The PDF of theta is 1 on [0,1]. The conditional PDF of x given theta is 1/theta on [0, theta]. The lecture explains why the goal of inference is to combine this prior and the observed data into a posterior via Bayes' rule.
Deriving the posterior for one observation (4:01)
Applying Bayes' rule, the numerator is the product of the prior and the conditional PDF, valid only where theta is in [0,1] and x is in [0, theta]. The normalizing denominator is computed by integrating over theta from x to 1, giving |log x|. The resulting posterior is 1/(theta * |log x|) on [x, 1].
Interpreting the posterior (8:07)
The lecture explains what the shape means: observing a delay of, say, half an hour rules out any theta below half an hour, and values of theta closer to the observed x are more likely because a smaller theta concentrates more probability near that value.
Extending to multiple observations (12:16)
With n independent dates, each delay x_1 through x_n is conditionally independent given theta. Their joint conditional PDF is the product of individual 1/theta terms, giving 1/theta^n, valid only when theta is at least as large as the maximum observed delay, x_bar.
The posterior with n data points (16:34)
Applying Bayes' rule again, the posterior becomes proportional to 1/theta^n on [x_bar, 1]. The lecture notes the normalizing constant can be computed the same way as before but is less important than the shape, which becomes steeper and more concentrated as n grows.
Before you watch
- Be comfortable with the original Romeo-and-Juliet uniform-delay setup from earlier in the course.
- Review Bayes' rule for continuous random variables and how to compute a posterior PDF from a prior and a conditional PDF.
- Know how to integrate simple PDFs like
1/thetato find normalizing constants.
Check your understanding
- Why is the posterior distribution restricted to the range
[x, 1]rather than the full[0,1]prior range? - How does observing a delay of
xrule out certain values of theta? - Why does conditional independence let you write the joint likelihood of
nobservations as a product? - Why does only the maximum observed delay matter when there are multiple observations?
- What happens to the shape of the posterior as the number of observations
nincreases?
Vocabulary
- parameter (noun)
- A fixed number that defines the shape of a distribution.
The unknown parameter theta sets the upper bound of the uniform distribution. - upper bound (noun)
- The highest value that a quantity is allowed to take.
Theta is the upper bound of Juliet's possible arrival delay. - two layers of randomness (phrase)
- A situation where one random variable's distribution depends on another random variable.
There are two layers of randomness: theta, and then the delay given theta. - prior (noun)
- The starting belief about an unknown quantity before seeing data.
The prior on theta is uniform between 0 and 1. - likelihood (noun)
- The probability or density of observing the data, given a particular value of the unknown parameter.
The likelihood of x given theta is 1/theta. - normalizing constant (noun)
- A number used to divide a formula so that the total probability adds up to 1.
The denominator in Bayes' rule is a normalizing constant. - rule out (phrasal verb)
- To decide that something is not possible.
Observing a delay of x rules out any theta smaller than x. - conditional independence (noun)
- The property that two events are independent once you know a third variable.
Given theta, the observed delays show conditional independence. - joint likelihood (noun)
- The combined probability of all observed data occurring together, for a given parameter value.
The joint likelihood of n delays is the product of individual likelihoods. - sufficient statistic (noun)
- A single summary of the data that captures everything needed to estimate the parameter.
The maximum observed delay is a sufficient statistic here. - sharpen (verb)
- To make more precise or more concentrated.
More data sharpens the posterior distribution. - concentrated (adjective)
- Gathered closely together around a particular value, rather than spread out.
The posterior becomes more concentrated as n grows. - PDF (probability density function) (noun)
- A function describing the relative likelihood of a continuous random variable taking a given value.
The PDF of x given theta is 1/theta on the interval [0, theta]. - integrate (verb)
- To calculate a total or area using calculus, by adding up infinitely small pieces.
We integrate over theta from x to 1 to find the normalizing constant. - steeper (adjective)
- Rising or falling more sharply.
The posterior curve becomes steeper with more observations. - review problem (phrase)
- A practice problem that revisits and applies material already taught.
This is a review problem reusing tools from Chapter 8. - twist (noun)
- An unexpected change to a familiar situation.
The upper bound being unknown is the twist added to the old problem. - delay (noun)
- The amount of extra time before something happens.
Juliet's arrival delay is modeled as a random variable. - arrival (noun)
- The act of reaching or getting to a place.
The delay measures the gap before Juliet's arrival. - range (noun)
- The set of all values a variable is allowed to take.
The posterior is restricted to the range [x, 1]. - extend (a result) (verb)
- To apply a finding to a broader or more general case.
The lecture extends the one-observation result to n observations. - narrow (a range) (verb)
- To reduce the set of possible values.
Each new observation helps narrow the range for theta.
Chapters
- 0:00 Introduction
- 1:03 Problem Description
- 2:20 Distributions
- 4:29 The posterior
- 8:01 Interpreting the posterior
- 10:28 Scaling factor
- 18:15 More Data
From the YouTube description
MIT 6.041SC Probabilistic Systems Analysis and Applied Probability, Fall 2013
View the complete course: http://ocw.mit.edu/6-041SCF13
Instructor: Jimmy Li
License: Creative Commons BY-NC-SA
More information at http://ocw.mit.edu/terms
More courses at http://ocw.mit.edu
← Lecture 22: Bayesian Statistical Inference II · Inferring a Parameter of the Uniform Distribution, Part 2 →
