Seyed Masoud Hosseini · Overview · Study log · Ideas · Transcript · RSS feed
Probability · Lecture 1 of 76 · 51:10
Lecture 1: Probability Models and Axioms
Study guide
What this lecture covers
This opening lecture of MIT's probability course answers a basic question: what does it mean to build a rigorous model of an uncertain situation? Before any formulas appear, the lecture lays the groundwork that the rest of the course depends on: how to describe every possible result of a random experiment, and how to attach numbers to those results in a way that behaves consistently.
By the end you can describe the sample space of a simple experiment, explain why outcomes must be mutually exclusive and collectively exhaustive, and use the three axioms of probability (non-negativity, normalization, and additivity) to compute probabilities of simple events, including under the discrete uniform law and for a continuous sample space where probability is measured as area.
Key ideas
- Sample space: the set of all possible outcomes of an experiment; outcomes must be mutually exclusive (only one happens) and collectively exhaustive (one of them always happens).
- Modeling choice: how much detail to include in a sample space is a judgment call, not a fixed procedure — irrelevant details are usually left out.
- Event: a subset of the sample space; an event "occurs" when the actual outcome falls inside it.
- Axioms of probability: probabilities are non-negative, the probability of the whole sample space is 1, and the probability of a union of disjoint events is the sum of their probabilities.
- Discrete uniform law: if all outcomes in a finite sample space are equally likely, the probability of an event is the count of its outcomes divided by the total count.
- Continuous probability as area: on a sample space like a square, a natural probability law assigns each region a probability equal to its area, so single points have probability zero.
- Countable additivity: the additivity axiom needs to be extended to infinite sequences of disjoint events, which is not automatic and can lead to contradictions if misapplied to uncountable collections.
Walkthrough
Setting up a sample space (11:11)
The lecture opens the technical material by defining an experiment as anything with an uncertain result, and the sample space as the set that lists every possible outcome. Two requirements matter: the list must be mutually exclusive and collectively exhaustive. A simple coin flip has a two-element sample space, heads or tails. A more elaborate example — flipping a coin while also tracking whether it is raining — shows that the amount of detail included in a sample space is a choice, not something forced by the problem; unrelated details can usually be dropped. The lecture then works through rolling a four-sided die twice, showing both a list-based sample space of ordered pairs and an equivalent tree diagram, where each path from root to leaf corresponds to one outcome. It also introduces a sample space that is infinite: a dart that can land anywhere inside a unit square, where every point is a possible outcome.
Assigning probabilities and the axioms (21:06)
Rather than assigning probabilities to individual outcomes (which would often force them to be zero, as with a dart hitting one exact point), the lecture assigns probabilities to events — subsets of the sample space. Three ground rules, the axioms, are introduced: probabilities are non-negative, the probability of the entire sample space equals 1, and if two events are disjoint (no common outcomes), the probability that either occurs is the sum of their individual probabilities. This additivity property is compared to how mass or area adds up over disjoint regions.
What the axioms give us (34:29)
The lecture shows that facts not explicitly listed as axioms can still be derived from them. For instance, that probabilities never exceed 1 follows from combining all three axioms with the idea that an event and its complement are disjoint and together make up the whole sample space. The additivity rule for two disjoint events is then extended, step by step, to the union of three, and then any finite number, of disjoint events — the total probability is just the sum of the individual probabilities. A related consequence: for a finite set of outcomes, the probability of the set equals the sum of the probabilities of its individual elements.
Discrete uniform law and dice examples (42:16)
Returning to the two-dice example, the lecture assigns equal probability, 1/16, to each of the 16 outcomes. This is the discrete uniform law: when a finite sample space has N equally likely outcomes and an event contains n of them, its probability is n/N. Several example questions are answered by identifying which outcomes belong to the event of interest and counting them, including the probability that the first roll is 1, that the sum of the two rolls is odd, and that the minimum of the two rolls equals 2. The same counting-then-dividing procedure is then applied to the continuous dart example, where probability equals area: the probability of landing exactly on one point is zero, and the probability that the coordinates sum to at most 1/2 is found by computing the area of a triangle.
Continuous models and countable additivity (46:46)
The lecture closes by testing the limits of the additivity axiom using an experiment of flipping a coin until the first head appears, where the outcome is any positive integer and the probability of waiting n flips is 2^-n. Finding the probability of an even outcome requires summing infinitely many individual probabilities, which the basic (finite) additivity axiom does not justify. This motivates the countable additivity axiom: probabilities of disjoint events can be summed over an infinite sequence of events, but not over arbitrary infinite collections such as all the points in a square (a distinction developed further in the next lecture).
Before you watch
- No earlier lectures in this course are required; this is the first one.
- Basic set notation (union, intersection, subset, complement) will make the sample-space and event language easier to follow.
Check your understanding
- Why must a sample space be both mutually exclusive and collectively exhaustive?
- Using only the three axioms, how would you show that the probability of an event is always at most 1?
- In the two-dice example, why does each of the 16 outcomes get probability
1/16, and how would that change if the die were unfair? - Why does a single point in the dart-board example have probability zero, and what does that imply about probability of the whole square?
- What goes wrong if you try to apply plain additivity to an uncountable union of single-point events?
Chapters
- 0:00 Intro
- 1:03 Administrative Details
- 2:42 Mechanics
- 4:59 Sections
- 5:42 Style
- 6:50 Why Probability
- 8:52 Class Details
- 9:51 Goals
- 11:11 Sample Space
- 15:19 Example
- 21:06 Assigning probabilities
- 25:47 Intersection and Union
- 28:59 Are these axioms enough
- 31:09 Union of 3 sets
- 34:29 Union of finite sets
- 36:04 Weird sets
- 42:16 Discrete uniform law
- 46:46 An example
From the YouTube description
MIT 6.041 Probabilistic Systems Analysis and Applied Probability, Fall 2010
View the complete course: http://ocw.mit.edu/6-041F10
Instructor: John Tsitsiklis
License: Creative Commons BY-NC-SA
More information at http://ocw.mit.edu/terms
More courses at http://ocw.mit.edu
