Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
Distributed Systems · Lecture 1 of 20 · 1:19:35
Lecture 1: Introduction
Study guide
What this lecture covers
This opening lecture asks why anyone builds distributed systems at all, given that a single computer is always simpler, and lays out the reasons: performance through parallelism, fault tolerance, physically separated problems, and security isolation. It then works through the core challenges that make distributed systems hard, namely concurrency, partial failure, and the difficulty of turning more computers into proportionally more performance.
As the first lecture in MIT's Distributed Systems course, it also sets expectations for how the class runs: reading a paper before each lecture, four programming labs building toward a sharded, fault-tolerant key-value store, and two exams. It closes with a detailed walkthrough of MapReduce, the case study for the week, showing how a simple map/reduce programming model can be executed reliably across thousands of machines. After watching, you should be able to explain what makes distributed systems hard, describe the course's lab sequence, and trace how a MapReduce job moves data from input files through map and reduce tasks to output.
Key ideas
- Distributed system: a set of cooperating computers communicating over a network to complete one coherent task.
- Scalability: adding more computers should yield a proportional increase in performance; in practice a single component (often a database) becomes a bottleneck and limits this.
- Partial failure: unlike a single computer, which mostly either works or is down, a system of many computers can have some parts working and others failing at the same time.
- Availability vs. recoverability: an available system keeps serving requests despite some failures; a recoverable system stops on failure but resumes correctly once repaired.
- Consistency: the rules that define what a read returns relative to earlier writes, especially when data is replicated; strong consistency guarantees the latest write is seen, weak consistency does not.
- Non-volatile storage and replication: the two main tools for fault tolerance, used to survive crashes and to keep serving requests when a copy fails.
- MapReduce: a framework where programmers write simple map and reduce functions and the framework handles distributing the work, moving data (the "shuffle"), and tolerating failures.
Walkthrough
Why build distributed systems, and why they're hard (0:01)
The lecture opens by defining a distributed system and arguing you should avoid building one unless a single computer genuinely cannot do the job. The four reasons people build them anyway are performance (parallelism across many CPUs and disks), fault tolerance (redundant computers), problems that are inherently spread across locations, and security isolation between mutually distrusting components. The lecture then explains why these systems are hard to build: concurrency introduces timing-dependent bugs, partial failures create unpredictable states that a single machine never has, and achieving the performance you expect from adding more machines takes careful design.
Course structure and labs (7:13)
The instructor describes the course mechanics: lectures built around case-study papers (one due before each class, with a short question and answer submitted beforehand), a midterm and final exam, and four programming labs. Lab 1 implements a simplified MapReduce, Lab 2 uses the Raft protocol for fault tolerance, Lab 3 builds a replicated key-value server on top of Raft, and Lab 4 shards that key-value store across multiple replicated groups for parallel performance. Students may substitute an original team project for Lab 4.
Scalability and fault tolerance (17:23)
Using a web server plus database example, the lecture shows how adding more web servers can scale a system smoothly until the shared database becomes the bottleneck, illustrating why scalability rarely extends indefinitely without redesign. It then turns to fault tolerance: at large scale, failures that would be rare on one machine become routine (roughly three failures a day across a thousand machines), so systems must be designed assuming something is always broken. The two central defenses are non-volatile storage, for recovering state after a crash, and replication, for continuing to serve requests when one copy fails.
Consistency and the cost of correctness (36:37)
Using a key-value store with put and get operations, the lecture shows how replication can lead a client to read stale data if a write reaches only one replica before a crash. It contrasts strong consistency, where a get is guaranteed to return the most recent put, with weaker consistency models that allow stale reads in exchange for avoiding expensive cross-replica communication. The lecture notes that placing replicas far apart for fault tolerance (different data centers or continents) makes strong consistency more expensive, since confirming a write may require tens of milliseconds of network round trips.
MapReduce as a case study (46:43)
The lecture introduces MapReduce as the framework Google built so that engineers who were not distributed-systems specialists could still run huge computations, such as indexing the web, across thousands of machines. A programmer writes only a map function, which turns an input file into key-value pairs, and a reduce function, which processes all values sharing one key; the framework handles splitting input, scheduling tasks across worker machines, and combining results. The word-count example shows how simple these functions can be: map emits (word, 1) for every word, and reduce sums the values for each word.
How a MapReduce job actually runs (1:02:55)
A master server assigns map and reduce tasks to worker machines. Map output is written to local disk rather than sent over the network immediately, and the system tries to run each map task on the same machine that stores its input in GFS (the Google File System), minimizing network traffic. The "shuffle" step, where each key's values are gathered from every worker onto the machine running its reduce task, is where most network communication happens, and in the 2004 paper this was the main performance bottleneck, since each machine's share of the network was only about 50 megabits per second. The lecture notes that modern data center networks, with multiple root switches, no longer force this constraint, and that Google eventually moved on from MapReduce.
Before you watch
- No prior lecture is required; this is the first lecture of the course.
- Basic familiarity with programming and with the idea of a network connecting computers is assumed.
- It helps to have skimmed the MapReduce paper, since the lecture discusses it in detail without re-explaining every detail from scratch.
Check your understanding
- Why does the lecture argue you should avoid building a distributed system if a single computer can do the job?
- What is the difference between an available system and a recoverable system?
- In the key-value store example, how can a client end up reading a stale value after a replica crash?
- Why did the "shuffle" step dominate MapReduce's runtime in the 2004 Google paper, and what trick did Google use to reduce network traffic during the map phase?
- What does it mean for a system to achieve scalable speedup, and why does it typically stop scaling at some point?
Vocabulary
- distributed system (noun)
- A group of computers working together over a network to do one job.
A distributed system can keep working even if one machine fails. - parallelism (noun)
- Doing many parts of a task at the same time.
Parallelism across many CPUs speeds up big computations. - fault tolerance (noun)
- The ability of a system to keep working despite failures.
Redundant computers give a system fault tolerance. - isolation (noun)
- Keeping something separate so it can't affect or be affected by others.
Security isolation keeps mutually distrusting programs apart. - concurrency (noun)
- Multiple tasks happening or being handled at overlapping times.
Concurrency introduces timing-dependent bugs in a program. - partial failure (noun)
- A situation where some parts of a system fail while others keep working.
Partial failure never happens on a single, simple computer. - bottleneck (noun)
- The part of a system that limits its overall speed or performance.
A shared database can become a bottleneck as traffic grows. - scalability (noun)
- A system's ability to handle more work as more resources are added.
Adding servers doesn't always give proportional scalability. - replication (noun)
- Keeping multiple identical copies of data on different machines.
Replication lets a service keep running when one copy fails. - consistency (noun)
- The rules that define what value a read should return after writes.
Strong consistency guarantees the latest write is seen. - stale (adjective)
- Out of date, not reflecting the most recent change.
A crashed replica can cause a client to read stale data. - non-volatile storage (noun)
- Storage that keeps its data even after power is lost.
Non-volatile storage lets a machine recover state after a crash. - sharded (adjective)
- Split into separate pieces stored across different servers.
The final lab builds a sharded, fault-tolerant key-value store. - master server (noun)
- A central machine that coordinates and assigns work to others.
A master server assigns map and reduce tasks to workers. - shuffle (noun)
- The step of moving data between machines so matching keys end up together.
The shuffle step caused most of MapReduce's network traffic. - worker (noun)
- A machine that carries out tasks assigned by a coordinator.
Each worker runs map or reduce tasks on its share of data. - index (verb)
- To build a searchable record of a large collection of data.
MapReduce was built to index the entire web. - round trip (noun)
- The time for a message to travel to another machine and a reply to return.
Confirming a write can require a network round trip. - case study (noun)
- A detailed example used to teach a general lesson.
MapReduce is used as this week's case study. - coherent (adjective)
- Logically connected and making sense as a whole.
Many computers cooperate to complete one coherent task. - redundant (adjective)
- Extra and able to replace something else if it fails.
Redundant computers provide fault tolerance. - recoverable (adjective)
- Able to return to correct operation after stopping.
A recoverable system resumes correctly once repaired.
Chapters
- 0:00 Distributed Systems
- 7:58 Course Overview
- 12:12 Programming Labs
- 17:33 Infrastructure for Applications
- 20:57 Topics
- 23:17 Scalability
- 29:21 Failure
- 31:33 Availability
- 37:15 Consistency
- 46:50 Map Reduce
- 50:04 MapReduce
- 57:13 Reduce
From the YouTube description
Lecture 1: Introduction
MIT 6.824: Distributed Systems (Spring 2020)
https://pdos.csail.mit.edu/6.824/
