Seyed Masoud Hosseini · Overview · Study log · Ideas · Transcript · RSS feed
Parallel Computing & CUDA
Stanford CS149 · AI · Fall 2026 · Planned
Lectures
- Lecture 1: Why Parallelism? Why Efficiency? (1:12:21)
- Lecture 2: A Modern Multi-Core Processor (1:16:13)
- Lecture 3: Multi-core Arch Part II and ISPC Programming Abstractions (1:16:18)
- Lecture 4: Parallel Programming Basics (1:17:14)
- Lecture 5: Work Distribution and Scheduling (1:17:39)
- Lecture 6: Locality, Communication, and Contention (1:17:24)
- Lecture 7: GPU Architecture and CUDA Programming (1:18:47)
- Lecture 8: Data-Parallel Thinking (1:17:48)
- Lecture 9: Distributed Data-Parallel Computing Using Spark (1:17:54)
- Lecture 10: Efficiently Evaluating DNNs on GPUs (1:20:26)
- Lecture 11: Cache Coherence (1:20:37)
- Lecture 12: Memory Consistency (1:19:15)
- Lecture 13: Fine-Grained Synchronization and Lock-Free Programming (1:15:47)
- Lecture 14: Midterm Review (1:13:12)
- Lecture 15: Domain-Specific Programming Languages (1:18:53)
- Lecture 16: Transactional Memory 1 (1:20:20)
- Lecture 17: Transactional Memory 2 (1:18:33)
- Lecture 18: Hardware Specialization (1:11:48)
- Lecture 19: Accessing Memory + Course Wrap Up (1:08:18)
Notes
No notes yet.
References
No references yet.
Study log
No log entries for this course yet.