Seyed Masoud Hosseini · Overview · Study log · Ideas · Transcript · RSS feed

Parallel Computing & CUDA

Stanford CS149 · AI · Fall 2026 · Planned

Lectures

  1. Lecture 1: Why Parallelism? Why Efficiency? (1:12:21)
  2. Lecture 2: A Modern Multi-Core Processor (1:16:13)
  3. Lecture 3: Multi-core Arch Part II and ISPC Programming Abstractions (1:16:18)
  4. Lecture 4: Parallel Programming Basics (1:17:14)
  5. Lecture 5: Work Distribution and Scheduling (1:17:39)
  6. Lecture 6: Locality, Communication, and Contention (1:17:24)
  7. Lecture 7: GPU Architecture and CUDA Programming (1:18:47)
  8. Lecture 8: Data-Parallel Thinking (1:17:48)
  9. Lecture 9: Distributed Data-Parallel Computing Using Spark (1:17:54)
  10. Lecture 10: Efficiently Evaluating DNNs on GPUs (1:20:26)
  11. Lecture 11: Cache Coherence (1:20:37)
  12. Lecture 12: Memory Consistency (1:19:15)
  13. Lecture 13: Fine-Grained Synchronization and Lock-Free Programming (1:15:47)
  14. Lecture 14: Midterm Review (1:13:12)
  15. Lecture 15: Domain-Specific Programming Languages (1:18:53)
  16. Lecture 16: Transactional Memory 1 (1:20:20)
  17. Lecture 17: Transactional Memory 2 (1:18:33)
  18. Lecture 18: Hardware Specialization (1:11:48)
  19. Lecture 19: Accessing Memory + Course Wrap Up (1:08:18)

Notes

No notes yet.

References

No references yet.

Study log

No log entries for this course yet.