Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Digital Design & Computer Architecture · Lecture 30 of 37 · 1:47:46

Lecture 24: Virtual Memory

Digital Design and Comp. Arch. - L24: Virtual Memory (Spring 2025) on YouTube

Study guide

What this lecture covers

This lecture, guest-taught by a SAFARI research group PhD student, answers a basic question: why don't programs ever see real physical memory addresses, and how does the system fake it for them? It sits in the course's memory-system arc, right after the caching material, and treats physical memory itself as a kind of cache for data that ultimately lives on disk.

By the end you should be able to explain why virtual memory exists, walk through how a virtual address is translated to a physical one using a multi-level page table, describe what happens on a page fault, and explain how TLBs and memory protection (including the Row Hammer attack) fit into the same hardware-software system.

Key ideas

  • Virtual memory: an abstraction that gives each process the illusion of a large, private, contiguous address space, while the operating system and hardware cooperatively map it onto a much smaller physical memory.
  • Page and frame: the virtual address space is split into fixed-size pages; the physical address space is split into frames of the same size. A page table maps pages to frames.
  • Page table: a per-process data structure, conceptually a dictionary, that stores the virtual-to-physical mapping plus a validity bit and permission bits for each page.
  • Multi-level page table: instead of one huge flat table, the table is organized as a tree of smaller tables so that unused parts of the virtual address space cost nothing in physical memory.
  • Page fault: what happens when a page table entry is invalid; a minor fault just needs a free physical frame, a major fault also needs to fetch the page from disk.
  • TLB (translation lookaside buffer): a small hardware cache inside the memory management unit that stores recently used translations so most accesses skip the page table walk.
  • Memory protection: page tables also carry read/write/execute permissions, which is what keeps one process from touching another process's memory or overwriting shared library code.

Walkthrough

Why programs need virtual memory (3:49)

The lecture opens with the ideal a programmer wants: free, infinite, zero-cost memory. Reality is a small, finite physical memory shared by multiple programs. If programs addressed physical memory directly, the programmer would have to worry about whether code and data fit, where to place them so two processes don't collide, and how to relocate or share data safely. Virtual memory removes this burden by giving each process its own address space and letting the system, not the programmer, decide where data actually lives.

Address translation and the page table (11:58)

The lecture works through a worked example: a 2 GB virtual address space, a much smaller physical memory, and 4 KB pages. It shows how to split an address into a virtual page number and a page offset, and how the page offset stays untranslated because it only locates data within a page. A page table entry supplies a validity bit and the physical frame number; if the entry is invalid, the requested page is on disk rather than in memory. Several example addresses are translated by hand to show the lookup.

Multi-level page tables (29:17)

A flat page table for a 64-bit address space would need an infeasible amount of memory per process. The fix is a hierarchical, tree-shaped page table: a top-level table with a fixed size points to second-level tables only for the parts of the address space actually in use. In x86-64, four levels of nine bits each are walked from a base register (CR3) down to a leaf entry, so an unused region of the address space never needs its lower-level tables allocated at all.

Page faults and the operating system's role (39:26)

When a virtual page is touched for the first time, the OS must "fault it in": find free physical memory and update the page table. A minor fault (an anonymous allocation with no backing file) is cheap. A major fault, where the data is backed by a file on disk, is expensive: the OS signals an I/O controller, which uses direct memory access (DMA) to move the block from disk into memory without involving the processor, and then the page table entry is updated and execution resumes. The lecture also covers the clock algorithm, a cheap approximation of least-recently-used, for choosing which page to evict when memory is full.

Memory protection and the Row Hammer attack (1:07:37)

Beyond translation, page tables carry access permissions per page, enforced at the same time as translation, which is what isolates processes and lets code pages be marked read-only. This section connects that mechanism to Row Hammer: repeatedly activating a DRAM row can flip bits in physically neighboring rows. An attacker can spray physical memory with page tables so that a flipped bit is likely to land inside one, turning a hardware reliability issue into a way to gain unauthorized access to physical memory.

Speeding up translation: TLBs and the MMU (1:22:50)

Since every memory access would otherwise need an extra access just to read the page table, modern cores include a memory management unit (MMU) with a translation lookaside buffer (TLB) that caches recent translations, page-table caches for intermediate tree levels, and a hardware page table walker. The lecture contrasts hardware-managed TLB misses (fast, but with a fixed page table format) against software-managed ones, as used in MIPS, which need an expensive exception but let the OS define its own page table layout.

Virtual memory overheads and current research (1:37:59)

The lecture closes with the presenter's own research: measuring how much time real workloads spend on memory allocation and address translation. Short-running workloads like LLM inference can spend around 32% of execution time simply allocating memory, and data-intensive workloads with irregular access patterns can spend around 26% of execution time on address translation because they overwhelm the TLB. These overheads motivate ongoing work on rethinking virtual memory for GPUs, unified CPU-GPU memory, and disaggregated memory systems.

Before you watch

  • Review the caching lectures earlier in the course; the page table is explicitly presented as a cache-like structure for pages that may live on disk.
  • Be comfortable converting between bits, bytes, and powers of two, since the lecture derives page table sizes from address-space arithmetic.
  • Know what a process and an address space are from earlier systems material; the lecture assumes you already understand basic process isolation.

Check your understanding

  1. Why can't a program simply use physical addresses directly, even on a system with more than enough physical memory for all running programs?
  2. Walk through what happens, step by step, when a virtual address is not found valid in any level of a four-level page table.
  3. Why does a multi-level page table save space compared to a single flat page table, and what has to be true for that saving to be large?
  4. What is the difference between a minor and a major page fault, and why is DMA used to service the major one?
  5. How does the Row Hammer attack turn a DRAM reliability problem into a way to compromise the page table?

Vocabulary

virtual memory (noun)
A system that gives each program the illusion of its own large memory, separate from the real physical memory.
Virtual memory lets a program act as if it has more space than actually exists.
physical memory (noun)
The real, limited memory hardware installed in a computer.
Physical memory is shared among all running programs.
address space (noun)
The full range of addresses a program can use to refer to memory.
Each process gets its own private address space.
page (noun)
A fixed-size chunk of a program's virtual memory.
A 4 KB page is a common size used in real systems.
frame (noun)
A fixed-size chunk of physical memory that can hold one page.
Each virtual page maps to one physical frame.
page table (noun)
A data structure that stores where each virtual page is actually located in physical memory.
The page table translates a virtual address into a physical one.
page table entry (noun)
One row of the page table describing a single page's mapping and status.
A page table entry includes a validity bit and the frame number.
multi-level page table (noun)
A page table organized as a tree of smaller tables to save space for unused memory regions.
A multi-level page table avoids allocating tables for unused address ranges.
page fault (noun)
An event that occurs when a requested virtual page is not currently mapped to physical memory.
A page fault triggers the operating system to load the missing page.
minor fault (noun)
A page fault that only needs a free memory frame, with no disk access.
A minor fault is much cheaper than fetching data from disk.
major fault (noun)
A page fault that requires reading the missing data from disk.
A major fault is expensive because it involves slow disk I/O.
direct memory access (DMA) (noun)
A method that lets a device move data to or from memory without using the processor.
DMA transfers the page from disk into memory during a major fault.
clock algorithm (noun)
A simple, low-cost way to approximate least-recently-used page replacement.
The clock algorithm helps choose which page to evict when memory is full.
memory protection (noun)
Rules enforced by hardware that stop one program from accessing another's memory.
Memory protection keeps processes isolated from each other.
translation lookaside buffer (TLB) (noun)
A small, fast cache that stores recently used virtual-to-physical address translations.
The TLB avoids a full page table lookup for most accesses.
memory management unit (MMU) (noun)
The hardware component responsible for translating virtual addresses to physical ones.
The MMU includes the TLB and a page table walker.
page table walker (noun)
Hardware that automatically searches the page table levels when a TLB lookup misses.
The page table walker fetches the entry when it's not in the TLB.
disaggregated memory (noun)
A system design where memory is physically separated from the processor and accessed over a network.
Disaggregated memory research explores new virtual memory designs.
illusion (noun)
Something that looks real but is not actually true.
Virtual memory creates the illusion of a large private address space.
contiguous (adjective)
Next to each other in a continuous, unbroken row.
Each process sees its address space as one contiguous block.
abstraction (noun)
A simplified way of viewing something complex, hiding the details underneath.
Virtual memory is an abstraction over the real physical memory.
validity bit (noun)
A single bit showing whether an entry currently holds usable data.
An invalid validity bit means the page is not in physical memory.
permission bit (noun)
A flag that controls whether a page can be read, written, or executed.
Permission bits stop code pages from being overwritten.
hierarchical (adjective)
Organized in ranked levels, from general to specific.
The page table is built as a hierarchical tree of smaller tables.
base register (noun)
A register holding the starting address used to begin a lookup.
CR3 is the base register that starts the page table walk.
anonymous allocation (phrase)
Memory given to a program that is not backed by any file on disk.
A minor fault often comes from an anonymous allocation.
evict (verb)
To remove something from a cache or memory to make room for something else.
The clock algorithm decides which page to evict.
isolate (verb)
To keep something separate so it cannot affect other things.
Memory protection isolates one process from another.
spray (verb)
To scatter something widely across an area.
An attacker can spray physical memory with page tables to increase their chance of success.
irregular access pattern (phrase)
A pattern of memory requests that does not follow a simple, predictable order.
Data-intensive workloads with irregular access patterns overwhelm the TLB.
unified memory (noun)
A design where CPU and GPU share the same memory space instead of separate ones.
Unified CPU-GPU memory is one direction of current research.
rethink (verb)
To think about something again in a new way.
Researchers are working to rethink virtual memory for modern workloads.
on the fly (idiom)
While something is happening, without stopping to plan in advance.
Translations are looked up on the fly during every memory access.

From the YouTube description

Digital Design and Computer Architecture, ETH Zürich, Spring 2025 (https://safari.ethz.ch/ddca/spring2025/)

Lecture 24: Virtual Memory
Lecturer: Prof. Onur Mutlu
Date: 23 May 2025

Lecture 24 Slides (pptx): https://safari.ethz.ch/ddca/spring2025/lib/exe/fetch.php?media=kanellok-ddca-2025-lecture-24-virtual-memory-before-lecture.pptx
Lecture 24 Slides (pdf): https://safari.ethz.ch/ddca/spring2025/lib/exe/fetch.php?media=kanellok-ddca-2025-lecture-24-virtual-memory-before-lecture.pdf

Recommended Reading:
====================
Intelligent Architectures for Intelligent Computing Systems
https://people.inf.ethz.ch/omutlu/pub/intelligent-architectures-for-intelligent-computingsystems-invited_paper_DATE21.pdf

A Modern Primer on Processing in Memory
https://people.inf.ethz.ch/omutlu/pub/ModernPrimerOnPIM_springer-emerging-computing-bookchapter21.pdf

RowHammer: A Retrospective
https://people.inf.ethz.ch/omutlu/pub/RowHammer-Retrospective_ieee_tcad19.pdf

RECOMMENDED LECTURE VIDEOS & PLAYLISTS:
========================================
Computer Architecture Fall 2021 Lectures Playlist:
https://www.youtube.com/watch?v=4yfkM_5EFgo&list=PL5Q2soXY2Zi-Mnk1PxjEIG32HAGILkTOF

Computer Architecture Fall 2022 Lectures Playlist:
https://www.youtube.com/watch?v=BIpPTqHK-Lc&list=PL5Q2soXY2Zi-cAls3cyauNzM7-74Eq31O

Digital Design and Computer Architecture Spring 2022 Livestream Lectures Playlist:
https://www.youtube.com/watch?v=cpXdE3HwvK0&list=PL5Q2soXY2Zi97Ya5DEUpMpO2bbAoaG7c6

Digital Design and Computer Architecture Spring 2021 Livestream Lectures Playlist:
https://www.youtube.com/watch?v=LbC0EZY8yw4&list=PL5Q2soXY2Zi_uej3aY39YB5pfW4SJ7LlN

Featured Lectures:
https://www.youtube.com/watch?v=jVYCchBGNVc&list=PL5Q2soXY2Zi8VrmOTz44l2WupethSdh-M&index=1

Interview with Professor Onur Mutlu:
https://www.youtube.com/watch?v=8ffSEKZhmvo&list=PL5Q2soXY2Zi8VrmOTz44l2WupethSdh-M&index=9

The Story of RowHammer Lecture:
https://www.youtube.com/watch?v=sgd7PHQQ1AI&list=PL5Q2soXY2Zi8D_5MGV6EnXEJHnV2YFBJl&index=39

Accelerating Genome Analysis Lecture:
https://www.youtube.com/watch?v=r7sn41lH-4A&list=PL5Q2soXY2Zi8D_5MGV6EnXEJHnV2YFBJl&index=41

Memory-Centric Computing Systems Tutorial at IEDM 2021:
https://www.youtube.com/watch?v=H3sEaINPBOE&list=PL5Q2soXY2Zi8D_5MGV6EnXEJHnV2YFBJl&index=35

Intelligent Architectures for Intelligent Machines Lecture:
https://www.youtube.com/watch?v=GTieZPY4Wmc&list=PL5Q2soXY2Zi8D_5MGV6EnXEJHnV2YFBJl&index=38

Computer Architecture Fall 2020 Lectures Playlist:
https://www.youtube.com/watch?v=c3mPdZA-Fmc&list=PL5Q2soXY2Zi9xidyIgBxUz7xRPS-wisBN

Digital Design and Computer Architecture Spring 2020 Lectures Playlist:
https://www.youtube.com/watch?v=AJBmIaUneB0&list=PL5Q2soXY2Zi_FRrloMa2fUYWPGiZUBQo2

Public Lectures by Onur Mutlu, Playlist:
https://www.youtube.com/watch?v=kgiZlSOcGFM&list=PL5Q2soXY2Zi8D_5MGV6EnXEJHnV2YFBJl

Computer Architecture at Carnegie Mellon Spring 2015 Lectures Playlist:
https://www.youtube.com/watch?v=zLP_X4wyHbY&list=PL5PHm2jkkXmi5CxxI7b3JCL1TWybTDtKq

Rethinking Memory System Design Lecture @stanfordonline :
https://www.youtube.com/watch?v=F7xZLNMIY1E&list=PL5Q2soXY2Zi8D_5MGV6EnXEJHnV2YFBJl&index=4

← Lecture 23b: Multi-Core Issues in Caching · Lecture 25: Prefetching II and Parting Thoughts →