Seyed Masoud Hosseini · Overview · Study log · Ideas · Transcript · RSS feed

Digital Design & Computer Architecture · Lecture 19 of 37 · 24:02

Lecture 15c: Load-Store Handling in Out-of-Order Execution

Digital Design and Comp. Arch. - L15c: Load-Store Handling in Out-of-Order Execution (Spring 2025) on YouTube

Study guide

What this lecture covers

This lecture explains why handling memory instructions is the hardest part of building an out-of-order execution engine, harder than the register renaming covered in earlier lectures. It answers a specific question: since a load or store's address isn't known until the instruction partially executes, how can hardware still schedule loads early and give them the correct value when an earlier store might write to the same address?

The lecture builds on the reorder buffer and register renaming material from prior lectures in this course. After watching, you should understand why memory dependences can't be resolved at decode time the way register dependences can, the three broad strategies for scheduling loads relative to unknown stores, and how the store queue's content-addressable search works to forward data from a store directly to a dependent load.

Key ideas

  • Memory disambiguation problem: determining whether a load depends on an earlier store (with an unknown address) requires comparing addresses that aren't available until execution.
  • Conservative approach: stall every load until all older stores have computed their addresses (or retired); correct but slow.
  • Aggressive approach: schedule loads immediately, assuming independence from prior stores, then check and recover (flush) if wrong.
  • Intelligent approach: predict load-store dependence based on past behavior, since the same dependence tends to repeat across loop iterations.
  • Store queue (store buffer): an in-order, hardware list of in-flight stores holding address, data, and validity bits, searched by later loads.
  • Store-to-load forwarding: when a load's address matches a valid store address, the load gets its value directly from that store instead of from memory.
  • Content-addressable search: the store queue search must match on address range and size, not just an exact address, because loads and stores can partially overlap.

Walkthrough

Why memory is harder than registers (0:03)

The lecture opens by contrasting registers and memory as sources of complexity in out-of-order machines. Register dependences are known statically, right after decode, so renaming can happen early and in order. Memory addresses are only known after an instruction partially executes, addresses are far larger than register identifiers, and in multiprocessors memory is shared across threads in a way registers are not.

The unknown-address problem (3:03)

Using a store followed by a load example, the lecture shows that if the store is stalled waiting on its source registers, its address is unknown when a younger, ready load wants to execute. The hardware cannot tell whether the load depends on that store, yet still wants high performance, which is the core tension the rest of the lecture addresses.

Conservative, aggressive, and predicted scheduling (7:07)

Three strategies are compared: stalling a load until all older stores resolve (conservative, low performance), scheduling loads immediately and recovering on misprediction (aggressive), and predicting dependence based on history (intelligent, used in real machines). A chart of instructions-per-cycle across benchmarks shows the conservative approach performing far worse than the aggressive one, with a large remaining gap to a perfect predictor, which is why simple prediction schemes capture most of the available performance.

Store-to-load forwarding and the load/store queues (14:11)

Even assuming all store addresses were known, the machine still needs a mechanism to check whether a load depends on a store and to forward that store's data to the load. This is handled by dedicated structures, typically a store queue and a load queue (combined into a single "memory ordering buffer" in the Pentium Pro), each entry tracking address, data, and separate validity bits for address and data.

The store-to-load forwarding search logic (17:15)

A load searches the store queue with a wide comparator against every in-flight store's address. Because a load can read bytes written by several different stores, the search must match by address range and size, not equality alone, and must select the youngest matching store for each byte. If no store covers a byte, the load must still access memory for it.

Before you watch

  • Be comfortable with the reorder buffer and register renaming, covered in the preceding out-of-order execution lectures.
  • Recall what content-addressable memory structures do, introduced in the previous lecture on scheduling logic.

Check your understanding

  1. Why can't memory addresses be renamed at decode time the way register operands can?
  2. What is the tradeoff between the conservative and aggressive approaches to scheduling loads relative to unknown stores?
  3. Why does a store-to-load forwarding search need to compare address ranges and sizes rather than just checking for an exact address match?
  4. Why do real processors keep the store buffer much smaller than the instruction window?

From the YouTube description

Digital Design and Computer Architecture, ETH Zürich, Spring 2025 (https://safari.ethz.ch/ddca/spring2025/)

Lecture 15c: Load-Store Handling in Out-of-Order Execution
Lecturer: Prof. Onur Mutlu
Date: 15 April 2025

Lecture 15c Slides (pptx): https://safari.ethz.ch/ddca/spring2025/lib/exe/fetch.php?media=onur-ddca-2025-lecture15c-load-store-handling-in-out-of-order-execution.pptx
Lecture 15c Slides (pdf): https://safari.ethz.ch/ddca/spring2025/lib/exe/fetch.php?media=onur-ddca-2025-lecture15c-load-store-handling-in-out-of-order-execution.pdf

Recommended Reading:
====================
Intelligent Architectures for Intelligent Computing Systems
https://people.inf.ethz.ch/omutlu/pub/intelligent-architectures-for-intelligent-computingsystems-invited_paper_DATE21.pdf

A Modern Primer on Processing in Memory
https://people.inf.ethz.ch/omutlu/pub/ModernPrimerOnPIM_springer-emerging-computing-bookchapter21.pdf

RowHammer: A Retrospective
https://people.inf.ethz.ch/omutlu/pub/RowHammer-Retrospective_ieee_tcad19.pdf

RECOMMENDED LECTURE VIDEOS & PLAYLISTS:
========================================
Computer Architecture Fall 2021 Lectures Playlist:
https://www.youtube.com/watch?v=4yfkM_5EFgo&list=PL5Q2soXY2Zi-Mnk1PxjEIG32HAGILkTOF

Computer Architecture Fall 2022 Lectures Playlist:
https://www.youtube.com/watch?v=BIpPTqHK-Lc&list=PL5Q2soXY2Zi-cAls3cyauNzM7-74Eq31O

Digital Design and Computer Architecture Spring 2022 Livestream Lectures Playlist:
https://www.youtube.com/watch?v=cpXdE3HwvK0&list=PL5Q2soXY2Zi97Ya5DEUpMpO2bbAoaG7c6

Digital Design and Computer Architecture Spring 2021 Livestream Lectures Playlist:
https://www.youtube.com/watch?v=LbC0EZY8yw4&list=PL5Q2soXY2Zi_uej3aY39YB5pfW4SJ7LlN

Featured Lectures:
https://www.youtube.com/watch?v=jVYCchBGNVc&list=PL5Q2soXY2Zi8VrmOTz44l2WupethSdh-M&index=1

Interview with Professor Onur Mutlu:
https://www.youtube.com/watch?v=8ffSEKZhmvo&list=PL5Q2soXY2Zi8VrmOTz44l2WupethSdh-M&index=9

The Story of RowHammer Lecture:
https://www.youtube.com/watch?v=sgd7PHQQ1AI&list=PL5Q2soXY2Zi8D_5MGV6EnXEJHnV2YFBJl&index=39

Accelerating Genome Analysis Lecture:
https://www.youtube.com/watch?v=r7sn41lH-4A&list=PL5Q2soXY2Zi8D_5MGV6EnXEJHnV2YFBJl&index=41

Memory-Centric Computing Systems Tutorial at IEDM 2021:
https://www.youtube.com/watch?v=H3sEaINPBOE&list=PL5Q2soXY2Zi8D_5MGV6EnXEJHnV2YFBJl&index=35

Intelligent Architectures for Intelligent Machines Lecture:
https://www.youtube.com/watch?v=GTieZPY4Wmc&list=PL5Q2soXY2Zi8D_5MGV6EnXEJHnV2YFBJl&index=38

Computer Architecture Fall 2020 Lectures Playlist:
https://www.youtube.com/watch?v=c3mPdZA-Fmc&list=PL5Q2soXY2Zi9xidyIgBxUz7xRPS-wisBN

Digital Design and Computer Architecture Spring 2020 Lectures Playlist:
https://www.youtube.com/watch?v=AJBmIaUneB0&list=PL5Q2soXY2Zi_FRrloMa2fUYWPGiZUBQo2

Public Lectures by Onur Mutlu, Playlist:
https://www.youtube.com/watch?v=kgiZlSOcGFM&list=PL5Q2soXY2Zi8D_5MGV6EnXEJHnV2YFBJl

Computer Architecture at Carnegie Mellon Spring 2015 Lectures Playlist:
https://www.youtube.com/watch?v=zLP_X4wyHbY&list=PL5PHm2jkkXmi5CxxI7b3JCL1TWybTDtKq

Rethinking Memory System Design Lecture @stanfordonline :
https://www.youtube.com/watch?v=F7xZLNMIY1E&list=PL5Q2soXY2Zi8D_5MGV6EnXEJHnV2YFBJl&index=4

← Lecture 15: Dataflow, Superscalar Execution and Branch Prediction · Lecture 16: Advanced Branch Prediction →