Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed
NLP with Deep Learning · Lecture 8 of 23 · 1:17:03
Lecture 8: Self-Attention and the Transformer
Study guide
What this lecture covers
This lecture asks why recurrent neural networks, even with attention bolted on, were replaced almost entirely by a new architecture built purely on attention: the Transformer. It starts by diagnosing two specific weaknesses of RNNs, long-range dependency learning and parallelization, then builds up self-attention as a replacement from first principles, fixing each gap it has (no sense of order, no nonlinearity, no way to prevent looking at the future) one at a time.
The second half assembles these pieces into the full Transformer decoder, adds multi-head attention, residual connections, and layer normalization, and then covers the encoder and encoder-decoder variants along with the original machine translation results. After watching, you should be able to explain why self-attention parallelizes better than recurrence, compute what a single self-attention operation does, and describe the role of each component in a Transformer block.
Key ideas
- Linear interaction distance: in an RNN, two words separated by many positions must interact through many sequential steps, making long-range dependencies hard to learn even with LSTMs.
- Non-parallelizability: an RNN's hidden state at each time step depends on the previous one, forcing O(sequence length) sequential operations that cannot be parallelized on a GPU.
- Self-attention as fuzzy lookup: each word produces a query, and attention softly matches that query against every word's key to produce a weighted average of value vectors, letting any two words interact directly regardless of distance.
- Query, key, value matrices: three learned weight matrices transform each word's embedding into a query, key, and value vector; attention scores come from dot products between queries and keys.
- Position representations: because self-attention has no inherent notion of order, position information (sinusoidal or learned) is added to word embeddings so the model can distinguish word order.
- Masking: for decoders and language modeling, attention scores to future positions are set to negative infinity before the softmax, so a word's representation cannot use information from words that come after it.
- Multi-head attention: running several smaller attention operations in parallel, each with its own query, key, and value matrices, lets the model attend to different kinds of relationships at once.
- Residual connections and layer normalization: adding a layer's input back to its output helps gradients flow during training, and layer normalization rescales each vector's values to stabilize training further.
Walkthrough
Why RNNs limit long-range learning and parallelization (4:08)
The lecture identifies two structural problems with RNNs. First, linear interaction distance: because RNNs process words one step at a time, two related words that are far apart in a sentence must interact through many applications of the recurrent weight matrix, making such dependencies hard to learn even with LSTMs' improved gradient flow. Second, non-parallelizability: computing the hidden state at time step 5 requires first computing steps 1 through 4, so the number of unparallelizable operations grows with sequence length, which wastes the parallel computation that GPUs are good at. Attention, previously used only to connect a decoder to an encoder, is introduced as a way to solve both problems by letting any word interact with any other word in a single, parallelizable step.
Self-attention as a fuzzy lookup (10:11)
Attention is described as a soft version of a key-value lookup table: instead of an exact match between a query and one key, a query is compared against every key, producing similarity scores that are turned into weights via softmax, and the output is a weighted sum of the corresponding values. The lecture illustrates this with a toy example, representing the word "learned" in a sentence by attending, to varying degrees, to other words in the same sentence such as "Stanford" and "CS224N."
Computing queries, keys, and values (14:15)
Formally, each word embedding is transformed by three separate learned matrices into a query, a key, and a value vector. Attention scores are the dot product of a query with every key in the sequence; a softmax over these scores gives attention weights, and the output for a word is the weighted sum of all value vectors using those weights. Using separate query and key matrices, rather than comparing raw embeddings directly, effectively gives a low-rank, learnable way of deciding what should attend to what, including whether a word should attend to itself.
Fixing self-attention's gaps: position, nonlinearity, and masking (21:23)
Plain self-attention has three problems that need fixing before it can replace RNNs. It has no notion of sequence order, since it operates on a set of vectors, so position vectors (either fixed sinusoidal patterns or learned parameters) are added to word embeddings at the input, though learned position vectors limit the model to a maximum sequence length. It has no nonlinearity, since stacking self-attention layers alone just re-averages value vectors, so a feed-forward network is applied independently to each word's output after attention. Finally, for tasks like language modeling where the model must not see future words, attention scores to future positions are masked to negative infinity before the softmax so their weight becomes zero; this masking is used in decoders but not in encoders, which are allowed to see the whole input.
Multi-head attention and scaled dot-product attention (41:44)
A single self-attention operation averages information for one reason at a time, but a word may need to attend to different other words for different reasons at once, such as syntactic role versus topical meaning. Multi-head attention runs several attention operations in parallel, each with its own smaller query, key, and value matrices projecting into a lower-dimensional space, and concatenates their outputs before a final linear transformation combines them. This computation is implemented efficiently as matrix operations over the whole sequence at once, called the sequence-stacked form. Because dot products between vectors grow large as dimensionality increases, which can shrink softmax gradients, scores are divided by a constant based on the model dimensionality, a fix called scaled dot-product attention.
Residual connections, layer normalization, and the full decoder block (58:00)
Two additional optimization tricks are needed for a full Transformer block. Residual connections add each sublayer's input back to its output, giving gradients a direct path through the network and helping avoid vanishing gradients. Layer normalization rescales each word vector to have roughly unit mean and standard deviation (computed independently per word, not shared across the batch or sequence), which helps stabilize training. A Transformer decoder block applies masked multi-head self-attention, then a residual connection and layer normalization, then a feed-forward layer, then another residual connection and layer normalization, and this block is repeated several times to form the full decoder.
Encoder, encoder-decoder, and Transformer results (1:09:09)
A Transformer encoder is nearly identical to the decoder but omits masking, allowing bidirectional context. The original "Attention Is All You Need" architecture is an encoder-decoder model, where the decoder adds a cross-attention step: its own vectors form the queries, while the encoder's output vectors supply the keys and values, letting every decoder position attend over the whole encoded source sentence. Transformers achieved competitive machine translation results while training far more efficiently than RNN-based systems, because parallelization allowed much more data and compute to be used, which later enabled the large-scale pre-training covered in the next lecture. The lecture closes by noting a major drawback: self-attention's compute and memory cost grows quadratically with sequence length, which remains an active area of research, alongside other proposed modifications to position representations and the architecture generally.
Before you watch
- Watch the previous lecture on attention in encoder-decoder machine translation, since this lecture builds directly on that attention mechanism and assumes you know softmax-weighted averaging.
- Be comfortable with matrix multiplication and how linear layers transform vectors, since query, key, and value projections are presented as matrix operations throughout.
- Review vanishing gradients and residual connections from the LSTM lecture, since they motivate why residual connections are used in the Transformer.
Check your understanding
- What two specific problems with RNNs does self-attention solve, and how does removing recurrence solve each one?
- Walk through how a self-attention output is computed for one word, from query/key/value projections to the final weighted sum.
- Why does self-attention need explicit position representations, and what is the tradeoff between sinusoidal and learned position embeddings?
- Why is masking used in a Transformer decoder but not in a Transformer encoder, and how is masking implemented mathematically?
- What is the purpose of having multiple attention heads instead of one, and what is the main computational drawback of self-attention as sequence length grows?
Vocabulary
- self-attention (noun)
- A mechanism where each word in a sequence attends to every other word in the same sequence.
Self-attention lets any two words interact directly, regardless of distance. - linear interaction distance (noun)
- The number of steps two words must pass through to interact in a sequential model.
Linear interaction distance makes long-range dependencies hard for RNNs. - parallelization (noun)
- Doing many calculations at the same time instead of one after another.
Parallelization on a GPU makes training much faster. - fuzzy lookup (noun)
- A search that returns a blend of results based on similarity, rather than one exact match.
Self-attention works like a fuzzy lookup across all words at once. - query (noun)
- A vector representing what a word is looking for in an attention operation.
Each word produces a query used to search for relevant information. - key (noun)
- A vector representing what a word offers, compared against queries in attention.
The query is matched against every word's key to score relevance. - value (noun)
- A vector holding the actual information passed along once attention weights are computed.
The final output is a weighted sum of value vectors. - position representation (noun)
- Information added to a word's embedding that encodes its position in the sequence.
Position representations let self-attention understand word order. - sinusoidal (adjective)
- Based on sine and cosine wave patterns.
Sinusoidal position encodings use wave patterns to represent word order. - masking (noun)
- Blocking a model from seeing certain information, such as future words.
Masking prevents the decoder from looking ahead at future words. - multi-head attention (noun)
- Running several attention operations in parallel to capture different relationships.
Multi-head attention lets the model attend to multiple patterns at once. - scaled dot-product attention (noun)
- Dot-product attention divided by a constant to keep gradients stable.
Scaled dot-product attention prevents scores from growing too large. - residual connection (noun)
- A shortcut that adds a layer's input directly to its output.
A residual connection helps gradients flow through deep networks. - layer normalization (noun)
- A technique that rescales a vector's values to stabilize training.
Layer normalization is applied after each sublayer in a Transformer. - feed-forward network (noun)
- A simple neural network layer that processes each input independently.
A feed-forward network adds nonlinearity after the attention layer. - cross-attention (noun)
- An attention mechanism where queries come from one sequence and keys/values come from another.
Cross-attention lets the decoder attend over the encoder's output. - quadratic (adjective)
- Growing in proportion to the square of a quantity.
Self-attention's cost grows quadratically with sequence length. - bidirectional context (noun)
- Information from both earlier and later parts of a sequence.
A Transformer encoder can use bidirectional context freely. - diagnose (verb)
- To identify the exact cause of a problem.
The lecture diagnoses two specific weaknesses of RNNs. - recurrence (noun)
- The property of a network reusing its own previous output as input.
Self-attention removes the need for recurrence entirely. - gradient flow (noun)
- How well training signals pass backward through a network's layers.
LSTMs improve gradient flow but don't fully solve long-range dependencies. - toy example (noun)
- A small, simplified example used to explain an idea clearly.
The lecture uses a toy example to illustrate self-attention. - learnable (adjective)
- Able to be adjusted automatically by training rather than fixed by hand.
The query and key matrices give a learnable way to decide focus. - syntactic role (noun)
- The grammatical function a word plays in a sentence.
One attention head might focus on a word's syntactic role. - dimensionality (noun)
- The number of values or features in a vector.
Large dimensionality can make dot products grow too big. - unit mean (phrase)
- An average value close to a standard size, often used loosely for zero or one.
Layer normalization rescales vectors toward roughly unit mean. - drawback (noun)
- A disadvantage or downside of something.
The main drawback of self-attention is its quadratic cost. - active area of research (phrase)
- A topic that researchers are still actively working to improve.
Reducing self-attention's cost remains an active area of research. - competitive (adjective)
- Good enough to compare favorably with other strong results.
Transformers achieved competitive machine translation results.
From the YouTube description
For more information about Stanford's Artificial Intelligence professional and graduate programs, visit: https://stanford.io/ai
This lecture covers:
1. From recurrence (RNN) to attention-based NLP models
2. The Transformer model
3. Great results with Transformers
4. Drawbacks and variants of Transformers
To learn more about this course visit: https://online.stanford.edu/courses/c...
To follow along with the course schedule and syllabus visit: http://web.stanford.edu/class/cs224n/
John Hewitt
https://nlp.stanford.edu/~johnhew/
Professor Christopher Manning
Thomas M. Siebel Professor in Machine Learning, Professor of Linguistics and of Computer Science
Director, Stanford Artificial Intelligence Laboratory (SAIL)
#naturallanguageprocessing #deeplearning
← Lecture 7: Attention and Choosing a Final Project · Lecture 9: Pretraining →
