Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Neural Networks: Zero to Hero · Lecture 8 of 10 · 42:40

Lecture 8: State of GPT

State of GPT | BRK216HFS on YouTube

Study guide

What this lecture covers

This talk, given at Microsoft Build, answers two questions: how are models like ChatGPT actually trained, and how should you use them once they exist? It's a conference presentation rather than a coded walkthrough, and it pairs well with the "Let's build GPT" lecture earlier in this course by zooming out from the Transformer architecture to the full training pipeline and to practical prompting technique.

The talk is split into two halves. The first covers the four-stage recipe for producing a GPT assistant: pretraining, supervised fine-tuning, reward modeling, and reinforcement learning from human feedback. The second half covers how to think about prompting these models effectively, including techniques like chain-of-thought prompting, retrieval augmentation, and tool use. After watching, you should be able to describe what happens at each training stage and apply several concrete prompting strategies to your own use of LLMs.

Key ideas

  • Pretraining dominates the compute budget: it uses internet-scale data, thousands of GPUs, and months of training, while the later fine-tuning stages use much smaller datasets and far less compute.
  • Tokenization: raw text is losslessly translated into integer tokens (for example via byte pair encoding) before it can feed into a Transformer.
  • Base models complete documents, they don't answer questions: a freshly pretrained model imitates whatever text pattern it's given rather than acting as a helpful assistant.
  • Supervised fine-tuning (SFT): a small, high-quality dataset of prompt-and-ideal-response pairs, written by human contractors, is used to continue language-model training and turn a base model into a basic assistant.
  • Reward modeling and RLHF: humans rank multiple model completions for the same prompt; a reward model learns to predict these rankings, and reinforcement learning then tunes the assistant to produce completions the reward model scores highly.
  • Transformers are token simulators: they spend roughly equal computation per token, don't know what they don't know, and can't revise earlier tokens, so techniques like "think step by step" work by spreading reasoning across more tokens.
  • Retrieval augmentation and tool use: loading relevant documents or calculator/code-interpreter access into the context window compensates for the model's imperfect recall and weak arithmetic.

Walkthrough

The four-stage training recipe (0:00)

Karpathy introduces pretraining, SFT, reward modeling, and RLHF as four stages that each have their own dataset, training objective, and resulting model, with pretraining set apart as the most compute-intensive by a wide margin.

Pretraining: data, tokenization, and scale (2:03)

Using LLaMA's published data mixture as an example (CommonCrawl, GitHub, Wikipedia, books, and more), he walks through tokenization, typical hyperparameters like vocabulary size and context length, and why LLaMA outperforms the larger GPT-3 despite fewer parameters: it was trained on far more tokens. He also shows what pretraining looks like in practice, with a small GPT trained on Shakespeare gradually producing more coherent text as loss decreases.

From base model to assistant: SFT (9:09)

Base models can be prompted cleverly (few-shot prompting, or framing a fake Q&A document) to behave somewhat like assistants, but this is unreliable. Supervised fine-tuning instead trains directly on a small set of high-quality prompt-and-ideal-response demonstrations written by contractors, producing a model that behaves like a real, if limited, assistant.

Reward modeling and reinforcement learning (13:11)

Rather than asking contractors to write ideal answers directly, this stage asks them to rank multiple completions from the SFT model, since comparing outputs is easier than generating them. A reward model is trained on these rankings, and reinforcement learning then optimizes the assistant's policy to produce completions the reward model scores highly, which is the process behind models like ChatGPT.

Why RLHF helps, and where it falls short (17:14)

RLHF models are generally preferred by human evaluators over SFT models, partly because comparing outputs is computationally and cognitively easier than generating them well. But RLHF models lose some output diversity compared to base models, which is why Karpathy still prefers base models for tasks like generating many varied examples from a small seed set.

Thinking like the model: prompting technique (20:17)

Using the example of writing a sentence comparing two states' populations, Karpathy contrasts the rich internal reasoning a human does with a Transformer's fixed compute-per-token budget. This motivates techniques such as "let's think step by step," self-consistency (sampling multiple times and picking the best), and asking the model to check its own work, since models don't automatically know when they've made a mistake.

Tools, retrieval, and constrained output (31:27)

Because models don't reliably know their own weaknesses, prompts can explicitly tell them to use a calculator or other tools. Retrieval-augmented generation, illustrated with LlamaIndex, loads relevant document chunks into the context window so the model can reference primary sources instead of relying on imperfect memory. Constrained prompting libraries like Microsoft's Guidance can force outputs into a specific format, such as valid JSON.

Fine-tuning and closing recommendations (34:30)

Parameter-efficient fine-tuning techniques such as LoRA make it cheaper to adapt a model by training only small pieces of it. Karpathy recommends defaulting to prompt engineering with GPT-4 first, only considering fine-tuning afterward, and treating RLHF as largely research territory. He closes with a caution about current LLM limitations, including bias, hallucination, and susceptibility to prompt injection, recommending human oversight for any high-stakes use.

Before you watch

  • Watching "Let's build GPT" earlier in this series first will make the architecture references (tokens, context length, Transformer layers) much more concrete.
  • Basic familiarity with what a language model does and what a probability distribution over tokens means is assumed.

Check your understanding

  1. Why does pretraining use the vast majority of the total compute in training a GPT assistant?
  2. What is the practical difference between the datasets used for supervised fine-tuning and reward modeling?
  3. Why might comparing two completions be an easier task for a human than writing one from scratch, and how does this asymmetry help RLHF?
  4. Give two concrete prompting techniques the lecture describes for getting a model to reason more effectively, and explain why each works.
  5. Why does the lecture recommend keeping humans in the loop rather than treating current LLMs as fully autonomous agents?

Vocabulary

assistant (noun)
A model trained to follow instructions and help a user, rather than just complete text.
The talk explains how a base model becomes a helpful assistant.
pretraining (noun)
The first, largest training stage where a model learns from huge amounts of general text.
Pretraining uses far more compute than the later stages.
compute budget (noun)
The total amount of computing power and time available for a task.
Pretraining uses almost the entire compute budget.
tokenization (noun)
The process of converting text into integer tokens a model can process.
Byte pair encoding is one common tokenization method.
base model (noun)
A model that has only been pretrained, without extra instruction-following training.
A base model just completes documents rather than answering questions.
supervised fine-tuning (SFT) (noun)
Training a model further on a small set of high-quality example answers.
SFT turns the base model into a basic assistant.
contractor (noun)
A hired worker who is not a permanent employee, often doing a specific task.
Contractors write the ideal prompt-and-response examples for SFT.
reward model (noun)
A model trained to predict how good a response is, based on human rankings.
The reward model learns to score completions the way humans would.
reinforcement learning (noun)
A training method where a model improves by getting feedback on its actions.
Reinforcement learning tunes the assistant using the reward model's scores.
RLHF (reinforcement learning from human feedback) (noun)
A training process that uses human preferences to guide reinforcement learning.
RLHF is the final stage behind models like ChatGPT.
rank (verb)
To put items in order from best to worst.
Humans rank several model completions for the same prompt.
diversity (noun)
The amount of variety among different outputs.
RLHF models lose some output diversity compared to base models.
few-shot prompting (noun)
Giving a model a few examples in the prompt to show it what kind of answer is wanted.
Few-shot prompting can make a base model act more like an assistant.
step by step (phrase)
One small stage at a time, in order.
Asking the model to think step by step improves its answers.
self-consistency (noun)
A technique of generating several answers and choosing the most common or best one.
Self-consistency samples multiple answers and picks the best.
retrieval-augmented generation (noun)
A technique that gives a model relevant documents to read before it answers.
Retrieval-augmented generation loads document chunks into the context window.
context window (noun)
The maximum amount of text a model can consider at once.
Relevant documents are loaded into the model's context window.
tool use (noun)
Letting a model call external programs, like a calculator, to help answer.
Tool use compensates for the model's weak arithmetic.
constrained output (noun)
Forcing a model's response to follow a specific format, like valid JSON.
Constrained output libraries force the model to produce valid JSON.
parameter-efficient fine-tuning (noun)
A method of adapting a model by training only a small part of it, saving cost.
LoRA is a popular parameter-efficient fine-tuning technique.
prompt engineering (noun)
The practice of carefully wording instructions to get better answers from a model.
The talk recommends prompt engineering before fine-tuning.
hallucination (noun)
When a model confidently states false information as if it were true.
Hallucination is listed as a current limitation of LLMs.
prompt injection (noun)
An attack where hidden instructions in input text try to hijack a model's behavior.
Prompt injection is mentioned as a security risk for LLMs.
high-stakes (adjective)
Describes a situation where mistakes could cause serious harm.
Human oversight is recommended for high-stakes uses of LLMs.

Chapters

From the YouTube description

Learn about the training pipeline of GPT assistants like ChatGPT, from tokenization to pretraining, supervised finetuning, and Reinforcement Learning from Human Feedback (RLHF). Dive deeper into practical techniques and mental models for the effective use of these models, including prompting strategies, finetuning, the rapidly growing ecosystem of tools, and their future extensions.

*Speakers:*
* Andrej Karpathy

*Session Information:*
This video is one of many sessions delivered for the Microsoft Build 2023 event. View the full session schedule and learn more about Microsoft Build at https://build.microsoft.com

BRK216HFS | English (US) | AI

#MSBuild

← Lecture 7: Let's Build GPT From Scratch · Lecture 9: Building the GPT Tokenizer →