Lecture 2 — building makemore
Sep 17, 2026 · Neural Networks: Zero to Hero · 2 h
What I did
Watched the second lecture of Neural Networks: Zero to Hero, where Karpathy builds makemore, a character-level model that makes up new names. He builds a bigram model twice: first by counting, then as a tiny neural network.
What I learned
- A bigram model predicts the next character from only the previous one. Counting how often each pair appears (a 27×27 table) and turning each row into probabilities is already "training".
- The loss is the negative log likelihood: the average of
-log(p)for the characters that really came next. Lower is better. - The neural network version gets the same result: one-hot inputs, a linear layer that outputs logits, then softmax turns them into probabilities.
- Smoothing (adding fake counts) or regularisation stops the model from giving zero probability to unseen pairs, which would make the loss infinite.
- Broadcasting bugs are silent: forgetting
keepdim=Truenormalises the wrong axis and still looks like it works.