Seyed Masoud Hosseini · Overview · Study log · Weekly summaries · Ideas · Search · Transcript · RSS feed

Study log

Lecture 2 — building makemore

Sep 17, 2026 · Neural Networks: Zero to Hero · 2 h

What I did

Watched the second lecture of Neural Networks: Zero to Hero, where Karpathy builds makemore, a character-level model that makes up new names. He builds a bigram model twice: first by counting, then as a tiny neural network.

What I learned

  • A bigram model predicts the next character from only the previous one. Counting how often each pair appears (a 27×27 table) and turning each row into probabilities is already "training".
  • The loss is the negative log likelihood: the average of -log(p) for the characters that really came next. Lower is better.
  • The neural network version gets the same result: one-hot inputs, a linear layer that outputs logits, then softmax turns them into probabilities.
  • Smoothing (adding fake counts) or regularisation stops the model from giving zero probability to unseen pairs, which would make the loss infinite.
  • Broadcasting bugs are silent: forgetting keepdim=True normalises the wrong axis and still looks like it works.

Day one — building micrograd from scratch

Sep 15, 2026 · Neural Networks: Zero to Hero · 2.5 h

What I did

Watched the first lecture of Neural Networks: Zero to Hero and rebuilt micrograd alongside it: a Value class that records the operations applied to it, a topological sort, and a backward() pass that fills in gradients.

What I learned

  • Backpropagation is just the chain rule applied recursively over a DAG, in reverse topological order.
  • Gradients must accumulate (+=), not overwrite. A node used twice gets a contribution from each path.
  • A neuron is tanh(w · x + b); an MLP is layers of those. Everything else is bookkeeping.

Open questions

  • How do real frameworks avoid building a Python object per scalar?
  • Why does tanh fall out of favour against ReLU in deeper networks?