I didn't get backpropagation the first time. Here's the version that finally made sense
By Seyed Masoud Hosseini · · Learning
My notes on Andrej Karpathy's two-and-a-half-hour micrograd lecture, the first course of my self-made bachelor. What a derivative really is, how backprop works by hand, the bug everyone hits, and a training loop in about twenty lines.
From my bachelor: Neural Networks: Zero to Hero · Lecture 1: The Spelled-Out Intro to Neural Networks and Backpropagation
The first lecture of my self-designed bachelor is Andrej Karpathy's The spelled-out intro to neural networks and backpropagation. It's 2 hours and 25 minutes long. I typed along with all of it, and at the end I had working code and only a rough idea of why it worked.
So I went back through it slowly, and this post is what I wish I had read before pressing play. If you're about to watch it, read this first. If you already watched it and felt lost somewhere around the one-hour mark, you're in good company.
One more honest note: English is not my first language, and Karpathy talks fast. Half my confusion on the first watch was words, not maths. That's why every lecture page on my bachelor tracker now has a vocabulary tab. For this one, learn derivative, gradient, chain rule and nudge before anything else.
What he builds
The whole lecture builds one small thing: micrograd, a tiny library that computes gradients automatically. The core engine is about a hundred lines of Python. On top of it sits a mini neural network library of about fifty more.
That's the point. PyTorch does the same job with far more code, but underneath it's the same idea. Once you've written it yourself, loss.backward() stops being magic.
Start with the derivative, not the network
He doesn't open with neurons. He opens with a boring function:
def f(x):
return 3*x**2 - 4*x + 5
At x = 3, f(x) is 20. Move x a tiny bit, to 3.001, and f becomes about 20.014. The output moved 0.014 for an input change of 0.001, so the slope is about 14.
That number is the derivative. Not a formula you memorise, but an answer to a very practical question: if I nudge this input a little, how much does the output move, and in which direction?
He does it numerically on purpose. Nobody writes out the symbolic derivative of a network with millions of parameters. What we want is that sensitivity number for every input, and the rest of the lecture is about getting it cheaply.
A number that remembers where it came from
Next comes a Value class. At first it just wraps a number. Then he teaches it + and *, and here's the clever part: every result remembers which values produced it, and with which operation.
a = Value(2.0)
b = Value(-3.0)
c = Value(10.0)
e = a * b # -6
d = e + c # 4
f = Value(-2.0)
L = d * f # -8
Because every value keeps pointers to its parents, the expression becomes a graph. He draws it with Graphviz, and seeing the boxes and arrows helped me more than any equation. A neural network is just a much bigger version of this picture.
Each Value also gets a grad field, starting at zero. It answers the question from the previous section: how much does the final output L change if I nudge this value?
Backprop by hand, once
This was the section I had to rewatch. He fills in every gradient manually, starting from the end.
L.gradis 1. NudgingLchangesLby the same amount.L = d * f, sod.gradis the value off(-2), andf.gradis the value ofd(4).
That's the multiplication rule. For x * y, the gradient of x is y, and the gradient of y is x. Each input gets the other one.
Then d = e + c. Addition has a local derivative of 1 for both inputs, so a plus node just passes the gradient through unchanged. e.grad and c.grad both become -2.
Last step: e = a * b. The multiply rule again, but now we also multiply by the gradient coming from above. That's the chain rule:
a.grad = b * e.grad = -3 * -2 = 6b.grad = a * e.grad = 2 * -2 = -4
And that's backpropagation. Seriously. Every node only needs its own small local derivative. Multiply it by the gradient arriving from above, pass the result down, and repeat until you reach the inputs.
He then checks each number by nudging the input by 0.001 and measuring, like at the start. That check matters. It's the difference between "I believe the formula" and "I've seen it work".
What I wrote on paper:
Plus passes the gradient through. Times swaps the inputs. The chain rule multiplies by whatever comes from above.
A neuron is the same thing with a squash at the end
A neuron multiplies each input by a weight, adds a bias, and squashes the result with tanh so the output lands between -1 and 1:
o = tanh(x1*w1 + x2*w2 + b)
He picks the numbers so the output comes out around 0.707. tanh can't be built from plus and times, so it gets its own local derivative, 1 - tanh(x)². With o ≈ 0.707 that's 1 - 0.5 = 0.5. From there, backprop to the weights works exactly as before.
What clicked for me here: a weight is just another input in the graph. It has a gradient like everything else. Training means changing the weights using their gradients. That's all.
Making it automatic
Doing this by hand is how you learn it, not how you use it. So each operation gets a small _backward function that knows its own local rule. Multiplication looks roughly like this:
def __mul__(self, other):
out = Value(self.data * other.data, (self, other), '*')
def _backward():
self.grad += other.data * out.grad
other.grad += self.data * out.grad
out._backward = _backward
return out
To call these in the right order he uses a topological sort: a list of the nodes where each one comes after everything it depends on. backward() sets the output's gradient to 1, walks that list in reverse, and calls _backward on each node.
The bug everyone hits
Look at the += above. His first version used =, and it breaks the moment a value is used twice:
a = Value(3.0)
b = a + a
b.backward()
The right answer for a.grad is 2, because a reaches b through two paths. With =, the second path overwrites the first and you get 1. Gradients accumulate. If you remember one bug from this lecture, make it this one.
From one neuron to a small network
With the engine done, the network code is short:
Neuron: some weights and a bias, random at the start.Layer: a list of neurons that all see the same inputs.MLP: layers chained one after another.
He builds MLP(3, [4, 4, 1]): three inputs, two hidden layers of four neurons, one output. That's 41 parameters. Tiny, but a real neural network.
Training, finally
The dataset is four examples of three numbers each, with targets 1, -1, -1, 1. The loss adds up the squared differences between predictions and targets. It's one number that says how wrong we are, and lower is better.
This loop is the part to memorise:
for step in range(20):
# 1. forward pass
preds = [model(x) for x in xs]
loss = sum((p - y)**2 for p, y in zip(preds, ys))
# 2. backward pass
for p in model.parameters():
p.grad = 0.0
loss.backward()
# 3. update
for p in model.parameters():
p.data += -0.05 * p.grad
Why the minus sign? The gradient points in the direction that makes the loss bigger. We want it smaller, so we step the other way. That's gradient descent. The 0.05 is the learning rate: too small and training crawls, too big and it jumps around and never settles.
Run it and the loss falls towards zero while the predictions move towards 1, -1, -1, 1.
The second bug, made live on camera
His first loop doesn't reset the gradients, and it still seems to work, which is exactly what makes it dangerous. Every backward() call adds to the old gradients (remember +=), so each update is based on the sum of all previous ones. He points out that this is one of the most common mistakes in real PyTorch code too. It's why zero_grad() exists.
What I'd tell myself before watching
- The first hour is the whole lecture. Derivatives, the
Valueclass and backprop by hand. If that part makes sense, the rest is packaging. - Do one backward pass on paper. Take
L = (a*b + c) * fand fill in every gradient yourself. Ten minutes, and worth more than a rewatch. - Type the code, don't just watch. Pause and guess what the next line will do.
- If the chain rule feels shaky, fix that first. 3Blue1Brown's calculus and neural network series are the best warm-up I know.
- Know the three rules cold. Plus passes the gradient through, times swaps the inputs, gradients accumulate.
Questions I still have
Writing this down also showed me what I don't know yet:
- Real networks have millions of parameters. One Python object per number can't be how PyTorch does it. (It isn't. It works on whole tensors at once, which later lectures get to.)
- Why do deeper networks prefer ReLU over
tanh? - How do you pick a learning rate, other than trial and error?
Those go on the list for the next lectures. For now, the thing that felt like magic is a hundred lines of code I can explain on a whiteboard, and that's a decent first day of university.
If you want to follow along, the lecture page has the video, a study guide, the vocabulary list and clickable timestamps. The rest of the course is on my bachelor tracker.