Backpropagation & Gradient Descent
Backpropagation & Gradient Descent
So, we have a machine with a collection of adjustable weights. But how does it know which ones to change? Imagine our cricket-playing machine again. It looks at the weather, the rain, the ground, and all the other information it has been given. It calculates its weighted combination. And it makes a prediction. **PLAY. ** But the correct answer was: **DON'T PLAY. ** The machine was wrong. Now comes the interesting part. We need a way to tell the machine not just that it was wrong, but **how wrong it was**. That measure of wrongness is called a **loss**. You can think of loss as a score that tells the machine how far its prediction is from the answer it was supposed to produce. A small mistake produces a small loss. A large mistake produces a large loss. And now the machine has something it didn't have before: A measurable signal telling it how badly it performed. But knowing that we are wrong isn't enough. Imagine discovering that you got the wrong answer on an exam. Knowing the answer is wrong doesn't tell you which part of your thinking caused the mistake. A neural network has the same problem. It may have thousands of connections. If the final answer is wrong, which connections are responsible? This is where **backpropagation** becomes important. The word itself gives us a clue. **Back. Propagation. ** The network first moves information forward. Inputs enter. They pass through the network. A prediction comes out. Then the error moves backward through the network. The system works backwards, tracing how much each connection contributed to the final mistake. Not all connections are equally responsible. Some barely affected the answer. Others had a much greater influence. Backpropagation helps calculate those contributions. And now we have something incredibly powerful. The machine knows: **What it predicted. ** **What it should have predicted. ** **How wrong it was. ** And approximately: **Which weights need to change, and in which direction. ** But how much should each weight change? This is where another idea enters. **Gradient descent. ** Imagine standing somewhere on a landscape covered in hills and valleys. Your goal is to reach the lowest point. But you can't see the entire landscape. You can only feel which direction slopes downward from where you are standing. So you take a small step downhill. Then another. And another. Eventually, you hope to reach a low point. Training a neural network works with a similar idea. The landscape represents the possible values of all those weights. The height represents the loss. The machine tries to move toward regions where the loss becomes smaller. That is the basic intuition behind **gradient descent**. And then the cycle begins again. Make a prediction. Measure the loss. Send information backwards. Adjust the weights. Make another prediction. Measure again. Adjust again. Millions of times. Sometimes the change is tiny. Sometimes it is larger. But gradually, the network begins to discover a configuration of weights that produces better answers. This is what we mean when we say: **The machine is learning. ** It isn't learning because someone manually rewrote its rules. It is learning because its internal parameters are repeatedly adjusted in response to experience. And this gives us a remarkable new way to think about neural networks. The network isn't simply a collection of neurons. It is a landscape of possibilities, constantly being reshaped by feedback. But there was a problem. We now had a way to train neural networks. We knew how to adjust their weights. So why didn't this immediately produce today's extraordinary AI? Why did we have to wait decades before neural networks became truly powerful? The answer lies outside the neural network itself. We needed something the early researchers simply didn't have enough of. **Data. ** **Computing power. ** And algorithms capable of taking advantage of both. And when those three forces finally began to converge... AI was about to wake up again.