Skip to main content

4.2. Backpropagate Through the Final Neuron

723

4. Backpropagation

Backpropagation means calculating:

dLdparameter\boxed{\frac{dL}{d\text{parameter}}}

for every weight and bias.

We start at the loss and move backward.

L
↓
y
↓
a1, a2, a3
↓
z1, z2, z3
↓
hidden weights and biases

The chain rule is the main mechanism.


5. Start at the Loss

Our loss is:

L=(target−y)2L=(target-y)^2

Therefore:

dLdy=−2(target−y)\frac{dL}{dy} = -2(target-y)

Here:

target=20target=20

and:

y=34y=34

So:

dLdy=−2(20−34)\frac{dL}{dy} = -2(20-34) dLdy=28\boxed{\frac{dL}{dy}=28}

The positive gradient tells us that increasing the prediction increases the loss at this point. The final neuron is:

y=w1a1+w2a2+w3a3+by=w_1a_1+w_2a_2+w_3a_3+b

We need gradients for:

w1,w2,w3,bw_1,w_2,w_3,b

Gradient of w1​

dydw1=a1\frac{dy}{dw_1}=a_1

Therefore:

dLdw1=dLdydydw1\frac{dL}{dw_1} = \frac{dL}{dy} \frac{dy}{dw_1} =28(9)=28(9) dLdw1=252\boxed{\frac{dL}{dw_1}=252}

Gradient of w2​

dLdw2=dLdya2\frac{dL}{dw_2} = \frac{dL}{dy}a_2 =28(9)=28(9) dLdw2=252\boxed{\frac{dL}{dw_2}=252}

Gradient of w3​

dLdw3=dLdya3\frac{dL}{dw_3} = \frac{dL}{dy}a_3 =28(6)=28(6) dLdw3=168\boxed{\frac{dL}{dw_3}=168}

Gradient of final bias​

Because:

dydb=1\frac{dy}{db}=1

we get:

dLdb=28\boxed{\frac{dL}{db}=28}

7. Continue Backward Through the Hidden Layer

Now we need gradients for:

a1,a2,a3a_1,a_2,a_3

For a1a_1:

dyda1=w1\frac{dy}{da_1}=w_1

Therefore:

dLda1=dLdydyda1\frac{dL}{da_1} = \frac{dL}{dy} \frac{dy}{da_1} =28(1)=28(1) dLda1=28\boxed{\frac{dL}{da_1}=28}

For a2a_2:

dLda2=28(2)\frac{dL}{da_2}=28(2) dLda2=56\boxed{\frac{dL}{da_2}=56}

For a3a_3:

dLda3=28(1)\frac{dL}{da_3}=28(1) dLda3=28\boxed{\frac{dL}{da_3}=28}

8. Backpropagate Through ReLU

ReLU is:

a=ReLU(z)=max⁡(0,z)a=ReLU(z)=\max(0,z)

Its derivative is:

dadz={1z>00z<0\frac{da}{dz} = \begin{cases} 1 & z>0\\ 0 & z<0 \end{cases}

Our values are:

z1=9,z2=9,z3=6z_1=9,\quad z_2=9,\quad z_3=6

All are positive.

Therefore:

da1dz1=1\frac{da_1}{dz_1}=1 da2dz2=1\frac{da_2}{dz_2}=1 da3dz3=1\frac{da_3}{dz_3}=1

So:

dLdz1=dLda1da1dz1=28(1)=28\frac{dL}{dz_1} = \frac{dL}{da_1} \frac{da_1}{dz_1} = 28(1) = 28 dLdz1=28\boxed{\frac{dL}{dz_1}=28}

Similarly:

dLdz2=56\boxed{\frac{dL}{dz_2}=56} dLdz3=28\boxed{\frac{dL}{dz_3}=28}

9. Backpropagate Into Hidden Weights

Neuron 1:

z1=w11x1+w12x2+b1z_1=w_{11}x_1+w_{12}x_2+b_1

For w11w_{11}:

dz1dw11=x1\frac{dz_1}{dw_{11}}=x_1

Therefore:

dLdw11=dLdz1dz1dw11\frac{dL}{dw_{11}} = \frac{dL}{dz_1} \frac{dz_1}{dw_{11}} =28(2)=28(2) dLdw11=56\boxed{\frac{dL}{dw_{11}}=56}

For w12w_{12}:

dLdw12=28(3)\frac{dL}{dw_{12}} = 28(3) dLdw12=84\boxed{\frac{dL}{dw_{12}}=84}

For b1b_1:

dLdb1=28\boxed{\frac{dL}{db_1}=28}

Neuron 2​

We have:

dLdz2=56\frac{dL}{dz_2}=56

Therefore:

dLdw21=56(2)=112\boxed{\frac{dL}{dw_{21}}=56(2)=112} dLdw22=56(3)=168\boxed{\frac{dL}{dw_{22}}=56(3)=168} dLdb2=56\boxed{\frac{dL}{db_2}=56}

Neuron 3​

We have:

dLdz3=28\frac{dL}{dz_3}=28

Therefore:

dLdw31=28(2)=56\boxed{\frac{dL}{dw_{31}}=28(2)=56} dLdw32=28(3)=84\boxed{\frac{dL}{dw_{32}}=28(3)=84} dLdb3=28\boxed{\frac{dL}{db_3}=28}

10. All Gradients

We have now backpropagated through the entire network.

Hidden layer​

dLdw11=56\frac{dL}{dw_{11}}=56 dLdw12=84\frac{dL}{dw_{12}}=84 dLdb1=28\frac{dL}{db_1}=28 dLdw21=112\frac{dL}{dw_{21}}=112 dLdw22=168\frac{dL}{dw_{22}}=168 dLdb2=56\frac{dL}{db_2}=56 dLdw31=56\frac{dL}{dw_{31}}=56 dLdw32=84\frac{dL}{dw_{32}}=84 dLdb3=28\frac{dL}{db_3}=28

Final neuron​

dLdw1=252\frac{dL}{dw_1}=252 dLdw2=252\frac{dL}{dw_2}=252 dLdw3=168\frac{dL}{dw_3}=168 dLdb=28\frac{dL}{db}=28

Every trainable parameter now has a gradient.


11. Weight and Bias Update

Backpropagation gave us the gradients.

Now gradient descent uses them to change the parameters.

The update rule is:

parameternew=parameterold−ηdLd(parameter)\boxed{ parameter_{new} = parameter_{old} - \eta \frac{dL}{d(parameter)} }

where η\eta is the learning rate.

Use:

learning_rate = 0.001

For example:

w1new=1−0.001(252)w_1^{new} = 1-0.001(252) w1new=0.748\boxed{w_1^{new}=0.748}

For w2w_2:

w2new=2−0.001(252)w_2^{new}=2-0.001(252) w2new=1.748\boxed{w_2^{new}=1.748}

For bb:

bnew=1−0.001(28)b^{new}=1-0.001(28) bnew=0.972\boxed{b^{new}=0.972}

The same update is performed for every parameter.