4.2 Why Activation Function?
Without an activation function, stacking layers isnt multiple functions, it can be simplified into a single linear function, we dont want that so we are sticking relu in the middle. ReLU introduces non-linearity.
A neuron can therefore calculate:
z=Wx+b
| ReLU stands for Rectified Linear Unit. | Sigmoid | Tanh |
|---|
| ReLU(x)=max(0,x) | σ(x)=1+e−x1 | tanh(x)=ex+e−xex−e−x |
ReLU(5)=5 ReLU(−3)=0 | σ(0)=0.5 σ(10)≈1 | tanh(−10)≈−1 tanh(10)≈1 |
return max(0, x) | return 1 / (1 + math.exp(-x)) | return math.tanh(x) |