Lesson 4 of 6
Lab 4 — Build a Neuron, Then a Network
Weights, bias, activation and gradient descent — worked by hand on a 2-input neuron, then scaled up.
Learn it
A neuron multiplies each input by a weight, adds them up, adds a bias, and squashes the result into a decision.
Training nudges the weights a little at a time in whichever direction reduces the error.
Stack neurons in layers and the network can learn curves and shapes a single neuron never could.
Key terms
- Weight
- How strongly an input influences the neuron's output.
- Bias
- A constant that shifts the decision boundary away from the origin.
- Activation function
- A non-linear squash (sigmoid, ReLU, tanh) applied to the weighted sum.
- Loss
- A number measuring how wrong the prediction is.
- Gradient descent
- Repeatedly stepping the weights downhill on the loss surface.
- Learning rate
- The size of each downhill step.
One update by hand
Neuron with w1 = 0.5, w2 = −0.4, b = 0.1, sigmoid activation, learning rate η = 0.5. Input x = (1, 1) with target y = 1.
- 11. Weighted sum: z = 0.5(1) + (−0.4)(1) + 0.1 = 0.2.
- 22. Activate: a = σ(0.2) = 1/(1+e^{-0.2}) ≈ 0.550.
- 33. Loss: Squared error = (0.550 − 1)² ≈ 0.203. We want it lower.
- 44. Gradient: ∂L/∂z = 2(a − y)·a(1 − a) = 2(−0.45)(0.550)(0.450) ≈ −0.223. For each input of 1, ∂L/∂w = −0.223.
- 55. Update: w1 ← 0.5 − 0.5(−0.223) ≈ 0.611; w2 ← −0.4 + 0.112 ≈ −0.288; b ← 0.1 + 0.112 ≈ 0.212. New z = 0.535, a ≈ 0.631 — closer to 1. That is learning.
A neuron that learns AND
pythonimport math, random
sig = lambda z: 1 / (1 + math.exp(-z))
data = [((0,0),0), ((0,1),0), ((1,0),0), ((1,1),1)]
w = [random.uniform(-1,1), random.uniform(-1,1)]
b, lr = 0.0, 0.5
for epoch in range(3000):
for (x1, x2), y in data:
a = sig(w[0]*x1 + w[1]*x2 + b)
d = 2*(a - y) * a * (1 - a) # dLoss/dz
w[0] -= lr * d * x1
w[1] -= lr * d * x2
b -= lr * d
for x, y in data:
print(x, y, round(sig(w[0]*x[0] + w[1]*x[1] + b), 3))Swap the targets to XOR (0,1,1,0) and this same code fails — a single neuron cannot separate XOR with one straight line.
Try it
What happens if you set lr = 50 in the code above?
pythonlr = 50
# ... same training loop ...Challenge
Do one full gradient-descent update by hand for a neuron with w1 = −0.2, w2 = 0.6, b = 0.0, sigmoid, η = 1.0, input (1, 0), target 0. Show z, a, the loss, the gradient and all three new parameters. Then explain in your own words why a single neuron cannot learn XOR and what the smallest network that can looks like.
Pick whichever way suits you — every mode earns the same bonus XP.
Write at least 40 more characters to submit.
Mark your own work
Guided walkthrough — 0/3 clues revealed
- Clue 1 locked — reveal it only if you get stuck.
- Clue 2 locked — reveal it only if you get stuck.
- Clue 3 locked — reveal it only if you get stuck.
Each clue costs 8 XP (never below 38 XP). You'd earn 75 XP right now.
Extension: Rewrite the code with one hidden layer of two neurons and get XOR working. Log the loss every 500 epochs.