
Deep learning sounds intimidating until you see that the core step is arithmetic you could do in a spreadsheet. We will build one artificial neuron in Excel, train it to separate two groups, and watch the numbers change. No Python needed.
In this article
The problem
A shop wants to predict whether a customer will return, using two numbers: visits last month and average bill (in ₹ thousands). Put ten rows of history on a sheet:
| A: Visits | B: Avg bill | C: Returned (1/0) |
|---|---|---|
| 1 | 0.5 | 0 |
| 4 | 1.2 | 1 |
| 2 | 0.8 | 0 |
| 6 | 2.0 | 1 |
| 5 | 0.9 | 1 |
(Use ten or twenty rows in practice; five fit here.)
Step 1: the neuron
A neuron multiplies each input by a weight, adds a bias, and squeezes the result between 0 and 1 with the sigmoid function. Put starting values in cells: w1 in F1 = 0.1, w2 in F2 = 0.1, bias in F3 = 0.
D2: =A2*$F$1 + B2*$F$2 + $F$3 ' raw score
E2: =1/(1+EXP(-D2)) ' prediction between 0 and 1
With tiny starting weights every prediction is around 0.5 — the neuron has no idea yet.
Step 2: measure the error
G2: =E2-C2 ' how wrong, with direction
H2: =G2^2 ' squared error
H20: =AVERAGE(H2:H11) ' the loss
Training means making that loss smaller.
Step 3: which way to nudge each weight
Calculus tells us how much the loss changes when a weight changes (the gradient). For a sigmoid neuron with squared error it comes out as simple products:
I2: =G2*E2*(1-E2) ' error signal for this row
J2: =I2*A2 ' gradient for w1
K2: =I2*B2 ' gradient for w2
Average J, K and I over all rows to get the gradient for w1, w2 and the bias.
Step 4: update and repeat
New weight = old weight − learning rate × gradient. With a learning rate of 0.5:
new w1: =F1 - 0.5*AVERAGE(J2:J11)
Copy the new values back into F1:F3 (paste values) and watch the loss drop. Doing it 20 times by hand gets boring — record a tiny macro that copies and pastes, or lay out each round on its own row so the sheet itself shows 100 rounds of training.
What you will see
- The loss falls quickly at first, then flattens.
- Weights grow in the direction that matters: visits usually gets the bigger weight.
- Too big a learning rate (try 5) makes the loss jump around instead of falling. That is the same problem real models have.
From one neuron to deep learning
A deep network is thousands or millions of these neurons in layers, each layer feeding the next. The maths is identical; the gradient is passed backwards through the layers (backpropagation) and computed by a GPU instead of a spreadsheet. Libraries like TensorFlow and PyTorch do the calculus automatically.
For the bigger picture, see machine learning vs deep learning.
Stuck on a step? Ask a question and the AI answers using this article.