One Algorithm. Every Machine Learning Model Traces Back to This.
From what supervised learning actually is, to cost functions and gradient descent — everything you need to understand and build linear regression models.
A Machine That Learns From Examples
Imagine you're a real estate agent. You've seen hundreds of houses. Over time, you develop an instinct — bigger house, higher price. You didn't memorise a formula. You learned from experience.
Machine learning does exactly this — but with math instead of instinct.
This is supervised learning: you give the machine labelled examples (inputs + correct answers), and it learns the pattern connecting them.
There are two types:
- Classification — predict a category. Is this email spam or not?
- Regression — predict a number. What will this house cost?
Linear regression is where regression begins. Simple, powerful, and the foundation of everything that follows.
What Linear Regression Actually Is
At its core, linear regression finds the best straight line through your data.
You have a dataset: house sizes (input) and their prices (output). Plot them. You'll see a rough upward trend. Linear regression finds the line that fits that trend best.
The line equation:
f(x) = wx + b
x— your input (house size)w— weight (slope of the line)b— bias (where the line crosses the y-axis)f(x)— prediction
Your model's job: find the right values of w and b so the line fits your data as closely as possible.
The Cost Function: Measuring How Wrong You Are
Any line can fit through data. The question is — how good is it?
The cost function measures this. It quantifies the error between your predictions and the actual values.
For linear regression, we use Mean Squared Error (MSE):
J(w,b) = (1/2m) × Σ(f(x⁽ⁱ⁾) - y⁽ⁱ⁾)²
Breaking this down:
m— number of training examplesf(x⁽ⁱ⁾)— your prediction for example iy⁽ⁱ⁾)— the actual value for example i(f(x⁽ⁱ⁾) - y⁽ⁱ⁾)²— squared error (squaring removes negatives and penalises large errors more)1/2m— average, divided by 2 for cleaner math later
The goal: minimise J(w,b). The smaller the cost, the better your line fits.
Think of it as a bowl-shaped surface. Every combination of w and b maps to a point on that bowl. The bottom of the bowl is where your error is lowest — that's where you want to be.
Gradient Descent: Walking Downhill
You know what you want (minimum cost). The question is how to get there.
Gradient descent is the algorithm that walks you downhill — toward the minimum — step by step.
The update rule:
w = w - α × (∂J/∂w)
b = b - α × (∂J/∂b)
α(alpha) — learning rate. How big each step is.∂J/∂w— partial derivative. Which direction is downhill forw?∂J/∂b— partial derivative. Which direction is downhill forb?
The derivatives for linear regression work out to:
∂J/∂w = (1/m) × Σ(f(x⁽ⁱ⁾) - y⁽ⁱ⁾) × x⁽ⁱ⁾
∂J/∂b = (1/m) × Σ(f(x⁽ⁱ⁾) - y⁽ⁱ⁾)
You repeat these updates until the cost stops decreasing — that's convergence.
The Learning Rate: The Most Important Hyperparameter
α controls step size. Get it wrong and everything breaks.
- Too large → you overshoot the minimum, bounce around, never converge
- Too small → you get there eventually, but painfully slowly
- Just right → smooth descent to the minimum
Typical starting values: 0.01, 0.001, 0.1. You experiment.
Building It in Code
python
import numpy as np # Training data X = np.array([1, 2, 3, 4, 5]) # house size y = np.array([2, 4, 5, 4, 5]) # price # Initialise w, b = 0.0, 0.0 alpha = 0.01 epochs = 1000 m = len(X) # Gradient descent for _ in range(epochs): y_pred = w * X + b error = y_pred - y dw = (1/m) * np.sum(error * X) db = (1/m) * np.sum(error) w -= alpha * dw b -= alpha * db print(f"w: {w:.2f}, b: {b:.2f}") print(f"Prediction for x=6: {w*6 + b:.2f}")
That's it. No library magic. The entire algorithm in 10 lines.
Multiple Features: The Real World
Real problems rarely have one input. A house price depends on size, location, number of rooms, age.
Multiple linear regression handles this:
f(x) = w₁x₁ + w₂x₂ + w₃x₃ + b
In vector form:
f(x) = w·x + b
The cost function and gradient descent work identically — just across more dimensions.
Where Linear Regression Is Actually Used
Finance — predicting stock returns, credit risk scoring, loan default probability.
Healthcare — predicting patient recovery time, drug dosage response, hospital readmission rates.
Real estate — property valuation models. Zillow's Zestimate started here.
Retail — demand forecasting, pricing models, inventory planning.
Engineering — predicting material stress, failure rates, system performance under load.
What Linear Regression Cannot Do
It assumes the relationship between input and output is a straight line. The real world often isn't.
When your data has curves, use polynomial regression. When you have complex non-linear patterns, you move toward decision trees, neural networks, or ensemble methods.
But here's what matters: every advanced model builds on this foundation. Understanding gradient descent here means you understand how neural networks train. The cost function here is the same concept used in every learning algorithm that follows.
The Line Is Just the Beginning
Linear regression seems simple. One line through data. But inside it lives everything:
- The idea that a model has parameters to learn
- The idea that a cost function measures error
- The idea that gradient descent optimises anything differentiable
Master this and you haven't just learned linear regression. You've learned how machines learn.
Everything else is a more complex version of this exact story.