linear-regression

One Algorithm. Every Machine Learning Model Traces Back to This.

From what supervised learning actually is, to cost functions and gradient descent — everything you need to understand and build linear regression models.

5 min read

A Machine That Learns From Examples

Imagine you're a real estate agent. You've seen hundreds of houses. Over time, you develop an instinct — bigger house, higher price. You didn't memorise a formula. You learned from experience.

Machine learning does exactly this — but with math instead of instinct.

This is supervised learning: you give the machine labelled examples (inputs + correct answers), and it learns the pattern connecting them.

There are two types:

  • Classification — predict a category. Is this email spam or not?
  • Regression — predict a number. What will this house cost?

Linear regression is where regression begins. Simple, powerful, and the foundation of everything that follows.

What Linear Regression Actually Is

At its core, linear regression finds the best straight line through your data.

You have a dataset: house sizes (input) and their prices (output). Plot them. You'll see a rough upward trend. Linear regression finds the line that fits that trend best.

The line equation:

f(x) = wx + b

  • x — your input (house size)
  • w — weight (slope of the line)
  • b — bias (where the line crosses the y-axis)
  • f(x) — prediction

Your model's job: find the right values of w and b so the line fits your data as closely as possible.

The Cost Function: Measuring How Wrong You Are

Any line can fit through data. The question is — how good is it?

The cost function measures this. It quantifies the error between your predictions and the actual values.

For linear regression, we use Mean Squared Error (MSE):

J(w,b) = (1/2m) × Σ(f(x⁽ⁱ⁾) - y⁽ⁱ⁾)²

Breaking this down:

  • m — number of training examples
  • f(x⁽ⁱ⁾) — your prediction for example i
  • y⁽ⁱ⁾) — the actual value for example i
  • (f(x⁽ⁱ⁾) - y⁽ⁱ⁾)² — squared error (squaring removes negatives and penalises large errors more)
  • 1/2m — average, divided by 2 for cleaner math later

The goal: minimise J(w,b). The smaller the cost, the better your line fits.

Think of it as a bowl-shaped surface. Every combination of w and b maps to a point on that bowl. The bottom of the bowl is where your error is lowest — that's where you want to be.

Gradient Descent: Walking Downhill

You know what you want (minimum cost). The question is how to get there.

Gradient descent is the algorithm that walks you downhill — toward the minimum — step by step.

The update rule:

w = w - α × (∂J/∂w)
b = b - α × (∂J/∂b)

  • α (alpha) — learning rate. How big each step is.
  • ∂J/∂w — partial derivative. Which direction is downhill for w?
  • ∂J/∂b — partial derivative. Which direction is downhill for b?

The derivatives for linear regression work out to:

∂J/∂w = (1/m) × Σ(f(x⁽ⁱ⁾) - y⁽ⁱ⁾) × x⁽ⁱ⁾
∂J/∂b = (1/m) × Σ(f(x⁽ⁱ⁾) - y⁽ⁱ⁾)

You repeat these updates until the cost stops decreasing — that's convergence.

The Learning Rate: The Most Important Hyperparameter

α controls step size. Get it wrong and everything breaks.

  • Too large → you overshoot the minimum, bounce around, never converge
  • Too small → you get there eventually, but painfully slowly
  • Just right → smooth descent to the minimum

Typical starting values: 0.01, 0.001, 0.1. You experiment.

Building It in Code

python

import numpy as np # Training data X = np.array([1, 2, 3, 4, 5]) # house size y = np.array([2, 4, 5, 4, 5]) # price # Initialise w, b = 0.0, 0.0 alpha = 0.01 epochs = 1000 m = len(X) # Gradient descent for _ in range(epochs): y_pred = w * X + b error = y_pred - y dw = (1/m) * np.sum(error * X) db = (1/m) * np.sum(error) w -= alpha * dw b -= alpha * db print(f"w: {w:.2f}, b: {b:.2f}") print(f"Prediction for x=6: {w*6 + b:.2f}")

That's it. No library magic. The entire algorithm in 10 lines.

Multiple Features: The Real World

Real problems rarely have one input. A house price depends on size, location, number of rooms, age.

Multiple linear regression handles this:

f(x) = w₁x₁ + w₂x₂ + w₃x₃ + b

In vector form:

f(x) = w·x + b

The cost function and gradient descent work identically — just across more dimensions.

Where Linear Regression Is Actually Used

Finance — predicting stock returns, credit risk scoring, loan default probability.

Healthcare — predicting patient recovery time, drug dosage response, hospital readmission rates.

Real estate — property valuation models. Zillow's Zestimate started here.

Retail — demand forecasting, pricing models, inventory planning.

Engineering — predicting material stress, failure rates, system performance under load.

What Linear Regression Cannot Do

It assumes the relationship between input and output is a straight line. The real world often isn't.

When your data has curves, use polynomial regression. When you have complex non-linear patterns, you move toward decision trees, neural networks, or ensemble methods.

But here's what matters: every advanced model builds on this foundation. Understanding gradient descent here means you understand how neural networks train. The cost function here is the same concept used in every learning algorithm that follows.

The Line Is Just the Beginning

Linear regression seems simple. One line through data. But inside it lives everything:

  • The idea that a model has parameters to learn
  • The idea that a cost function measures error
  • The idea that gradient descent optimises anything differentiable

Master this and you haven't just learned linear regression. You've learned how machines learn.

Everything else is a more complex version of this exact story.