10 min read
Demystifying Linear Regression
Finding the line that fits: A gentle introduction to linear regression
Linear regression often sounds like something that only a stats‑savvy person can handle, but the idea behind it is actually very old and straightforward. Imagine you have a bunch of measurements. You notice they seem to be drifting in a particular direction, and you want to guess what the next one will look like. The classic, simplest way to answer that is to draw a straight line that best fits the cloud of points.
The short version
Linear regression finds the straight-line pattern in a set of data and uses it to guess at values you have not measured yet.
Think of
the stove-top dial
After years of cooking, you’ve got a feel for how far you need to turn the dial to get just the right amount of heat, so you set it and walk away. That intuitive link between how much you crank the knob (input) and how much flame you get (output) is essentially a straight line: a relationship you’ve “fitted” by hand through countless burnt meals.
What linear regression actually does
Suppose you have the heights and weights of a group of people. Plotted out, they look more like a cloud than a line: taller people tend to weigh more, but plenty of them do not. Linear regression draws the line that best summarises the lean of that cloud. The same move works on stock prices or on how long a delivery will take, which is most of why it turns up everywhere.
A line through human data is (almost) never going to be right about any particular person. It is only ever a reasonable fit, and reasonable is usually enough to act on. Which leaves the question of what reasonable means, and who gets to measure it.
- Model
- A formula that takes an input, such as height, and returns a predicted output, such as weight.
- Linear relationship
- A connection where a fixed change in one variable produces a proportional fixed change in another, drawing a straight line. The line does not need to pass through zero.
The short version
We square our misses so they cannot cancel each other out, which gives us a referee that punishes big mistakes far more harshly than small ones.
Think of
the tailor's suit
A tailor fitting a suit. The residuals are the gaps and extra bunches of fabric between cloth and body. A suit cut to every single postural quirk on Tuesday will not fit properly on Wednesday. Zero residuals usually indicate overfitting rather than a good model.
Which line wins, and who decides
Infinitely many lines pass through any cloud of points, and most of them are terrible. To pick one you need a referee: a rule that takes a candidate line and hands back a single number saying how badly it misses. That rule is the loss function, and it works by measuring the gap between each prediction and what actually happened.
The standard referee is mean squared error (MSE). For every point, you take the vertical distance between the actual value and the line (the residual), square it, and compute the average across all data points. Fitting the line simply means searching for the slope and intercept that drive this average score as low as possible.
- Residual
- The difference between an actual observed value and the value the model predicted.
- Mean squared error (MSE)
- The average of all those differences after squaring them, which penalises the larger misses more heavily.
- Loss function
- A formula that measures how bad a model is. Fitting the model means driving this number as low as it will go.
Figure 1: eighteen synthetic people
0.35kg / cm
72.0kg
Sum of squares
1,189
Because errors are squared, a single distant outlier can exert heavy leverage over the slope, pulling the line away from the main cluster to reduce that single large penalty.
The short version
Several inputs blended into one formula, each carrying its own weight. This is how you get from toy examples to house prices.
Think of
the soup
Features are the ingredients in a soup. The salt and broth alter the taste significantly, while a pinch of parsley barely registers. The recipe has the same structure either way, but individual quantities dictate the outcome.
When one variable is not enough
Most real-world outcomes depend on multiple variables. A house price depends on floor area, bedroom count, and location. Multiple regression takes these factors together and calculates a single prediction, weighting each input by its estimated impact.
Model accuracy depends heavily on choosing appropriate features. While height and weight are both informative, Body Mass Index divides weight by height squared to capture body density.
Using height squared demonstrates an important property of linear regression: it is linear in its parameters, not necessarily in its raw inputs. A model can include squared terms, logarithmic scales, or interaction terms (such as multiplying two features together), allowing it to fit non-linear curves while preserving standard linear optimization.
- Predictor variable
- An individual input used to make a prediction, such as size, location or age.
- Multiple regression
- A model that handles more than one predictor variable at the same time. While often called multivariate regression in machine learning writing, statisticians reserve that term for models with multiple output variables.
- Featurization
- Turning raw, real-world data into the numbers a model can actually take as input.
The short version
Hold some of the data back so you can find out whether the model learned the pattern or just memorised the answers.
Think of
the old exam paper
A student who memorises the answers to an old exam paper will score full marks on that practice test and fail the real exam. A model needs generalization: enough grasp of the underlying pattern to make accurate predictions on unseen data.
The test the model has not seen
The failure that ruins more models than anything else is one that has quietly memorised its training data. It scores beautifully on the numbers it was fitted to and falls apart on anything new. The defence is almost embarrassingly simple: before fitting anything, put some of the data in a drawer. Train on the rest. When the model is finished, open the drawer and see how it does on rows it has never been shown.
Consider the election-formula problem. Someone claims a formula built from 31 economic variables that predicted every presidential election from 1928 to 2020. Count the elections in that range and you get twenty-four of them. With 31 knobs to turn and only 24 results to match, a perfect fit is not evidence of anything, it is arithmetic. You could hit all twenty-four using 31 variables drawn from the price of tea. The formula fits the history books flawlessly and still tells you nothing about the next election.
- Held-out data
- A portion of the data, often 20%, kept back from the model during training and used for the final test.
- Generalization error
- The prediction error on unseen data. The difference between generalization error and training error is known as the generalization gap.
- Overfitting
- When a model fits its training data closely and produces nonsense on anything new.
The short version
A model can score text by assigning a number to each word and adding them up: positive words raise the rating, negative words pull it down.
Think of
the bucket of magnetic poetry
A bag of words treats a review like a bucket: words like enjoyable and remarkable add points to the score, while terrible and disappointing subtract them.
Teaching a machine to read a review
Linear regression is not limited to physical measurements like height or square footage. It works just as well on words. Imagine predicting a movie review’s star rating (from 1 to 5) based entirely on the words it contains.
Every distinct word becomes its own feature with its own weight. The baseline intercept might start at 3 stars. Words like “masterpiece” or “delightful” carry positive weights that nudge the prediction upward, while words like “boring” or “unwatchable” carry negative weights that drag it down. To predict a score, the model simply counts the words in the review, multiplies each by its weight, and sums them up.
Most words in a dictionary—like the, table, or pencil—carry no emotional weight at all. When you inspect a trained model, the vast majority of word weights bunch tightly around zero. The model automatically learns to ignore the filler and concentrate its attention on the handful of words that actually signal sentiment.
Single words alone can miss context: “not bad” contains the word “bad”, but means something positive. To catch these phrases, models group adjacent words into pairs (bigrams) like “not good” or “highly recommended”, treating each pair as its own independent feature.
- Bag of words
- Representing text by counting how many times each word appears, ignoring sentence structure and grammar.
- Feature weight
- The number assigned to each word indicating how strongly (and in which direction) it influences the final prediction.
- N-gram
- A sequence of adjacent words treated as a single token—allowing the model to tell the difference between "good" and "not good". Example: "pretty good" is a bigram.
The short version
Under the hood, libraries use optimization algorithms to roll down the loss landscape and find the best-fitting line in seconds.
Think of
the foggy valley
Gradient descent is like walking down a foggy mountain. You cannot see the bottom of the valley, but you can feel which way the ground slopes under your feet. Step in the steepest downhill direction, and you eventually reach the lowest point of loss.
The tools that do the arithmetic
In simple examples with just one or two inputs, mathematicians can calculate the best slope and intercept in a single direct step using closed-form formulas. But as datasets grow to millions of rows and hundreds of features, computing the answer in one giant leap becomes impractical.
Instead, modern machine learning libraries find the winning line by rolling downhill. Starting with random weights, an algorithm called gradient descent repeatedly checks the loss referee, calculates which direction would reduce the error, and nudges the weights slightly downhill. Step by step, it navigates the fog until the error can go no lower.
Figure 2: descending the loss landscape
Loss at current slope
1,015
Once the model is fitted, you need a way to gauge its quality. While Mean Squared Error tells you the raw size of the misses in data units, R² (the coefficient of determination) normalizes that score between 0 and 1. An R² of 0 means the line is no better than guessing the average outcome every time; an R² of 0.85 means the model has captured 85% of the variation in the data.
- Gradient descent
- An iterative method that repeatedly adjusts model weights downhill to find the lowest possible loss.
- R-squared (R²)
- A score measuring how well the line fits the data compared to simply guessing the overall average.