What are we trying to predict? Features and targets
Every machine learning model starts with two decisions, what to predict and what to tell the model. The first is the target, the second the features, and real SPFL results show why the features matter as much as the model.
Beginner Part 1 of Machine Learning Through Football
Contents
The football question
Celtic play on Saturday. Before a machine learning model can tell us anything about the result, what do we have to decide?
Two things, before any maths:
- What are we trying to predict?
- What information are we going to give the model to help it?
The thing we're trying to predict is the target. The information we give the model is the features. Everything else in machine learning is built on top of those two choices.
The concept
A model learns from examples. Each example is one match: its features, then the target, the answer we already know because the match has been played.
| Value | |
|---|---|
| Home | Yes |
| Recent xG | 1.84 |
| Recent xGA | 0.91 |
| Opponent strength | 73 |
| Rest days | 6 |
| Target: result | Win |
Made-up example numbers. Repeat that across hundreds or thousands of matches, and the model starts looking for relationships between the features and the target: sides with high recent xG and a weak opponent win more often, say.
Once it has learned those relationships, you give it a match that hasn't been played. It gets the features, and it has to supply the target itself. Written down:
$$\hat{y} = f(\mathbf{x})$$
In plain football
- x is the match's features: at home, recent xG 1.84, and so on.
- f is what the model learned from all the old matches. How it learns is gradient descent.
- ŷ ("y-hat") is its prediction: win, or a 62% chance of winning. The hat means "estimated". The real result, y, comes at full time.
A football example
Before a model can use that row, every feature has to be a number. "Home = Yes" becomes 1, "No" becomes 0. The row becomes a vector, just like a midfielder's match in five numbers:
$$\mathbf{x} = (1,\ 1.84,\ 0.91,\ 73,\ 6)$$
In plain football
- 1: we're at home.
- 1.84 and 0.91: expected goals for and against over recent matches.
- 73: the opponent's strength rating, on whatever scale we choose (here, out of 100).
- 6: days since our last match.
Stack a thousand of those rows and you have a matrix, one match per row and one feature per column, as in a team as a matrix. The targets, one per row, sit alongside it in a column of their own.
The target decides the kind of problem
The same match can be asked about in different ways, and each gives a different target:
| Question | Target | Problem |
|---|---|---|
| Did Celtic win? | 1 or 0 | Classification |
| Home, draw or away? | H, D or A | Classification |
| How many goals? | 0, 1, 2, … | Regression |
Classification predicts a category; regression predicts a number. Choosing between them isn't a technical detail. "Did we win?" lumps a draw in with a defeat, which a manager fighting relegation would never do. Pick the target that matches the decision you actually care about. Three possible results at once is the territory of the Multinomial distribution.
Do the features carry information?
A feature earns its place only if it tells the model something about the target. Real results show what that looks like.
Take every Scottish Premiership match from 2000/01 to 2025/26, and one feature: each side's points per game over its previous five league matches, the home side's figure minus the away side's. Skip the first five matches of each season, when there's no run of form to use. That leaves 5,086 matches. Across all of them, the home side won 44.5%, drew 23.2% and lost 32.3%. (Whether form matters as much as fans think is a myth of its own.)
Now split them into five bands by the form gap: the away side far better (by a point a game or more), better, about level (within 0.4 of each other), then the home side better and far better. The bands hold between 650 and 1,871 matches each.
One number, known before kick-off, moves the chance of a home win from 25% to 65%. That's a feature carrying real information. It isn't a model yet: nothing here combines form with anything else. But it's exactly the kind of relationship a model goes looking for.
Choosing the right features
Choosing the right features can matter just as much as choosing the model. The cleverest algorithm there is can't make up for poor information. If all I tell it is that Celtic had 65% possession, I shouldn't be surprised when it struggles to understand the match. Possession alone says nothing about chances, finishing, or whether the other side were happy to sit in and counter.
The feature that cheats
There's a second trap, and it catches experienced analysts too. Possession, shots and the half-time score are all measured during the match. A model predicting the result before kick-off can't have them.
Try it with the same Premiership seasons. Here's the half-time score as a feature, across 5,879 matches:
It looks like a brilliant feature: a side leading at the break wins nearly four times in five. It's also useless on Saturday morning, because nobody knows the half-time score yet. Train a model with it and the model looks superb on old matches, then falls apart on the next one.
This is called data leakage: information from after the moment of prediction sneaking into the features. The test is simple. For every feature, ask: would I know this at the moment I have to make the prediction? Recent form, yes. Rest days, yes. The half-time score, no.
Why it matters
- The target defines the problem. Win or not, three results, or goals scored: each needs a different kind of model and answers a different question.
- Features are where football knowledge goes in. A coach's sense of what matters, form, fatigue, the quality of the opposition, becomes the columns of the data.
- Good features beat clever models. One well-chosen number moved a home win from 25% to 65%. No algorithm can find a pattern the features don't contain.
- Leakage makes bad models look brilliant. A model built on information from the future looks unbeatable until you use it.
Which leaves the obvious next question. We can give the model hundreds of old matches, but if we use all of them to teach it, how do we know it has actually learned anything, rather than just memorised the results?
Limitations
- Features only capture what you measure. Injuries, a new manager, a derby atmosphere: if it isn't a column, the model can't see it.
- Features overlap. Recent xG and recent points both measure how well a side is playing. Adding the second tells the model less than you'd think.
- Relationships change. What mattered in 2005 may matter less now; a model trained on old seasons carries old football with it.
- Some things no feature explains. None of the features above would have predicted Celtic getting beaten by our biggest rivals. I definitely didn't have that one in the model.
Try it yourself
Pick your team's next match. Write down five features you'd give a model to predict it, and for each one, ask the leakage question: would I know this before kick-off? Then choose your target: win or not, the three results, or goals scored. Which would be most useful to you as a fan, and which to the manager?
Reproduce the analysis
The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. Then:
import csv, glob
from collections import Counter, defaultdict
from datetime import datetime
POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
form, halftime = Counter(), Counter()
for path in glob.glob("SC0_*.csv"):
with open(path, encoding="latin-1") as f:
games = [g for g in csv.DictReader(f) if g.get("FTR") in POINTS]
games.sort(key=lambda g: datetime.strptime(g["Date"], "%d/%m/%Y" if len(g["Date"]) == 10 else "%d/%m/%y"))
pts = defaultdict(list)
for g in games:
h, a, r = g["HomeTeam"], g["AwayTeam"], g["FTR"]
if len(pts[h]) >= 5 and len(pts[a]) >= 5: # form known before kick-off
gap = (sum(pts[h][-5:]) - sum(pts[a][-5:])) / 5
band = (gap <= -1) + (gap < -0.4) + (gap <= 0.4) + (gap < 1) # 4 = home far better ... 0 = away far better
form[4 - band, r] += 1
if g.get("HTR") in POINTS:
halftime[g["HTR"], r] += 1
pts[h].append(POINTS[r][0])
pts[a].append(POINTS[r][1])
for name, table, rows in [("form band (0 = away far better)", form, range(5)), ("half time", halftime, "HDA")]:
print(name)
for k in rows:
n = sum(table[k, r] for r in "HDA")
print(f" {k}: {n} matches", *(f"{r} {table[k, r] / n:.1%}" for r in "HDA"))
Further reading
- Supervised learning, Google for Developers. Features, labels (their word for targets) and how a model learns from them, with interactive examples.
- Feature (machine learning), Wikipedia. What counts as a feature, how categories are turned into numbers, and feature selection.
- Leakage (machine learning), Wikipedia. The ways future information sneaks into training data, and how to spot it.