Eight stats, one story. PCA and dimensionality reduction
Football stats overlap. Principal component analysis finds the few directions that carry most of the information. On 26 seasons of Scottish Premiership data, one component built from eight stats explains 56% of the variation and tracks points per game better than any single stat.
Advanced Part 9 of Linear Algebra Through Football
Contents
The football question
A modern data provider can give you 40 or more numbers for every player and every team. Shots, shots on target, expected goals, corners, touches in the box, passes into the final third… Read them all and you have a 40-page scouting report. Do we really need all of them, or are many of them saying the same thing?
Mostly, they're saying the same thing. A team that takes a lot of shots also tends to hit the target a lot and win a lot of corners. Principal component analysis (PCA) finds that shared story and writes it as one number.
The concept
Think of each team's season as a vector of stats. With eight stats, every team is a point in eight dimensions. PCA looks for new directions through that cloud of points:
- The first principal component is the direction along which teams differ the most.
- The second is the direction of greatest difference that's left, at right angles to the first.
- And so on, one for each original stat.
Each direction is a set of weights, one per stat. A team's score on a component is the weighted sum of its (standardised) stats: the same dot product as a player rating. The difference is that nobody chooses the weights. The data does.
If the first two or three components capture most of the differences between teams, you can describe teams with those few numbers instead of all eight. That's dimensionality reduction.
A real example: 26 Premiership seasons
For every Scottish Premiership team from 2000/01 to 2025/26 (310 team-seasons, from football-data.co.uk), take eight per-match averages:
- shots, shots on target and corners, both for and against
- fouls committed and yellow cards
These overlap. Across team-seasons, shots and shots on target have a correlation of 0.72, and shots and corners 0.73. Knowing one tells you a lot about the other two.
Standardise each stat (as in the distance part), so 10 shots and 2 yellow cards a game count on the same scale, then run PCA:
| Component | Share of variation | Running total |
|---|---|---|
| 1 | 55.8% | 55.8% |
| 2 | 14.0% | 69.7% |
| 3 | 11.9% | 81.6% |
| 4 to 8 | 18.4% | 100% |
In plain football
- The first component alone carries more than half of all the differences between these 310 team-seasons.
- The first two carry 70%. Eight numbers per team can be squeezed into two while keeping most of the information.
- The last five together carry less than a fifth: mostly detail and noise.
Component 1: dominance
Here are the weights PCA found for the first component:
| Stat | Weight |
|---|---|
| Shots | +0.39 |
| Shots on target | +0.40 |
| Corners | +0.42 |
| Shots against | −0.37 |
| On target against | −0.37 |
| Corners against | −0.40 |
| Fouls | −0.14 |
| Yellow cards | −0.25 |
In plain football
- Plus weights on everything you do in the opponent's half, minus weights of about the same size on everything they do in yours.
- So a high score means you create more than you allow: territorial dominance. A low score means you're pinned back.
- Fouls and yellows lean slightly negative: dominant sides foul a little less, perhaps because they have the ball more.
Nobody told PCA about results, but this component tracks them closely. Its correlation with points per game is 0.87. The best single stat, shots on target, manages 0.80. Combining eight overlapping stats into one well-chosen direction beats any one of them. The four highest scores in the data all belong to Celtic, led by 2021/22. The lowest is Hamilton in 2019/20.
Component 2: fouls
The second component is almost all fouls committed (+0.85) and yellow cards (+0.41), with little weight on anything else. It separates teams that foul a lot from teams that don't, whatever their level. Its correlation with points per game is −0.05: essentially none. How physical a team is and how good it is are two separate stories, and PCA has pulled them apart.
Component 3 and beyond
The third component explains another 12%, but it's harder to name: it mixes yellow cards with more shots at both ends and fewer shots on target conceded. That's normal. The first components are usually clear, later ones less so, and there's no rule that says every component must mean something in football terms. A common rule of thumb keeps components whose eigenvalue (see below) is above 1: here that's the first two.
Show the mathsEigenvectors, eigenvalues and the share of variation. Optional.
Standardise the data matrix \(X\) (\(n\) team-seasons by \(p\) stats) column by column to get \(Z\). Its correlation matrix is
$$C = \frac{1}{n} Z^{\mathsf T} Z$$
a \(p \times p\) matrix: entry \((j, k)\) is the correlation between stats \(j\) and \(k\).
The principal components are the eigenvectors of \(C\): the directions \(\mathbf{v}\) that \(C\) only stretches, without turning,
$$C\,\mathbf{v} = \lambda\,\mathbf{v}$$
The stretch factor \(\lambda\), the eigenvalue, is the variance of the team scores along that direction. With standardised data the eigenvalues add up to \(p\), the number of stats, so component \(i\)'s share of the variation is
$$\frac{\lambda_i}{\lambda_1 + \dots + \lambda_p}$$
Here the eigenvalues are 4.46, 1.12, 0.95, 0.62, 0.31, 0.26, 0.22 and 0.06, which add up to 8, and 4.46 ÷ 8 = 55.8%.
Each team's scores on the components are \(Z V\), where the columns of \(V\) are the eigenvectors: one dot product per team per component. An eigenvector's overall sign is arbitrary. We chose signs so that component 1 rises with shots and component 2 rises with fouls.
Why it matters
- Scouting: 40 player metrics become five or six components, each a style or skill (ball progression, box presence, defensive work), and players can be compared on those.
- Similarity: the distance and cosine measures from earlier in the series work better on components than on dozens of overlapping stats, where one quality counted five times dominates the answer.
- Pictures: two components make a map, like the one above, where a data set with dozens of columns can be seen at a glance.
- Modelling: fewer, uncorrelated inputs make prediction models more stable.
PCA doesn't throw information away at random. It keeps the directions where teams differ most and drops the ones where they barely differ.
Limitations
- The names are ours. PCA produces weights, not labels. "Dominance" and "fouls" are our reading of the weights, and a later component may have no clean football meaning.
- Straight lines only. PCA finds linear combinations of stats. Relationships that curve, or only matter above a threshold, can be missed.
- Most variation isn't most important. A component can be large and irrelevant to winning, like fouls here, while a small one matters. Check components against the outcome you care about.
- Standardising is a choice. Without it, the stats with the biggest numbers, such as fouls and shots, would dominate the components, as in the scale trap.
- Team-seasons, not players. These are season averages for whole teams. Player-level PCA uses the same method on different data, and needs enough minutes per player to be stable.
- Some seasons are short. Team-seasons with fewer than 30 matches carrying these stats are left out, and 2019/20 was cut short by the pandemic.
Reproduce the analysis
The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. Then, with numpy installed:
import csv, glob
from collections import defaultdict
import numpy as np
STATS = ["shots", "shots against", "on target", "on target against",
"corners", "corners against", "fouls", "yellows"]
NEED = ("FTHG", "FTAG", "HS", "AS", "HST", "AST", "HC", "AC", "HF", "AF", "HY", "AY")
totals = defaultdict(lambda: np.zeros(len(STATS) + 2)) # stats, points, matches
for path in glob.glob("SC0_*.csv"):
with open(path, encoding="latin-1") as f:
for r in csv.DictReader(f):
if not all(r.get(c) for c in NEED):
continue
v = {c: int(r[c]) for c in NEED}
for team, us, them in ((r["HomeTeam"], "H", "A"), (r["AwayTeam"], "A", "H")):
gf, ga = v[f"FT{us}G"], v[f"FT{them}G"]
pts = 3 if gf > ga else 1 if gf == ga else 0
totals[path, team] += [v[us + "S"], v[them + "S"], v[us + "ST"], v[them + "ST"],
v[us + "C"], v[them + "C"], v[us + "F"], v[us + "Y"], pts, 1]
rows = [t for t in totals.values() if t[-1] >= 30]
X = np.array([t[:-2] / t[-1] for t in rows])
ppg = np.array([t[-2] / t[-1] for t in rows])
Z = (X - X.mean(0)) / X.std(0)
vals, vecs = np.linalg.eigh(np.corrcoef(Z, rowvar=False))
vals, vecs = vals[::-1], vecs[:, ::-1] * np.sign(vecs[0, ::-1])
print(len(rows), "team-seasons")
print("share of variation:", np.round(vals / vals.sum(), 3))
print("component 1 weights:", {s: round(float(w), 2) for s, w in zip(STATS, vecs[:, 0])})
print("component 1 vs points per game:", round(np.corrcoef(Z @ vecs[:, 0], ppg)[0, 1], 2))
Further reading
- Eigenvectors and eigenvalues, 3Blue1Brown. What it means for a matrix to only stretch a direction, with pictures.
- A tutorial on principal component analysis, Jonathon Shlens. A clear, step-by-step derivation of PCA from first principles.
- Principal component analysis, Wikipedia. The method, its history and its many variants.