Skip to content

One good season or a good model? Cross-validation

A single test season can flatter a model or bury it. Cross-validation tests it again and again on different slices of the data; for football, that means walking forward through the seasons. 23 SPFL seasons show why it matters.

Beginner Part 6 of Machine Learning Through Football

Contents

The football question

We've trained the model, tested it on matches it hasn't seen, and the results look pretty decent. So are we done?

No, not quite. What if we just happened to test it on a season that suited the model?

In one season, the team-strength model from training data and test data got 57.9% of match outcomes right. That sounds good. But football never sits still. Managers come and go, players leave, teams go up and down, tactics shift, and even home advantage moves about over time. One season is one set of circumstances.

The concept

Cross-validation means testing a model more than once, on different slices of the data, and looking at all the results together rather than trusting one.

The standard version is k-fold cross-validation. Shuffle the matches, cut them into k groups (often 5 or 10), then take turns: train on all the groups but one, test on the one left out, and repeat until every group has been the test set once. You get k scores instead of one.

For football, I wouldn't just shuffle 23 seasons of matches into random groups. A model predicting 2018/19 shouldn't be learning from games played years later. Time matters, and a shuffle lets the future leak into training, the trap from training data and test data.

So football uses walk-forward validation instead. Train on what we knew at the time, test on what happened next, move the window forward and go again.

K-fold takes turns with shuffled groups, so a test group can sit before matches the model trained on. Walk-forward keeps time in order: three seasons to learn from, the next one to test, then slide along a season.

A football example

The team-strength model rates each side by its points per game over the three previous Scottish Premiership seasons, and picks whichever side rates higher. Walk it forward from 2003/04 to 2025/26: 23 separate tests, each season predicted using only the three before it.

Each season's share of results called right. Same model, same method, every season: the score moves around a lot, but it stays above always picking a home win in 21 seasons of 23.
Team strength Always a home win
Average over 23 seasons 49.9% 43.5%
Best season 57.9% 51.8%
Worst season 38.6% 38.6%
Typical swing (standard deviation) 4.0 points 3.2 points

On average the model beats the home-win baseline by 6.4 percentage points, so it is adding something. It did so in 21 of the 23 seasons, which is stronger evidence than any single score.

Its best season was 57.9% (2022/23) and its worst was 38.6% (2012/13): the same approach, with a very different outcome. 2012/13 was the first season after Rangers dropped out of the top flight, and the pecking order the model had learned from the previous three seasons no longer held; that's a likely reason, though not one tested here. The home-win baseline's own worst season, also 38.6%, was 2020/21, played almost entirely behind closed doors (see home advantage).

How much of the swing is just luck?

Some of that season-to-season movement would happen even if the model's true skill never changed. A season is only 228 matches, and as in training data and test data, luck alone makes an accuracy near 50% wobble by about

$$\text{SE} = \sqrt{\frac{0.5 \times 0.5}{228}} \approx 3.3 \text{ points}$$

In plain football

  • 228 is the number of matches in a Premiership season.
  • 3.3 points is how much a model's score would typically move from season to season from luck alone, even if nothing about the model or the league changed.
  • The model's real swing is 4.0 points, only a little more. Most of the difference between a 58% season and a 48% season is the luck of which matches went which way.

That's the real case for cross-validation. One season's score is mostly signal plus a big dose of luck. Average over 23 seasons and the luck largely cancels out: the 6.4-point gap over the baseline is good to about ±1.6 points. We can trust that number in a way we can't trust 57.9%.

Why it matters

My first question, when someone tells me their model is 58% accurate, is: 58% on what? One season? One test set? Or consistently across lots of different periods?

"The model hit 57.9% in one season" and "across 23 seasons it averaged 49.9%, and moved around quite a bit from one year to the next" are very different statements. The second one isn't as flashy, but it tells you far more about the model.

  • One split can flatter or bury a model. Tested only on 2022/23, team strength looks excellent; tested only on 2012/13, it looks no better than guessing.
  • Consistency is the evidence. Beating the baseline in 21 seasons of 23 says far more than the best season.
  • In football, keep time in order. Walk forward, never shuffle: the model must only ever learn from the past.
  • Cross-validation is for choosing, too. Comparing two models, or picking how complex to make one, should use the average over many folds, not one lucky split.

Limitations

  • Walk-forward tests are few. 23 seasons is 23 scores; that's plenty for an average, but the spread is still rough.
  • Old seasons may not look like new ones. A model that did well in 2005 may matter less for 2027. Some analysts weight recent folds more.
  • The window is a choice. Three training seasons suits this model; a different window gives a different set of scores.
  • Accuracy may not be the right score at all. It only asks whether the favourite won; it ignores how confident the model was.

With football, one good Saturday doesn't tell you much. I want to know what happens when Saturday keeps coming round.

Try it yourself

Pick a simple rule for your league, such as "the side higher in the table wins". Check how often it was right in each of the last five seasons, not just the last one. What's the best season, the worst, and the average? Which would you quote if you were selling the rule, and which is the truth?

Reproduce the analysis

The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. Then:

import csv
from collections import Counter
from math import sqrt
from statistics import mean, pstdev, stdev

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}

def season(s):
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        return [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]

def team_strength(train):  # points per game in training; new teams get the average
    pts, games = Counter(), Counter()
    for r in train:
        for team, p in zip((r["HomeTeam"], r["AwayTeam"]), POINTS[r["FTR"]]):
            pts[team] += p
            games[team] += 1
    avg = sum(pts.values()) / sum(games.values())
    ppg = lambda t: pts[t] / games[t] if games[t] else avg
    return lambda r: "H" if ppg(r["HomeTeam"]) >= ppg(r["AwayTeam"]) else "A"

def accuracy(predict, rows):
    return sum(predict(r) == r["FTR"] for r in rows) / len(rows)

# walk forward: train on the three seasons before, test on the next, then move along a season
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
data = {s: season(s) for s in names}
model, home = [], []
for i in range(3, len(names)):
    predict = team_strength(data[names[i - 3]] + data[names[i - 2]] + data[names[i - 1]])
    test = data[names[i]]
    model.append(accuracy(predict, test))
    home.append(accuracy(lambda r: "H", test))
    print(f"20{names[i][:2]}/{names[i][2:]}  {len(test)} matches  team strength {model[-1]:.1%}  home win {home[-1]:.1%}")

gap = [m - h for m, h in zip(model, home)]
for name, x in (("team strength", model), ("home win", home), ("gap", gap)):
    print(f"{name:13} mean {mean(x):.1%}  sd {pstdev(x):.1%}  best {max(x):.1%}  worst {min(x):.1%}")
print("seasons the model beat a home win:", sum(g > 0 for g in gap), "of", len(gap))
print(f"luck alone, 228 matches at 50%: sd {sqrt(0.25 / 228):.1%}")
print(f"margin on the 23-season average gap: ±{1.96 * stdev(gap) / sqrt(len(gap)):.1%}")

Further reading

Get new pieces by email

An email when something new is published, and the occasional update. Unsubscribe in one click. How your email is used.