Skip to content

Has the model learned, or just memorised? Training data and test data

A model that has seen the answers will always look brilliant. Holding matches back to test it, and choosing which ones, is how you find out whether it can predict a match it hasn't seen. Real SPFL seasons show the difference.

Beginner Part 2 of Machine Learning Through Football

Contents

The football question

Say I hand a model data from 1,000 football matches and ask it to learn how to predict the result. It learns from all 1,000, then I test it on the same 1,000 and it gets almost every one right.

Has it learned anything?

Not necessarily. It's like giving someone the answers before an exam and then being impressed when they get 100%. The model has already seen every result. What we actually want to know is how it does on a match it has never seen, because that's the only kind of match it will ever be asked to predict.

This follows on from features and targets: once you've chosen what to predict and what to tell the model, you need a fair way to check it.

The concept

Split the matches before the model sees any of them:

  • Roughly 800 matches go into the training set, which is what the model learns from.
  • The other 200 become the test set. The model never sees them while it's learning.

Once it has learned from the 800, ask it to predict the 200, and compare its predictions with what really happened. The simplest score is accuracy: the share of results it got right.

The test set asks the question that matters: can it generalise, or has it just memorised? A model that has learned something real about football does about as well on the 200 as on the 800. A model that has memorised does brilliantly on the 800 and falls apart on the 200.

A football example

Real Scottish Premiership results, split the way the model would be used: train on 2022/23, 2023/24 and 2024/25 (684 matches), and test on 2025/26 (228 matches). Four models, each predicting home win, draw or away win:

  • Always pick a home win. Learns nothing. It's the baseline every model has to beat.
  • Memorise every match. Stores each training result by date and teams. For any match it hasn't stored, it guesses a home win.
  • Memorise each fixture. For every pairing, such as Celtic at home to Hearts, it predicts whatever result happened most often in training.
  • Team strength. Rates each team by its points per game in training, and picks whichever side rates higher. A team new to the league gets the average rating.
Accuracy on the three seasons each model learned from, against the season it hadn't seen. Only team strength keeps most of its training score, and it's the only one to beat always picking a home win on the test season.

The memoriser gets 100% on the training matches: it has literally stored the answers. On 2025/26 it has nothing stored, falls back on home wins every time, and scores exactly what that dumb baseline scores, 44.7%.

The fixture memoriser is sneakier. It looks like it has found something, 65.8% on training. But a handful of past meetings at one ground is mostly noise, and on the new season it does worse than always picking a home win. It learned the training data a bit too well.

Team strength is the only one that holds up: 54.1% on the seasons it learned from, 51.8% on the season it didn't. Its test score is lower, as a test score nearly always is, and that lower number is the honest one.

Why not just pick 200 matches at random?

For most data, a random 800/200 split is fine. For football, it isn't, because matches happen in order.

A random split can let information from the future leak into the training data. Suppose one test match is Celtic v Hibs in October. A random split would leave most of the rest of that season in training, including every match from November to May. The model gets to learn how strong each side turned out to be, from results that hadn't happened yet on the day it's supposedly predicting.

On Saturday morning, nobody has that information. So for football I'd train on the 2022/23, 2023/24 and 2024/25 seasons and test on 2025/26. Train on the past, test on the future, because that's how the model will actually be used.

It's the same leakage trap as a feature like the half-time score, from features and targets. There, the future sneaks in through a column. Here, it sneaks in through the split.

How big should the test set be?

2025/26 is only 228 matches, and a small test set gives a wobbly score. For an accuracy near 50%, the typical wobble, the standard error, is

$$\text{SE} = \sqrt{\frac{p\,(1 - p)}{n}} = \sqrt{\frac{0.5 \times 0.5}{228}} \approx 0.033$$

In plain football

  • p is the accuracy, about 0.5. Each prediction is right or wrong, like a shot going in or not in the Bernoulli distribution.
  • n = 228 is how many test matches there are.
  • 0.033 is 3.3 percentage points. Double it for a rough 95% range: team strength's 51.8% could really be anywhere from about 45% to 58%.

That range is wide enough to swallow the gap between team strength (51.8%) and always picking a home win (44.7%). One season can't settle it.

So test on more future seasons. Walk forward through the history: train on three seasons, test on the next, then move along a season and do it again. From 2003/04 to 2025/26 that's 23 test seasons and 5,195 matches, every one predicted using only the past:

Model Test accuracy, 5,195 matches
Always pick a home win 43.5%
Team strength 49.9%

With 5,195 matches, the wobble is about 0.7 points, so a 95% range of roughly ±1.4. A gap of 6.4 points is real. A simple rating of how good each team has been does predict results better than blindly backing the home side. Testing again and again like this is a form of cross-validation.

Why it matters

  • Training scores flatter. Every model looks better on the data it learned from. The 100% memoriser is the extreme case, but even team strength drops 2.3 points.
  • The test set is the model's first competitive match. It's the only score that says anything about next Saturday.
  • In football, time decides the split. Train on the past and test on the future, or the model gets to peek at results that hadn't happened.
  • More test matches, more trust. One season of 228 matches can't separate models a few points apart; 23 seasons can.

Limitations

  • Accuracy is a blunt score. It only asks whether the most likely result came in. A model that says 51% home win and one that says 90% get the same credit when the home side wins. Scores that grade the probabilities are in evaluating prediction models.
  • Look at the test set once. If you keep tweaking a model until its test score improves, you've quietly trained on the test set. Analysts keep a third slice, a validation set, for tweaking.
  • The future may not look like the past. A new manager, a new ball, a rule change: a model tested on last season can still be caught out by the next.
  • Draws are hard. None of these models ever predicts a draw, and draws are nearly a quarter of all results.

A model that shines on the data it learned from and struggles on anything new has a name: it's overfitted. The fixture memoriser is the textbook case. World-class in training, then the competitive matches started. It's one of the most important ideas in machine learning.

Try it yourself

Take your team's results from last season. Write down a simple rule for predicting a match, such as "we win at home against anyone who finished below us". Check how often the rule was right last season. Then use it to predict this season's matches so far, and see how often it's right now. Which number is higher, and which one would you trust?

Reproduce the analysis

The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. Then:

import csv
from collections import Counter, defaultdict

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}

def season(s):
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        return [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]

def accuracy(predict, rows):
    return sum(predict(r) == r["FTR"] for r in rows) / len(rows)

def home_win(train):
    return lambda r: "H"

def memorise_matches(train):  # every training match, by date and teams; anything new: home win
    seen = {(r["Date"], r["HomeTeam"], r["AwayTeam"]): r["FTR"] for r in train}
    return lambda r: seen.get((r["Date"], r["HomeTeam"], r["AwayTeam"]), "H")

def memorise_fixtures(train):  # most common result in each home-away pairing
    seen = defaultdict(Counter)
    for r in train:
        seen[r["HomeTeam"], r["AwayTeam"]][r["FTR"]] += 1
    return lambda r: seen[r["HomeTeam"], r["AwayTeam"]].most_common(1)[0][0] if (r["HomeTeam"], r["AwayTeam"]) in seen else "H"

def team_strength(train):  # points per game in training; new teams get the average
    pts, games = Counter(), Counter()
    for r in train:
        for team, p in zip((r["HomeTeam"], r["AwayTeam"]), POINTS[r["FTR"]]):
            pts[team] += p
            games[team] += 1
    avg = sum(pts.values()) / sum(games.values())
    ppg = lambda t: pts[t] / games[t] if games[t] else avg
    return lambda r: "H" if ppg(r["HomeTeam"]) >= ppg(r["AwayTeam"]) else "A"

train = season("2223") + season("2324") + season("2425")
test = season("2526")
for model in (home_win, memorise_matches, memorise_fixtures, team_strength):
    predict = model(train)
    print(f"{model.__name__:18} training {accuracy(predict, train):.1%}  test {accuracy(predict, test):.1%}")

# walk forward: train on three seasons, test on the next, from 2003/04 to 2025/26
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
tests = [(season(names[i - 3]) + season(names[i - 2]) + season(names[i - 1]), season(names[i])) for i in range(3, len(names))]
n = sum(len(te) for _, te in tests)
for model in (home_win, team_strength):
    right = sum(accuracy(model(tr), te) * len(te) for tr, te in tests)
    print(f"{model.__name__:18} walk-forward {right / n:.1%} of {n} matches")

Further reading

Get new pieces by email

An email when something new is published, and the occasional update. Unsubscribe in one click. How your email is used.